Skip to main content
Prime Inference provides one OpenAI-compatible API for both hosted and gateway models. Check the model catalog for current availability, serving type, and pricing; they vary by model.

Hosted and gateway models

Both use the same Prime API key and endpoint, and you select a model by its catalog ID. Pricing, limits, and supported features are set per model and shown in the catalog, so check them before choosing a model, especially if where a model runs matters for your workload.

Get started

  1. Create an API key with the Inference permission in the API keys guide.
  2. Set PRIME_API_KEY in your environment. Keep it on the server, not in browser code or a public repository.
  3. Send a chat completion with a model ID from the catalog. The example below uses GLM-5.3.

OpenAI Python SDK

To list model IDs programmatically, call GET https://api.pinference.ai/api/v1/models or run prime inference models. Use the returned model ID exactly as shown. See the Models API and Chat Completions API for endpoint details.

Team billing

Requests bill the API key owner’s personal account unless you include X-Prime-Team-ID. To use team credits, obtain your team ID from the team profile or prime teams list, and add the header to each direct API request:
With the OpenAI SDK, set default_headers={"X-Prime-Team-ID": "your-team-id"} on the client. prime config set-team-id applies to CLI operations, not to direct HTTP or SDK requests. Check usage and balances in the billing dashboard. If a request returns insufficient_funds, check the balance of the account selected by the key and optional team header.

Next steps