Hosted and gateway models
Both use the same Prime API key and endpoint, and you select a model by its catalog ID. Pricing, limits, and supported features are set per model and shown in the catalog, so check them before choosing a model, especially if where a model runs matters for your workload.
Get started
- Create an API key with the Inference permission in the API keys guide.
- Set
PRIME_API_KEYin your environment. Keep it on the server, not in browser code or a public repository. - Send a chat completion with a model ID from the catalog. The example below uses GLM-5.3.
OpenAI Python SDK
GET https://api.pinference.ai/api/v1/models or run prime inference models. Use the returned model ID exactly as shown. See the Models API and Chat Completions API for endpoint details.
Team billing
Requests bill the API key owner’s personal account unless you includeX-Prime-Team-ID. To use team credits, obtain your team ID from the team profile or prime teams list, and add the header to each direct API request:
default_headers={"X-Prime-Team-ID": "your-team-id"} on the client. prime config set-team-id applies to CLI operations, not to direct HTTP or SDK requests. Check usage and balances in the billing dashboard. If a request returns insufficient_funds, check the balance of the account selected by the key and optional team header.
Next steps
- GLM-5.3: use a model hosted on Prime Inference.
- LoRA adapter deployments: serve adapters from Hosted Training runs.
- Troubleshooting: resolve billing errors.
- Evaluate models with Prime Lab.