Get Started
Build on Pearl Inference Platform
OpenAI-compatible inference supporting the latest open-weight models, get started in minutes
Pearl Inference runs production inference for a curated set of open-weight models on the Pearl network. You bring any OpenAI SDK (or plain HTTP), point it at our base URL, and pay per token from a prepaid balance, no GPU provisioning, no contracts, no switching costs.
Get started in minutes
Quickstart
Create an API key and make your first request in a few minutes.
Playground
Chat with every model in the browser before writing any code.
Choosing a model
Pick the right model for coding, chat, vision, or reasoning.
Not sure where to start? Pick a model with the model selection guide, run the quickstart, and skim Concepts to understand keys, credits, and billing.
What you can build
Chat completions
OpenAI-compatible /chat/completions with streaming on every model.
Function calling
Connect models to your tools and APIs, supported across the catalog.
Structured outputs
JSON mode for reliable, machine-readable responses.
Reasoning models
GLM-5.3 and the DeepSeek V4 models expose their chain of thought in a separate field.
Vision
Image to text with GLM-5.3 Flash.
Long context
Up to 1M tokens of context with DeepSeek and GLM models.
The platform today
The platform currently serves chat-model inference through a deliberately curated catalog of latest open weights models spanning reasoning, coding, general chat, and vision. Every model is served on the same OpenAI-compatible endpoint and billed per token against your organization's prepaid credits.