Get Started
Build with Pearl Inference
OpenAI-compatible inference for open-weight models, a flat 10% below Together AI, on simple prepaid credits.
Pearl Inference runs production inference for a curated set of open-weight models on the Pearl network. You bring any OpenAI SDK (or plain HTTP), point it at our base URL, and pay per token from a prepaid balance, no GPU provisioning, no contracts, no switching costs.
Get started in minutes
Quickstart
Create an API key and make your first request in a few minutes.
Playground
Chat with every model in the browser before writing any code.
Choosing a model
Pick the right model for coding, chat, vision, or reasoning.
Not sure where to start? Pick a model with the model selection guide, run the quickstart, and skim Concepts to understand keys, credits, and billing.
What you can build
Chat completions
OpenAI-compatible /chat/completions with streaming on every model.
Function calling
Connect models to your tools and APIs, supported across the catalog.
Structured outputs
JSON mode for reliable, machine-readable responses.
Reasoning models
GLM-5.2 and the DeepSeek V4 models expose their chain of thought in a separate field.
Vision
Send images alongside text with Gemma 4 31B Instruct.
Long context
Up to 1M tokens of context with DeepSeek V4 Pro.
The platform today
The platform currently serves chat-model inference through a deliberately curated catalog of 5 production models spanning reasoning, coding, general chat, and vision. Every model is served on the same OpenAI-compatible endpoint and billed per token against your organization's prepaid credits.