The managed inference platform for open models.
Run open-weight models, your fine-tunes, and your distillations in production, without operating the infrastructure. Low-latency realtime inference on European GPUs, behind a single OpenAI-compatible API. Residency without the latency tax. Flex and async windows are there for background work that isn't blocking a user. Serverless per-token to dedicated deployments; every request logged and auditable, every model pinned.
Trusted by
The fastest and most efficient way to run open models.
An async scheduler with real queues and backpressure at the core. Every request is recorded and dispatched transactionally, routed over per-model queues, and executed by workers that drive vLLM and SGLang as in-process libraries, not model servers behind a load balancer. The same design that keeps latency low keeps the GPUs full: performance and efficiency aren't a trade-off here.
Not just a design claim. Requesty's Model Discovery analytics compare every provider behind their gateway on real traffic. These are their numbers, not ours, as of mid-August 2026.
Everything between your model and production.
One platform for all of your inference, open models or your own, tuned first for the traffic where a user is waiting. A scheduler and hardware-agnostic runtime built from the ground up, not a wrapper around someone else's cloud. You bring the workload (and optionally the weights); we handle serving, scaling, versioning, logging, and the audit trail.
Catalog models, fine-tunes, and distillations on one platform.
Serve open-weight models: GLM, DeepSeek V4 Flash, ThinkingCap, Qwen3.6 35B, Qwen3-VL, and Kimi K3 today, with DeepSeek V4 Pro, Gemma 4, GPT-OSS, Nemotron, and Qwen3.8 next. Or upload your own weights. Fine-tunes and distilled models run on the same API, the same infrastructure, and the same audit trail as catalog models. vLLM and SGLang models run out of the box; for anything else, we stand up the runtime.
Tuned for the request a user is waiting on.
Most usage runs on realtime sync endpoints (`/v1/chat/completions`, `/v1/responses`, `/v1/messages`), and that is what the scheduler is tuned for. Independent gateway data puts us at the highest throughput of any provider on every model we serve there. Background work has somewhere to go too: a 24h async window for batch and background jobs, and `service_tier: "flex"` for discounted tokens at lower priority. One OpenAI-compatible API for all of it; most clients switch with a base-URL change.
Pinned versions, named on every request.
Closed-API providers move the model under a stable name, and the same prompt returns different output next week. On sference you address an exact version rather than a "latest" alias, and every response records the version that served it, so the checkpoint behind a workload is always something you can point at. Rerun your evals against a newer version before you move to it. Your own models work the same way.
Sovereign infrastructure, without the latency penalty.
Every request runs on European GPUs, outside US CLOUD Act jurisdiction, architecturally rather than by policy. EU residency is usually assumed to cost you speed; measured on live gateway traffic, the European route is the fast one. Request-level logging and a full audit trail, configurable retention; GDPR, DORA, and EU AI Act ready, DPA included. If your customers ask where their data goes, you have a real answer.
What teams run on sference.
Coding assistants your team uses internally, and agent-powered features your customers touch. Both run on open models with full per-request provenance, and your own fine-tunes slot into either.
Coding and internal assistants
Copilots, code review, and coding agents on open models, alongside the internal chat surfaces your team already runs, like Open WebUI. Streaming, tool calls, and your own fine-tunes on the hot path.
Features powered by agents
Product features built on interactive and background agents. The turn a user waits on runs realtime; the steps behind it run on async windows, under one API and one audit trail.
One integration brings thousands of end-users through your API, with audit trails their security officer can verify.
Moving to open models is a project. We staff it with you.
Teams don't stay on closed APIs because they love the pricing. They stay because migration feels risky. Our forward-deployed engineers de-risk it: evals on your traffic, model selection, and migration support until the workload runs in production.
Benchmark on your real workload
We run your actual prompts across candidate open models (and your fine-tunes, if you have them) side by side against your current provider. You get an eval report with quality, latency, and cost per workload, not a leaderboard screenshot.
Migrate without a rewrite
The API is OpenAI-compatible, so most clients switch with a base-URL change. Our engineers work with your team on prompt adjustments, structured-output schemas, and SDK integration into your existing pipelines.
Optimize as you scale
Right-size the model per workload, route background jobs to async windows, and move hot paths onto dedicated capacity as volume grows. Pinned versions mean upgrades happen on your schedule, validated by reruns of the same evals.
Start serverless. Scale to dedicated.
Per-token pricing with no minimums to get started, reserved capacity when a workload earns it, and enterprise agreements when procurement gets involved. Our scheduler keeps GPUs busy around the clock, so you pay for tokens, not for the idle capacity that self-managed instances and on-prem clusters quietly bill you for.
Pay per token.
Shared endpoints across the full catalog. No minimum spend, no upfront commitment. Realtime by default, with flex and a 24h async window for batch and background work on the same endpoints.
Reserved capacity, private endpoints.
Your models on GPUs reserved for you: predictable latency, guaranteed throughput, no noisy neighbors. Billed per GPU-hour on monthly commitments, the right shape once a workload runs around the clock.
Annual agreements, custom terms.
Committed-spend discounts, volume pricing, custom SLAs, custom model hosting (fine-tunes and distillations), security review support, DPA and procurement-friendly terms. Includes forward-deployed onboarding: our engineers benchmark and migrate your workloads with your team.
| Model | Input | Output | Cached |
|---|---|---|---|
| Kimi K3 | $3.00 | $15.00 | $0.45 |
| GLM 5.2 | $1.20 | $4.20 | $0.26 |
| GLM 5.3 | $1.20 | $4.20 | $0.26 |
| DeepSeek V4.1 Flash | $0.50 | $1.50 | $0.05 |
| GLM 5.3 Flash | $0.20 | $0.60 | $0.07 |
| DeepSeek V4 Flash (0731) | $0.28 | $0.56 | $0.07 |
| DeepSeek V4 Procoming soon | — | — | — |
| Gemma 4 31Bcoming soon | — | — | — |
| NVIDIA Nemotron 3 Supercoming soon | — | — | — |
| OpenAI GPT-OSS 120Bcoming soon | — | — | — |
| Qwen3.8coming soon | — | — | — |
Serverless rates at realtime priority, confirmed per model when your account is provisioned.
Compliance answers you can hand to a security review.
We run on European infrastructure, which makes some questions easy: request-level logging, a full audit trail, exportable reports, DPA included. When your customer's compliance officer asks, give them a dashboard link, not a “we take security seriously” PDF.
Processed in Europe
Guaranteed architecturally, not by policy. Outside US CLOUD Act jurisdiction.
DORA & AI Act aligned
Deployer obligations begin August 2026. Transparency and traceability baked in.
BYOM with the same guarantees
Your fine-tune runs under the same audit trail as any catalog model.
Zero training on your data
Ever. Contractually and technically.
Things engineers and CTOs ask.
Your models, running in production this week.
Spin up an account and make your first API call.