The managed inference platform for open models.

Run open-weight models, your fine-tunes, and your distillations in production, without operating the infrastructure. Low-latency realtime inference on European GPUs, behind a single OpenAI-compatible API. Residency without the latency tax. Flex and async windows are there for background work that isn't blocking a user. Serverless per-token to dedicated deployments; every request logged and auditable, every model pinned.

Trusted by

01The engine

The fastest and most efficient way to run open models.

An async scheduler with real queues and backpressure at the core. Every request is recorded and dispatched transactionally, routed over per-model queues, and executed by workers that drive vLLM and SGLang as in-process libraries, not model servers behind a load balancer. The same design that keeps latency low keeps the GPUs full: performance and efficiency aren't a trade-off here.

Measured on live gateway traffic

Not just a design claim. Requesty's Model Discovery analytics compare every provider behind their gateway on real traffic. These are their numbers, not ours, as of mid-August 2026.

kimi-k3
76 tok/sthroughput
426 ms average latency
Next-fastest provider: 36 tok/s
deepseek-v4-flash
157 tok/sthroughput
861 ms average latency
Next-fastest provider: 111 tok/s
glm-5.2
211 tok/sthroughput
3.1 s average latency
Next-fastest provider: 131 tok/s
02The platform

Everything between your model and production.

One platform for all of your inference, open models or your own, tuned first for the traffic where a user is waiting. A scheduler and hardware-agnostic runtime built from the ground up, not a wrapper around someone else's cloud. You bring the workload (and optionally the weights); we handle serving, scaling, versioning, logging, and the audit trail.

Any open model, including yours

Catalog models, fine-tunes, and distillations on one platform.

Serve open-weight models: GLM, DeepSeek V4 Flash, ThinkingCap, Qwen3.6 35B, Qwen3-VL, and Kimi K3 today, with DeepSeek V4 Pro, Gemma 4, GPT-OSS, Nemotron, and Qwen3.8 next. Or upload your own weights. Fine-tunes and distilled models run on the same API, the same infrastructure, and the same audit trail as catalog models. vLLM and SGLang models run out of the box; for anything else, we stand up the runtime.

Realtime first

Tuned for the request a user is waiting on.

Most usage runs on realtime sync endpoints (`/v1/chat/completions`, `/v1/responses`, `/v1/messages`), and that is what the scheduler is tuned for. Independent gateway data puts us at the highest throughput of any provider on every model we serve there. Background work has somewhere to go too: a 24h async window for batch and background jobs, and `service_tier: "flex"` for discounted tokens at lower priority. One OpenAI-compatible API for all of it; most clients switch with a base-URL change.

Production discipline

Pinned versions, named on every request.

Closed-API providers move the model under a stable name, and the same prompt returns different output next week. On sference you address an exact version rather than a "latest" alias, and every response records the version that served it, so the checkpoint behind a workload is always something you can point at. Rerun your evals against a newer version before you move to it. Your own models work the same way.

European by architecture

Sovereign infrastructure, without the latency penalty.

Every request runs on European GPUs, outside US CLOUD Act jurisdiction, architecturally rather than by policy. EU residency is usually assumed to cost you speed; measured on live gateway traffic, the European route is the fast one. Request-level logging and a full audit trail, configurable retention; GDPR, DORA, and EU AI Act ready, DPA included. If your customers ask where their data goes, you have a real answer.

03Use cases

What teams run on sference.

Coding assistants your team uses internally, and agent-powered features your customers touch. Both run on open models with full per-request provenance, and your own fine-tunes slot into either.

01Workload

Coding and internal assistants

Copilots, code review, and coding agents on open models, alongside the internal chat surfaces your team already runs, like Open WebUI. Streaming, tool calls, and your own fine-tunes on the hot path.

IDE, CI, and internal chat·realtime·per-token pricing
02Workload

Features powered by agents

Product features built on interactive and background agents. The turn a user waits on runs realtime; the steps behind it run on async windows, under one API and one audit trail.

Interactive + background·realtime + async·per-token pricing
Built for regulated verticals
FinTechLegalTechHealthTechInsureTechAI/ML teams

One integration brings thousands of end-users through your API, with audit trails their security officer can verify.

04Onboarding

Moving to open models is a project. We staff it with you.

Teams don't stay on closed APIs because they love the pricing. They stay because migration feels risky. Our forward-deployed engineers de-risk it: evals on your traffic, model selection, and migration support until the workload runs in production.

1

Benchmark on your real workload

We run your actual prompts across candidate open models (and your fine-tunes, if you have them) side by side against your current provider. You get an eval report with quality, latency, and cost per workload, not a leaderboard screenshot.

2

Migrate without a rewrite

The API is OpenAI-compatible, so most clients switch with a base-URL change. Our engineers work with your team on prompt adjustments, structured-output schemas, and SDK integration into your existing pipelines.

3

Optimize as you scale

Right-size the model per workload, route background jobs to async windows, and move hot paths onto dedicated capacity as volume grows. Pinned versions mean upgrades happen on your schedule, validated by reruns of the same evals.

05Pricing

Start serverless. Scale to dedicated.

Per-token pricing with no minimums to get started, reserved capacity when a workload earns it, and enterprise agreements when procurement gets involved. Our scheduler keeps GPUs busy around the clock, so you pay for tokens, not for the idle capacity that self-managed instances and on-prem clusters quietly bill you for.

Serverless

Pay per token.

Shared endpoints across the full catalog. No minimum spend, no upfront commitment. Realtime by default, with flex and a 24h async window for batch and background work on the same endpoints.

Dedicated

Reserved capacity, private endpoints.

Your models on GPUs reserved for you: predictable latency, guaranteed throughput, no noisy neighbors. Billed per GPU-hour on monthly commitments, the right shape once a workload runs around the clock.

Enterprise

Annual agreements, custom terms.

Committed-spend discounts, volume pricing, custom SLAs, custom model hosting (fine-tunes and distillations), security review support, DPA and procurement-friendly terms. Includes forward-deployed onboarding: our engineers benchmark and migrate your workloads with your team.

Serverless rates · per 1M tokens
ModelInputOutputCached
Kimi K3$3.00$15.00$0.45
GLM 5.2$1.20$4.20$0.26
GLM 5.3$1.20$4.20$0.26
DeepSeek V4.1 Flash$0.50$1.50$0.05
GLM 5.3 Flash$0.20$0.60$0.07
DeepSeek V4 Flash (0731)$0.28$0.56$0.07
DeepSeek V4 Procoming soon
Gemma 4 31Bcoming soon
NVIDIA Nemotron 3 Supercoming soon
OpenAI GPT-OSS 120Bcoming soon
Qwen3.8coming soon

Serverless rates at realtime priority, confirmed per model when your account is provisioned.

06The European alternative

Compliance answers you can hand to a security review.

We run on European infrastructure, which makes some questions easy: request-level logging, a full audit trail, exportable reports, DPA included. When your customer's compliance officer asks, give them a dashboard link, not a “we take security seriously” PDF.

Processed in Europe

Guaranteed architecturally, not by policy. Outside US CLOUD Act jurisdiction.

DORA & AI Act aligned

Deployer obligations begin August 2026. Transparency and traceability baked in.

BYOM with the same guarantees

Your fine-tune runs under the same audit trail as any catalog model.

Zero training on your data

Ever. Contractually and technically.

EU AI Act readyGDPRDORA readyEuropean data residencyDPA included
07FAQ

Things engineers and CTOs ask.

08Get started

Your models, running in production this week.

Spin up an account and make your first API call.