NewClaude Opus 5.5 is now 90% below list

The same AI models.
Up to 90% cheaper.

Claude, GPT, Gemini and 300+ more, served from spare capacity at a fraction of list price. Same weights, same outputs, one OpenAI‑compatible API.

Claude Opus 5.5
$2.50$25.0090% off
GPT-5
$1.00$10.0090% off
Gemini 2.5 Pro
$1.50$10.0085% off
Output price per 1M tokens, zurelay vs. provider list price

Requests

JM
Requests · 24h+12.4%
1,284,309
Median latency−18 ms
212 ms
AllFallbacks 3Errors 0Last 15 min
ModelRouted toLatency
Claude Opus 5.5Vertex468 ms
GPT-5Azure388 ms
Claude Sonnet 5.5Bedrock412 ms
Gemini 2.5 ProVertex534 ms
Claude Opus 5.5Bedrock452 ms
DeepSeek V3.2Fireworks174 ms
GPT-5Azure402 ms
Kimi K2DeepInfra263 ms
tokens routed dailySeptember 2026
2.4B
average saving vs. list priceLast 30 days
84%
median routing overheadLast 30 days
14ms
API uptimeLast 90 days
99.99%

Where the 90% comes from.

zurelay never swaps your model for a cheaper one. The savings come from how the same model is bought, routed and cached.

Claude Opus 5.5Output, per 1M tokens
  1. List price

    $25.00

    What the model costs when you buy it directly.

  2. Spare capacity, bought in bulk

    $5.00

    Clouds and labs reserve more GPU time than they use. zurelay buys those idle hours far below list and serves your requests on them.

  3. The cheapest healthy route

    $3.25

    Prices move by the minute. Every request goes to whichever provider is cheapest right now, within your latency and data rules.

  4. Caching that pays you back

    $2.50

    Repeated prompts and shared prefixes are cached automatically, so you never pay twice for the same tokens.

You pay $2.50 instead of $25.0090% less, same model

Estimate your savings

Mostly using
Today$20,000/mo
With zurelay$2,000/mo

You keep

$216,000 a year

Start saving

Estimate uses the rates in the price table. Your mix may vary.

How routing works

One request in. The best route out.

zurelay sits between your app and every provider. It knows who is fast, who is cheap and who is down right now, not last week.

  1. Score

    Every provider serving your model is ranked on live price, latency and error rate, refreshed every few seconds.

  2. Route

    Your request goes to the winner. Pin providers, regions or data policies, and zurelay only picks from routes that qualify.

  3. Recover

    If a provider errors or stalls, zurelay retries on the next-best route before your user notices anything.

Optimize for
zurelay
Anthropic
$3.00/M310 ms
Bedrock
$2.75/M355 ms
VertexRouted
$2.50/M420 ms
Azure
$2.90/M390 ms
Cheapest healthy routevertex/us-east5420 ms90% under list

Pay a fraction for the exact same model.

Every price below is what you pay through zurelay, next to the provider’s list price. Same weights and same outputs, served from spare capacity.

  • Claude Opus 5.5
    anthropic/claude-opus-5.5
    $0.50 / $2.50
    $5.00 / $25.00
    460 ms
    90%
  • GPT-5
    openai/gpt-5
    $0.13 / $1.00
    $1.25 / $10.00
    380 ms
    90%
  • Claude Sonnet 5.5
    anthropic/claude-sonnet-5.5
    $0.45 / $2.25
    $3.00 / $15.00
    410 ms
    85%
  • Gemini 2.5 Pro
    google/gemini-2.5-pro
    $0.19 / $1.50
    $1.25 / $10.00
    520 ms
    85%
  • GPT Image 1
    openai/gpt-image-1
    $0.006 /image
    $0.042
    6.2 s
    85%
  • Veo 3
    google/veo-3
    $0.060 /second
    $0.40
    48 s
    85%
  • Grok 4
    xai/grok-4
    $0.48 / $2.40
    $3.00 / $15.00
    430 ms
    84%
Illustrative rates, per 1M tokens (input / output) unless noted. 7 of 17 shown.Browse all 300+ models

Change one line. Keep everything else.

zurelay speaks the OpenAI API, so your SDK, prompts, tools and streaming code stay exactly as they are. Swap the base URL and every model is one string away.

Works with what you already use

  • Streaming
  • Tool calling
  • Structured outputs
  • Vision & audio
  • Prompt caching
  • Batch jobs
import OpenAI from "openai";
const client = new OpenAI({
− baseURL: "https://api.openai.com/v1",
+ baseURL: "https://api.zurelay.com/v1",
apiKey: process.env.ZURELAY_API_KEY,
});
const stream = await client.chat.completions.create({
model: "anthropic/claude-opus-5.5",
messages: [{ role: "user", content: "Summarize this PR." }],
stream: true,
});
Response headers200 OK
x-zurelay-routevertex/us-east5
x-zurelay-latency184 ms
x-zurelay-saved$0.0549 (90%)

When a provider goes down, your users never find out.

Failover, data boundaries, budgets and observability ship on every plan. No sidecars, no extra dashboards.

req_8f2c…a91 · POST /v1/chat/completions253 ms · 200 OK
0100200300
Anthropic
503
Bedrock
200 · first token 182 ms
Your app
one clean response

Automatic fallbacks

When a provider errors, times out or rate-limits, zurelay retries on the next-best route. Your app gets one clean response.

Enforce zero retention
  • Bedrockus-east-1eligible
  • Vertexeurope-west4eligible
  • OpenAI30-day retentionexcluded

212 of 318 routes qualify

Zero data retention

Flip one switch and zurelay only routes to providers contractually bound not to store prompts or outputs.

prod-web$412 / $1,000
eval-batch$948 / $1,000
staging$38 / $100

eval-batch hit 95%. Alert sent to #infra.

Spend limits per key

Hard caps and alerts for every key, so an eval loop can never eat the production budget.

EU only
  • eu-west-1Dublin18 ms
  • eu-central-1Frankfurt23 ms
  • us-east-1blocked
  • ap-south-1blocked

Region pinning

Keep traffic inside the EU, the US or any region list you define, enforced per key.

p50 latency · 24h212 ms now
24h agonow

Observability built in

Latency, cost and routing decisions for every request, exportable to your own stack via OpenTelemetry.

Route your first request
in five minutes.

Get a key, change one URL, and pay a fraction of list price from the very first token.