← Atlas Gateway
Atlas Alloy · Benchmark Report

Engineered to be right.

Most models are graded on how good their answers sound. Alloy is built for the harder standard: being correct. On tasks with a single checkable answer, Alloy matched or beat the strongest frontier models we tested — while completing every single request.

Latest full evaluation: 2026-07-07 · every result below was measured against the same live production API our customers use.

100%
verifiable-task accuracy — Alloy (low & medium thinking) and Alloy Pro (high thinking), exact-match graded
287/287
benchmark requests completed. Zero failures, zero excluded runs.
< $0.001
effective cost per answer on paid plans — flat pricing, no token metering

Accuracy where it can be proven

25 tasks with one objectively correct answer — counting, numeric comparison traps, predicting code output, calendar math, probability — graded by exact match against the live API. No partial credit, no retries.

alloy · low thinking perfect
100%
alloy · medium thinking perfect
100%
alloy-pro · high thinking perfect
100%
alloy · code sessions
96%
alloy-pro · medium thinking
96%
gpt-5.5
100%
claude-opus-4.8
96%
Bar scale 90–100% to make differences visible. Baselines in grey are solo frontier models, also served through the gateway.

Quality against the strongest solo frontier model

16 open-ended prompts across facts, reasoning, code, and writing — each answer compared blind against claude-opus-4.8 by two independent grader models with answer positions swapped to cancel ordering bias. 0.50 is parity with the frontier baseline; above it means graders preferred the Alloy answer.

alloy-pro · medium thinking
0.56
alloy · medium thinking
0.55
alloy-pro · high thinking
0.55
alloy · code sessions
0.52
alloy · low thinking speed-tuned
0.40
gpt-5.5
0.58
Scale 0–0.60. Low thinking deliberately trades open-ended polish for speed — it still holds 100% verifiable accuracy above. For maximum answer quality, use alloy at medium thinking or alloy-pro.

Speed, in context

Alloy answers are reasoned through, not blurted. The extra seconds buy the accuracy above — and stay well inside interactive range.

ModelMedian responsep95Requests completed
alloy · low thinking12.9s34.3s100%
alloy · medium thinking13.6s94.7s100%
alloy · code sessions15.2s28.2s100%
alloy-pro · medium thinking14.7s33.4s100%
alloy-pro · high thinking19.4s49.7s100%
claude-opus-4.8 baseline3.7s6.2s100%
gpt-5.5 baseline4.1s14.9s100%

Performance per dollar

Frontier APIs bill by the token — long answers cost more, and multi-pass quality costs multiples. Atlas is flat: a request is a request, and the reasoning compute behind an Alloy answer is on us, not on your bill.

Free
$0 / month

200 requests every 7 days, full model catalog access. Try Alloy on real work before paying anything.

Pro · $9 / month
< $0.001 / answer

3,000 requests per rolling 7-day window — up to ~12,800 answers a month. Every one of them can be a full Alloy deliberation at no extra charge.

Premium · $29 / month
< $0.0007 / answer

10,000 requests per rolling 7-day window — up to ~43,000 answers a month. Frontier-grade output at a per-answer cost token pricing can't touch.

Reliability

0
failed or errored requests across the full evaluation (287 requests)
112/112
open-ended generations completed
175/175
verifiable-task runs completed

Methodology

Both evaluations run against the live production API — the exact endpoint customers use, with no special routing. Verifiable tasks are graded by exact match; a timeout counts as a failure. Open-ended quality uses blind pairwise grading: two independent grader models score each comparison with answer positions swapped to cancel ordering bias. Baselines (claude-opus-4.8, gpt-5.5) are public frontier models served through this gateway, and their results are published unchanged — including where they lead. We re-run and refresh this page periodically. What we don't publish: internal system architecture, prompt text, or any customer data — results only.

Start free See pricing Read the docs

Figures from the 2026-07-07 evaluation run. Questions? See the API docs or model catalog.