Most models are graded on how good their answers sound. Alloy is built for the harder standard: being correct. On tasks with a single checkable answer, Alloy matched or beat the strongest frontier models we tested — while completing every single request.
Latest full evaluation: 2026-07-07 · every result below was measured against the same live production API our customers use.
25 tasks with one objectively correct answer — counting, numeric comparison traps, predicting code output, calendar math, probability — graded by exact match against the live API. No partial credit, no retries.
16 open-ended prompts across facts, reasoning, code, and writing — each answer compared blind against claude-opus-4.8 by two independent grader models with answer positions swapped to cancel ordering bias. 0.50 is parity with the frontier baseline; above it means graders preferred the Alloy answer.
Alloy answers are reasoned through, not blurted. The extra seconds buy the accuracy above — and stay well inside interactive range.
| Model | Median response | p95 | Requests completed |
|---|---|---|---|
| alloy · low thinking | 12.9s | 34.3s | 100% |
| alloy · medium thinking | 13.6s | 94.7s | 100% |
| alloy · code sessions | 15.2s | 28.2s | 100% |
| alloy-pro · medium thinking | 14.7s | 33.4s | 100% |
| alloy-pro · high thinking | 19.4s | 49.7s | 100% |
| claude-opus-4.8 baseline | 3.7s | 6.2s | 100% |
| gpt-5.5 baseline | 4.1s | 14.9s | 100% |
Frontier APIs bill by the token — long answers cost more, and multi-pass quality costs multiples. Atlas is flat: a request is a request, and the reasoning compute behind an Alloy answer is on us, not on your bill.
200 requests every 7 days, full model catalog access. Try Alloy on real work before paying anything.
3,000 requests per rolling 7-day window — up to ~12,800 answers a month. Every one of them can be a full Alloy deliberation at no extra charge.
10,000 requests per rolling 7-day window — up to ~43,000 answers a month. Frontier-grade output at a per-answer cost token pricing can't touch.
Both evaluations run against the live production API — the exact endpoint customers use, with no special routing. Verifiable tasks are graded by exact match; a timeout counts as a failure. Open-ended quality uses blind pairwise grading: two independent grader models score each comparison with answer positions swapped to cancel ordering bias. Baselines (claude-opus-4.8, gpt-5.5) are public frontier models served through this gateway, and their results are published unchanged — including where they lead. We re-run and refresh this page periodically. What we don't publish: internal system architecture, prompt text, or any customer data — results only.
Figures from the 2026-07-07 evaluation run. Questions? See the API docs or model catalog.