Design thesis: don't show the reader a frontier — let them push on it. Pick a workload, a budget, and a taste for smarts, and the page computes your personal pareto pick from the same ranked[] rows, live, showing what you'd have to give up to buy the next intelligence step.
All selection runs client-side over the printed frontier data. † = UNVERIFIED modeled cap; sub offerings are priced per-turn from published tables (capped purchases modeled at 8×). No backend, no inference — the frontier is the interface.
1 · What's your budget per turn?
Costs below $0.01 are sub-cent; the exact box takes full precision. Note: OpenRouter costs are listed per turn under the agent-chat shape (blended 30-day traces), not per token — that is the honest unit here.
2 · Which workload?
3 · What's your taste for smarts?
"Value" picks by intelligence-per-dollar; the log variant is for when a 200× cost spread makes raw ratios meaningless.
4 · Your pick:
| offering | x | y | status at your budget | dominated by |
|---|
"status" is recomputed at your budget ceiling: reachable = cost ≤ budget and not dominated at the current ceiling; dominated-by is the same list the frontier reports.
5 · What the next intelligence step would cost you
Sources: AA Intelligence Index v4.1.1 (current scale only — older articles use
renormalized scales and are never mixed in) · GPQA Diamond AA-run · SWE-bench Verified via
vals.ai independent same-harness (Mini-SWE-agent); openlm.ai rows are aggregator/vendor-reported ·
prices: OpenCode Go + ClinePass published tables and the live OpenRouter catalog ·
workload shapes: agent-chat from real 30-day traces, coding-session from OpenCode Go published
request patterns. † caps are UNVERIFIED models: ClinePass assumed $35; $20-tier subscriptions
modeled at 8× purchase price ($160) per community reporting. Y-axis labels follow each frontier's
per-row secondary_label (note: math-smarts-coding's axes string says SWE-bench-Verified
but its rows carry GPQA-Diamond values — the label shown is the authoritative one). Models without
verified scores are excluded-with-reason (muse-glimmer; GLM-5.3/Flash and muse-spark lack
verifiable GPQA — notable because GLM-5.3 leads the general index). ≈-near is computed client-side
as: within 10% of the nearest strictly-better-or-tie row on the frontier's primary axis (cost or
speed); it reproduces the generator's (+X% off) chips in 13 of 14 cases — the one divergence is
qwen3-30b-a3b-2507 on chat-metered-responsiveness, where the generator reports +8% off and this
rule finds no dominator within 10% of its speed.