Variant D — model dossiers

Design thesis: the frontier asks “which models win?” — but a buyer's real question is “who is this model for?” So: one dossier card per model, built bottom-up from its actual frontiers, with exact numbers, a per-frontier cost-slash-line (dot on a 10×-spaced ruler), its nearest rival at every frontier, and a computed verdict sentence.

Each card is computed from ranked[] — no hand-written copy. † = UNVERIFIED modeled cap. on frontier · within 10% of frontier · · dominated.

Model dossiers

Default view: frontier residents and models near at least one frontier. Dominated-only models remain available in the comparison table below.

Two-model head-to-head

Pick any two models; every frontier they share is printed side by side, with deltas.

vs

All models · ranked by frontier presence

Sources: AA Intelligence Index v4.1.1 (current scale only — older articles use renormalized scales and are never mixed in) · GPQA Diamond AA-run · SWE-bench Verified via vals.ai independent same-harness (Mini-SWE-agent); openlm.ai rows are aggregator/vendor-reported · prices: OpenCode Go + ClinePass published tables and the live OpenRouter catalog · workload shapes: agent-chat from real 30-day traces, coding-session from OpenCode Go published request patterns. † caps are UNVERIFIED models: ClinePass assumed $35; $20-tier subscriptions modeled at 8× purchase price ($160) per community reporting. Y-axis labels follow each frontier's per-row secondary_label (note: math-smarts-coding's axes string says SWE-bench-Verified but its rows carry GPQA-Diamond values — the label shown is the authoritative one). Models without verified scores are excluded-with-reason (muse-glimmer; GLM-5.3/Flash and muse-spark lack verifiable GPQA — notable because GLM-5.3 leads the general index). ≈-near is computed client-side as: within 10% of the nearest strictly-better-or-tie row on the frontier's primary axis (cost or speed); it reproduces the generator's (+X% off) chips in 13 of 14 cases — the one divergence is qwen3-30b-a3b-2507 on chat-metered-responsiveness, where the generator reports +8% off and this rule finds no dominator within 10% of its speed.