Item response theory analysis
2PL IRT fit separately per benchmark.
θ, difficulty (b), and discrimination (a, modeled as
log(a) ∼ Normal(0, 0.5)) are fit jointly via MCMC, not point-estimated separately.
Diagnostics (divergences, r̂, ESS) were checked for every
benchmark before trusting its output — worst case so far: 0 divergences,
max r̂ 1.006.
Model ability (θ) and item difficulty on one shared logit scale. The grey strip is a histogram of item difficulty; dots below are models, colored by vendor, positioned at their posterior mean θ, with an 89% credible interval whisker.
One point per item, posterior means. Saturated items (every model right, or every model wrong) are left off entirely — Muted points have a wide discrimination posterior (uncertain) or a confidently low one — hover for which.
How much statistical information this benchmark's items carry at each ability level, summed over reliably-estimated items only. Ticks along the bottom mark where your actual models' θ estimates sit.
Items unlikely to be pulling their weight right now, and why.
Correlation of model θ across every benchmark pair with ≥4 shared models. High correlation = ability on one predicts ability on the other. Faded cells are computed from fewer than 6 shared models — hover for the exact count; treat these as noise, not a real relationship.
Every benchmark is fit separately, so raw θ scales aren't comparable -- each is only identified up to its own linear transform. Linked here onto MMLU-Pro's scale using the models that took every benchmark as anchors (mean/sigma equating: rescale so those models' θ mean and spread match on every benchmark, the standard method for placing separately-calibrated tests on one scale via a shared examinee group). Difficulty and discrimination are transformed consistently with θ, and the curves recomputed on the aligned scale, not just re-drawn. Benchmarks here range from 22 curve-eligible items (AIME 2024) to 5,613 (MMLU-Pro), and raw summed information mechanically tracks item count, not how informative a typical item is. So every benchmark is rarefied down to the smallest one's item count: resampled without replacement many times, summed per draw. The line is the median draw; the shaded band is the [10th, 90th] percentile across draws — how much a same-length test from this benchmark's item pool would vary just from which items got picked. Full item counts are in the legend.
Every item is tagged against 8 capability dimensions (math, science, logic, code, reading & instruction-following, factual recall, multi-step reasoning, other) using a hand-authored Q-matrix keyed off each item's category/subcategory — not a per-item model fit yet. Loadings are soft priors (0-1, not mutually exclusive), meant to be refined later by an actual multidimensional-IRT fit; treat this page as descriptive, not a finished result. MacBench and MacBench Ablations are excluded (here and everywhere else on this page) — both are substantially visual benchmarks (micrographs, XRD plots, patent figures) and none of these 8 dimensions capture that; tagging them "science"/"other" would misrepresent what actually makes those items hard rather than just approximate it.
Mean skill loading across every item in the benchmark (all items, not just non-saturated -- this describes what's being tested, independent of what the fit found discriminating). Darker = more central to that benchmark.
Each item assigned to its single highest-loading skill, then counted per benchmark. Coarser than the fingerprint above (throws away every item's secondary skills) but easier to read as "what fraction of this benchmark is fundamentally a math test."
For items with a meaningful (>0.15) loading on that skill, the loading-weighted mean fitted discrimination and difficulty (non-saturated items only, from the per-benchmark 2PL fits -- these are today's UNIdimensional posterior means, not skill-specific estimates, so read this as "how discriminating are math-flavored items on average across every benchmark that has any," not as a per-skill ability measurement).