← Leaderboard

Item response theory analysis

Ability, difficulty & latent structure

2PL IRT fit separately per benchmark.

θ, difficulty (b), and discrimination (a, modeled as log(a) ∼ Normal(0, 0.5)) are fit jointly via MCMC, not point-estimated separately. Diagnostics (divergences, , ESS) were checked for every benchmark before trusting its output — worst case so far: 0 divergences, max 1.006.

Wright map

Model ability (θ) and item difficulty on one shared logit scale. The grey strip is a histogram of item difficulty; dots below are models, colored by vendor, positioned at their posterior mean θ, with an 89% credible interval whisker.

Ability / item difficulty (θ, logit scale)

Discrimination vs. difficulty

One point per item, posterior means. Saturated items (every model right, or every model wrong) are left off entirely — Muted points have a wide discrimination posterior (uncertain) or a confidently low one — hover for which.

Discrimination (a)
Item difficulty (b, logit scale)

Test information

How much statistical information this benchmark's items carry at each ability level, summed over reliably-estimated items only. Ticks along the bottom mark where your actual models' θ estimates sit.

Information I(θ)
Ability (θ, logit scale)

Flagged items

Items unlikely to be pulling their weight right now, and why.