Accuracy weights every item equally. IRT estimates item difficulty and discrimination from the pattern of correct and incorrect responses.
Background / Rationale
Why use IRT for benchmark analysis?
Accuracy reports how many answers were correct. It does not show which items created the score difference, which items still separate models, or where a benchmark measures precisely. IRT uses the full model-by-item response matrix to answer those questions.
An item is most informative near the ability level where its outcome changes. The same item can distinguish one group of models but not another.
Summed item information shows where a benchmark separates models and where it has gaps or has saturated for the models being compared.
Background / Walkthrough
From scored answers to benchmark information.
Walk through the four quantities used in the analysis. Each step builds directly on the one before it.
| Model | I1 | I2 | I3 | I4 | I5 |
|---|---|---|---|---|---|
| A | 1 | 1 | 0 | 1 | 0 |
| B | 1 | 0 | 0 | 1 | 0 |
| C | 1 | 1 | 1 | 1 | 0 |
| D | 1 | 0 | 1 | 0 | 0 |
1 = correct · 0 = incorrect
Start with the response matrix
A response matrix has one row per model and one column per item. The entry \(Y_{mi}\) records whether model \(m\) answered item \(i\) correctly.
IRT uses the pattern across the whole matrix—not only each row’s total score.
Model the probability of a correct answer
The response model connects each observed 0 or 1 to three estimated quantities.
- \(\theta_m\)
- Ability of model \(m\)
- \(b_i\)
- Difficulty of item \(i\)
- \(a_i\)
- How sharply item \(i\) separates nearby ability levels
Convert sensitivity into information
Item information is high when an item’s outcome changes sharply near \(\theta\). It is low when nearly every model succeeds, nearly every model fails, or the item barely discriminates.
Sum information across the benchmark
The benchmark information curve adds the contribution of every item. More information means lower uncertainty at that ability level. A gap means the benchmark has few useful items there.
Background / Theory
How IRT turns model responses into comparable measurements.
For model \(m\) and item \(i\), a two-parameter logistic IRT model estimates the probability of a correct response from model ability \(\theta_m\), item difficulty \(b_i\), and item discrimination \(a_i\).
Ability relative to difficulty
When \(\theta_m=b_i\), predicted success is 50%. Higher \(\theta_m\) increases success; higher \(b_i\) decreases it. The parameter \(a_i\) controls the slope.
Information peaks near difficulty
An item gives the most information about models whose \(\theta\) lies near \(b_i\). Very easy or very hard items add little precision for those models.
Benchmark information sums item information
A benchmark’s information curve is the sum of its item information curves. Its peak identifies the ability level measured most precisely.
Shared models link benchmark scales
We fit each benchmark separately, then align its \(\theta\) scale using the models evaluated on every benchmark. This makes the curve locations comparable.
Limit: \(\theta\) is not a universal intelligence score. It is relative to the included models and items and depends on the scoring and linking model.