About this page. This validity documentation is maintained by the ELPAC team at EUROCONTROL as a living methodological record. It is not an ICAO endorsement and does not by itself constitute a regulatory determination. National civil aviation authorities remain responsible for licensing decisions and for determining whether ELPAC results satisfy the ICAO language-proficiency requirements in their jurisdiction.
Last reviewed: 18 June 2026.
Paper 1 — Rasch (1PL) item response theory
Classical Test Theory tells us the scores are reliable. Rasch tells us whether candidates and items can be placed on the same measurement scale at all. The diagnostics below say they can, with the limits a screening test is allowed to have — fit indices are within acceptable bounds, person and item separation are acceptable, and item difficulties span the right range (Bond & Fox, 2015; Green, 2013).
Candidates calibrated
16,240
Extreme scorers (all right / all wrong) excluded from θ estimation as required by Rasch.
Person reliability
0.71–0.77
Rasch analogue of α; typically slightly lower because it accounts for measurement error in θ. Values ≥ 0.70 are acceptable for a screening test.
Mean infit MSQ
0.98
Bond and Fox (2015, p. 270) describe 0.5–1.5 as productive for measurement and 0.7–1.3 as the more demanding band. The aggregate sits comfortably inside the tighter band on every form.
Calibration summary
On every operational form the Rasch model converges with acceptable person and item separation reliabilities and well-behaved fit. The form-by-form numbers below confirm that candidates and items can be placed on a single, common scale.
One row per operational form. Item separation indicates how reliably the form distinguishes items along the difficulty continuum; person separation does the same for candidates.
| Form | N (kept) | Person rel. | Marginal rel. | Item sep. | Mean b | b range | Infit | Outfit | Misfit items |
|---|---|---|---|---|---|---|---|---|---|
| ATC A | 2,082 (−21) | 0.71 · acceptable | 0.71 | 15.0 | 0.00 | -3.5 … 1.9 | 0.98 | 0.93 | 8 / 45 (1↑ / 7↓) |
| ATC B | 2,021 (−26) | 0.77 · acceptable | 0.77 | 16.5 | 0.00 | -3.3 … 1.7 | 0.99 | 0.97 | 4 / 45 (2↑ / 2↓) |
| ATC C | 2,069 (−23) | 0.72 · acceptable | 0.72 | 15.7 | 0.00 | -4.2 … 2.2 | 0.99 | 0.95 | 2 / 45 (0↑ / 2↓) |
| ATC D | 1,342 (−21) | 0.72 · acceptable | 0.72 | 12.1 | 0.00 | -2.7 … 3.8 | 0.98 | 0.90 | 16 / 45 (4↑ / 12↓) |
| Pilot A | 3,628 (−21) | 0.76 · acceptable | 0.76 | 22.6 | 0.00 | -3.8 … 4.1 | 0.98 | 0.95 | 9 / 45 (2↑ / 7↓) |
| Pilot B | 3,552 (−45) | 0.76 · acceptable | 0.76 | 19.8 | 0.00 | -2.9 … 3.2 | 0.98 | 0.93 | 10 / 45 (2↑ / 8↓) |
| Pilot C | 1,364 (−25) | 0.74 · acceptable | 0.74 | 7.8 | 0.00 | -5.9 … 3.6 | 0.96 | 0.90 | 11 / 45 (3↑ / 8↓) |
Following Bond and Fox (2015, ch. 12), infit/outfit MSQ in 0.5–1.5 is productive for measurement and 0.7–1.3 is the tighter, preferred band; values above the band indicate noisy (under-fitting) items, values below it indicate over-predictable (over-fitting) items that contribute little new information. A small number of margin items per form is normal and expected for an operational bank, and is the routine target of bank maintenance.
Person–item maps (Wright maps)
Candidate ability and item difficulty sit on the same scale, and the bulk of items is targeted at the borderline candidate — exactly what a screening test at the Level 4 floor needs. Precision is strongest where the pass/fail decision is taken.
Candidate ability (θ, left) and item difficulty (b, right) on the same logit axis. Bond and Fox (2015, ch. 4) treat targeting — the alignment of person and item distributions — as the central diagnostic of how well a test measures its intended population. On every form, item difficulties sit below the candidate ability mass: Paper 1 is calibrated to be passable by the target population, but precision is strongest below the typical candidate and weakest at the top — which has direct consequences for any decision taken near the upper end of the scale.
ATC A — person–item map
N = 2,082Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.
ATC B — person–item map
N = 2,021Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.
ATC C — person–item map
N = 2,069Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.
ATC D — person–item map
N = 1,342Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.
Pilot A — person–item map
N = 3,628Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.
Pilot B — person–item map
N = 3,552Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.
Pilot C — person–item map
N = 1,364Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.
Test information functions
Each form measures most precisely in the zone where the pass/fail decision is taken. Above the pass mark precision tapers — a known and accepted limit of a screening test, not a defect in the operational decision.
Each form measures most precisely in the part of the ability range where the pass/fail decision is taken. Precision tapers above the pass mark — Green (2013, p. 78) treats this as acceptable for a screen, not as a defect. The test information function I(θ) is the formal measure of this, with SEM(θ) = 1 / √I(θ) on the logit scale (Bond & Fox, 2015, ch. 4).
ATC A — test information & SEM
45 itemsATC B — test information & SEM
45 itemsATC C — test information & SEM
45 itemsATC D — test information & SEM
45 itemsPilot A — test information & SEM
45 itemsPilot B — test information & SEM
45 itemsPilot C — test information & SEM
45 itemsItem characteristic curves
Items behave the way the model expects: easier items produce a higher chance of a correct answer at every ability level, and the three reference items differ only in difficulty, not in shape. The scale behaves consistently across items.
Three reference items per form — easiest, median and hardest — illustrating the Rasch ICC shape. All ICCs share the same logistic slope; they differ only in horizontal position. Bond and Fox (2015, ch. 12) point out that this equal-slope constraint is the price of Rasch's specific objectivity: the CTT point-biserials show some variation across items, so a 2PL would fit marginally better, but only Rasch supports invariant person and item measures on a common scale.
ATC A — item characteristic curves
3 reference itemsEasiest (b = -3.491), median (b = 0.128), hardest (b = 1.87). Curves shift right with item difficulty.
ATC B — item characteristic curves
3 reference itemsEasiest (b = -3.292), median (b = 0.084), hardest (b = 1.699). Curves shift right with item difficulty.
ATC C — item characteristic curves
3 reference itemsEasiest (b = -4.182), median (b = 0.402), hardest (b = 2.183). Curves shift right with item difficulty.
ATC D — item characteristic curves
3 reference itemsEasiest (b = -2.736), median (b = -0.106), hardest (b = 3.747). Curves shift right with item difficulty.
Pilot A — item characteristic curves
3 reference itemsEasiest (b = -3.834), median (b = 0.011), hardest (b = 4.12). Curves shift right with item difficulty.
Pilot B — item characteristic curves
3 reference itemsEasiest (b = -2.887), median (b = -0.315), hardest (b = 3.158). Curves shift right with item difficulty.
Pilot C — item characteristic curves
3 reference itemsEasiest (b = -5.867), median (b = 0.161), hardest (b = 3.567). Curves shift right with item difficulty.
Per-item fit (expandable)
The large majority of items fit the model in the preferred band. A small number of margin items is identified per form; these are the routine targets of item-bank maintenance, not defects in the operational decision.
Full Rasch item table per form. Verdict applies the Bond & Fox (2015) fit bands: excellent 0.8–1.2, acceptable 0.7–1.3, productive (review) 0.5–1.5, otherwise misfit.
ATC A — 45 items (8 flagged)
| Item | b | SE | Infit | Outfit | Verdict |
|---|---|---|---|---|---|
| 2000380 | -3.49 | 0.24 | 0.90 | 0.55 | productive (review) |
| 2000381 | -2.48 | 0.15 | 0.91 | 0.62 | productive (review) |
| 2000382 | 0.13 | 0.06 | 0.94 | 0.94 | excellent |
| 2000383 | 0.23 | 0.06 | 0.97 | 0.96 | excellent |
| 2000384 | -1.90 | 0.12 | 0.92 | 0.75 | acceptable |
| 2000385 | -1.35 | 0.09 | 0.91 | 0.69 | productive (review) |
| 2000300 | 1.07 | 0.05 | 1.03 | 1.03 | excellent |
| 2000301 | 1.20 | 0.05 | 1.20 | 1.27 | acceptable |
| 2000302 | 0.89 | 0.05 | 0.98 | 0.96 | excellent |
| 2000309 | -0.77 | 0.07 | 0.94 | 0.88 | excellent |
| 2000310 | 1.40 | 0.05 | 1.12 | 1.16 | excellent |
| 2000311 | 0.33 | 0.06 | 0.93 | 0.89 | excellent |
| 2000197 | -0.87 | 0.08 | 0.94 | 0.84 | excellent |
| 2000198 | -1.99 | 0.12 | 0.94 | 0.91 | excellent |
| 2000199 | -1.34 | 0.09 | 0.85 | 0.56 | productive (review) |
| 2000322 | 0.18 | 0.06 | 1.15 | 1.14 | excellent |
| 2000323 | 1.20 | 0.05 | 1.03 | 1.05 | excellent |
| 2000324 | -0.33 | 0.07 | 0.88 | 0.75 | acceptable |
| 2000325 | -0.60 | 0.07 | 0.81 | 0.57 | productive (review) |
| 2000326 | 0.43 | 0.05 | 0.91 | 0.88 | excellent |
| 2000342 | 0.03 | 0.06 | 1.12 | 1.27 | acceptable |
| 2000343 | 0.17 | 0.06 | 0.82 | 0.72 | acceptable |
| 2000344 | 1.44 | 0.05 | 1.19 | 1.27 | acceptable |
| 2000345 | -0.19 | 0.06 | 1.06 | 1.14 | excellent |
| 2000346 | -1.21 | 0.09 | 0.91 | 0.75 | acceptable |
| 2000348 | 0.46 | 0.05 | 0.97 | 0.95 | excellent |
| 2000349 | 0.51 | 0.05 | 0.97 | 0.94 | excellent |
| 2000350 | -0.68 | 0.07 | 1.03 | 1.16 | excellent |
| 2000351 | 1.66 | 0.05 | 1.11 | 1.15 | excellent |
| 2000352 | 0.04 | 0.06 | 1.03 | 1.02 | excellent |
| 2000353 | 0.15 | 0.06 | 0.92 | 0.86 | excellent |
| 2000354 | 0.68 | 0.05 | 1.03 | 1.02 | excellent |
| 2000355 | 1.50 | 0.05 | 1.08 | 1.10 | excellent |
| 2000356 | -1.04 | 0.08 | 0.88 | 0.69 | productive (review) |
| 2000358 | 1.14 | 0.05 | 0.98 | 0.95 | excellent |
| 2000395 | 0.09 | 0.06 | 1.00 | 0.97 | excellent |
| 2000396 | 0.99 | 0.05 | 0.89 | 0.87 | excellent |
| 2000397 | 0.04 | 0.06 | 1.01 | 0.95 | excellent |
| 2000398 | 1.87 | 0.05 | 1.13 | 1.20 | excellent |
| 2000400 | 0.81 | 0.05 | 1.07 | 1.04 | excellent |
| 2000271 | -0.04 | 0.06 | 0.86 | 0.76 | acceptable |
| 2000272 | -1.21 | 0.09 | 0.98 | 0.91 | excellent |
| 2000273 | -0.16 | 0.06 | 0.82 | 0.66 | productive (review) |
| 2000274 | -0.06 | 0.06 | 0.96 | 0.90 | excellent |
| 2000275 | 1.07 | 0.05 | 1.24 | 1.31 | productive (review) |
ATC B — 45 items (4 flagged)
| Item | b | SE | Infit | Outfit | Verdict |
|---|---|---|---|---|---|
| 2000454 | -3.00 | 0.17 | 0.92 | 0.86 | excellent |
| 2000455 | -3.29 | 0.19 | 0.91 | 0.81 | excellent |
| 2000456 | 1.56 | 0.05 | 1.07 | 1.08 | excellent |
| 2000457 | -0.05 | 0.06 | 1.05 | 1.12 | excellent |
| 2000458 | -0.62 | 0.07 | 1.08 | 1.20 | acceptable |
| 2000459 | -1.50 | 0.09 | 0.95 | 0.92 | excellent |
| 2000317 | 1.03 | 0.05 | 0.97 | 0.96 | excellent |
| 2000318 | 1.02 | 0.05 | 0.97 | 0.95 | excellent |
| 2000319 | 0.65 | 0.05 | 1.03 | 1.00 | excellent |
| 2000314 | 0.21 | 0.05 | 0.99 | 0.96 | excellent |
| 2000315 | -0.53 | 0.06 | 1.00 | 0.92 | excellent |
| 2000316 | -0.94 | 0.07 | 0.88 | 0.71 | acceptable |
| 2000201 | 1.18 | 0.05 | 0.99 | 0.98 | excellent |
| 2000202 | -2.00 | 0.10 | 0.92 | 0.82 | excellent |
| 2000203 | -0.09 | 0.06 | 0.91 | 0.81 | excellent |
| 2000460 | 0.81 | 0.05 | 0.98 | 0.95 | excellent |
| 2000461 | -0.61 | 0.06 | 1.00 | 1.08 | excellent |
| 2000462 | -1.37 | 0.08 | 0.90 | 0.70 | productive (review) |
| 2000463 | 0.17 | 0.05 | 1.07 | 1.06 | excellent |
| 2000464 | 0.78 | 0.05 | 1.25 | 1.32 | productive (review) |
| 2000465 | 0.96 | 0.05 | 0.92 | 0.89 | excellent |
| 2000466 | -0.79 | 0.07 | 0.82 | 0.63 | productive (review) |
| 2000467 | -1.20 | 0.08 | 0.90 | 0.86 | excellent |
| 2000468 | 1.70 | 0.05 | 1.13 | 1.23 | acceptable |
| 2000469 | 1.07 | 0.05 | 1.04 | 1.06 | excellent |
| 2000470 | -0.49 | 0.06 | 1.24 | 1.64 | misfit |
| 2000471 | 0.49 | 0.05 | 0.99 | 1.00 | excellent |
| 2000472 | 1.32 | 0.05 | 1.01 | 1.02 | excellent |
| 2000473 | 0.90 | 0.05 | 1.05 | 1.05 | excellent |
| 2000474 | 0.01 | 0.06 | 0.99 | 0.94 | excellent |
| 2000475 | 0.39 | 0.05 | 0.92 | 0.92 | excellent |
| 2000476 | 1.44 | 0.05 | 1.15 | 1.20 | excellent |
| 2000477 | -0.20 | 0.06 | 1.08 | 1.13 | excellent |
| 2000478 | -0.48 | 0.06 | 0.91 | 0.84 | excellent |
| 2000479 | -1.07 | 0.07 | 0.90 | 0.71 | acceptable |
| 2000482 | -0.36 | 0.06 | 0.92 | 0.87 | excellent |
| 2000483 | -0.21 | 0.06 | 0.93 | 0.88 | excellent |
| 2000484 | 1.46 | 0.05 | 1.15 | 1.20 | acceptable |
| 2000485 | 1.30 | 0.05 | 0.96 | 0.95 | excellent |
| 2000486 | 0.08 | 0.05 | 1.11 | 1.22 | acceptable |
| 2000487 | 0.15 | 0.05 | 0.94 | 0.87 | excellent |
| 2000488 | -0.55 | 0.06 | 0.92 | 0.79 | acceptable |
| 2000489 | 0.99 | 0.05 | 0.90 | 0.87 | excellent |
| 2000490 | -0.43 | 0.06 | 0.84 | 0.71 | acceptable |
| 2000491 | 0.09 | 0.05 | 0.92 | 0.86 | excellent |
ATC C — 45 items (2 flagged)
| Item | b | SE | Infit | Outfit | Verdict |
|---|---|---|---|---|---|
| 2000493 | -4.18 | 0.30 | 0.90 | 0.39 | misfit |
| 2000494 | -3.15 | 0.18 | 0.94 | 1.11 | excellent |
| 2000495 | -0.15 | 0.06 | 0.99 | 1.00 | excellent |
| 2000496 | -0.42 | 0.06 | 0.96 | 0.92 | excellent |
| 2000497 | -0.16 | 0.06 | 1.03 | 1.04 | excellent |
| 2000498 | -0.58 | 0.07 | 0.92 | 0.82 | excellent |
| 2000499 | -0.41 | 0.06 | 0.97 | 0.96 | excellent |
| 2000500 | 0.78 | 0.05 | 1.11 | 1.11 | excellent |
| 2000501 | -1.76 | 0.10 | 0.91 | 0.77 | acceptable |
| 2000502 | 0.36 | 0.05 | 1.24 | 1.26 | acceptable |
| 2000503 | 1.03 | 0.05 | 1.09 | 1.09 | excellent |
| 2000504 | -0.23 | 0.06 | 1.01 | 1.02 | excellent |
| 2000505 | -1.57 | 0.09 | 0.95 | 0.85 | excellent |
| 2000506 | -2.18 | 0.12 | 0.85 | 0.46 | misfit |
| 2000507 | 0.44 | 0.05 | 1.08 | 1.13 | excellent |
| 2000508 | 1.88 | 0.05 | 1.01 | 1.02 | excellent |
| 2000509 | 0.92 | 0.05 | 0.97 | 0.98 | excellent |
| 2000510 | -0.30 | 0.06 | 0.91 | 0.80 | acceptable |
| 2000511 | 1.07 | 0.05 | 1.11 | 1.15 | excellent |
| 2000512 | 0.48 | 0.05 | 1.04 | 1.06 | excellent |
| 2000513 | -0.92 | 0.07 | 1.02 | 1.01 | excellent |
| 2000514 | -1.58 | 0.09 | 0.93 | 0.82 | excellent |
| 2000515 | 0.40 | 0.05 | 1.00 | 0.96 | excellent |
| 2000516 | 0.41 | 0.05 | 0.94 | 0.94 | excellent |
| 2000517 | 2.18 | 0.05 | 0.99 | 1.01 | excellent |
| 2000518 | 0.76 | 0.05 | 1.04 | 1.05 | excellent |
| 2000519 | 0.64 | 0.05 | 0.91 | 0.88 | excellent |
| 2000520 | 0.38 | 0.05 | 1.01 | 1.02 | excellent |
| 2000521 | 0.82 | 0.05 | 0.94 | 0.92 | excellent |
| 2000522 | 0.95 | 0.05 | 1.05 | 1.06 | excellent |
| 2000523 | 1.14 | 0.05 | 0.88 | 0.86 | excellent |
| 2000524 | 0.61 | 0.05 | 0.97 | 0.96 | excellent |
| 2000525 | 0.48 | 0.05 | 0.90 | 0.89 | excellent |
| 2000526 | 0.04 | 0.05 | 0.88 | 0.85 | excellent |
| 2000527 | 0.72 | 0.05 | 0.89 | 0.87 | excellent |
| 2000529 | -1.12 | 0.08 | 0.92 | 0.76 | acceptable |
| 2000530 | 0.20 | 0.05 | 0.99 | 0.98 | excellent |
| 2000531 | 0.63 | 0.05 | 1.03 | 1.05 | excellent |
| 2000532 | -1.76 | 0.10 | 0.92 | 0.71 | acceptable |
| 2000533 | 0.22 | 0.05 | 0.94 | 0.93 | excellent |
| 2000536 | 0.60 | 0.05 | 1.06 | 1.10 | excellent |
| 2000537 | 1.67 | 0.05 | 1.02 | 1.03 | excellent |
| 2000538 | -0.24 | 0.06 | 1.05 | 1.03 | excellent |
| 2000539 | 0.47 | 0.05 | 1.15 | 1.21 | acceptable |
| 2000540 | 0.41 | 0.05 | 1.06 | 1.12 | excellent |
ATC D — 45 items (16 flagged)
| Item | b | SE | Infit | Outfit | Verdict |
|---|---|---|---|---|---|
| 2100009 | -2.11 | 0.21 | 0.91 | 0.79 | acceptable |
| 2100010 | -1.94 | 0.20 | 0.96 | 0.88 | excellent |
| 2100011 | -0.86 | 0.13 | 0.97 | 0.89 | excellent |
| 2100012 | -1.41 | 0.16 | 0.95 | 0.76 | acceptable |
| 2100013 | -2.74 | 0.28 | 0.92 | 0.59 | productive (review) |
| 2100014 | -2.26 | 0.23 | 0.93 | 0.47 | misfit |
| 2100015 | -1.32 | 0.15 | 1.01 | 1.36 | productive (review) |
| 2100016 | 1.90 | 0.06 | 0.91 | 0.89 | excellent |
| 2100017 | -0.02 | 0.09 | 1.15 | 1.38 | productive (review) |
| 2100018 | -1.54 | 0.16 | 0.96 | 0.88 | excellent |
| 2100019 | -1.48 | 0.16 | 0.94 | 0.68 | productive (review) |
| 2100020 | -0.80 | 0.12 | 1.00 | 1.01 | excellent |
| 2100021 | -0.51 | 0.11 | 0.90 | 0.73 | acceptable |
| 2100022 | -1.41 | 0.16 | 1.02 | 1.08 | excellent |
| 2100023 | -0.96 | 0.13 | 0.90 | 0.62 | productive (review) |
| 2100024 | 2.14 | 0.06 | 0.96 | 0.94 | excellent |
| 2100025 | -0.07 | 0.09 | 0.91 | 0.69 | productive (review) |
| 2100026 | -1.68 | 0.17 | 0.86 | 0.44 | misfit |
| 2100027 | -0.56 | 0.11 | 0.87 | 0.50 | productive (review) |
| 2100028 | 1.59 | 0.06 | 0.95 | 0.93 | excellent |
| 2100077 | -0.11 | 0.10 | 0.96 | 0.72 | acceptable |
| 2100078 | 3.75 | 0.07 | 1.02 | 1.31 | productive (review) |
| 2100079 | 0.69 | 0.08 | 1.12 | 1.21 | acceptable |
| 2100080 | -0.19 | 0.10 | 0.92 | 0.88 | excellent |
| 2100081 | 1.47 | 0.07 | 1.15 | 1.18 | excellent |
| 2100034 | -1.89 | 0.19 | 0.94 | 1.04 | excellent |
| 2100035 | 1.58 | 0.06 | 1.20 | 1.27 | acceptable |
| 2100036 | 2.35 | 0.06 | 1.36 | 1.53 | misfit |
| 2100037 | -0.60 | 0.11 | 0.92 | 0.69 | productive (review) |
| 2100038 | 0.87 | 0.07 | 0.88 | 0.77 | acceptable |
| 2100039 | 0.11 | 0.09 | 1.14 | 1.13 | excellent |
| 2100040 | -0.89 | 0.13 | 0.88 | 0.55 | productive (review) |
| 2100041 | -0.42 | 0.11 | 0.87 | 0.59 | productive (review) |
| 2100042 | -0.84 | 0.12 | 0.94 | 0.88 | excellent |
| 2100043 | 0.06 | 0.09 | 0.89 | 0.65 | productive (review) |
| 2100045 | 0.73 | 0.07 | 0.89 | 0.79 | acceptable |
| 2100046 | 3.25 | 0.06 | 1.09 | 1.29 | acceptable |
| 2100047 | 0.27 | 0.09 | 0.94 | 0.98 | excellent |
| 2100048 | 0.99 | 0.07 | 0.95 | 0.89 | excellent |
| 2100049 | 0.60 | 0.08 | 1.09 | 1.14 | excellent |
| 2100052 | -0.02 | 0.09 | 0.91 | 0.71 | acceptable |
| 2100053 | 1.37 | 0.07 | 1.07 | 1.08 | excellent |
| 2100054 | -0.28 | 0.10 | 0.90 | 0.64 | productive (review) |
| 2100055 | 1.16 | 0.07 | 0.96 | 0.89 | excellent |
| 2100056 | 2.02 | 0.06 | 1.10 | 1.17 | excellent |
Pilot A — 45 items (9 flagged)
| Item | b | SE | Infit | Outfit | Verdict |
|---|---|---|---|---|---|
| 2100153 | -0.49 | 0.06 | 0.91 | 0.79 | acceptable |
| 2100154 | -0.48 | 0.06 | 1.03 | 1.14 | excellent |
| 2100155 | -1.44 | 0.08 | 0.97 | 0.94 | excellent |
| 2100156 | -0.74 | 0.06 | 1.04 | 1.13 | excellent |
| 2100157 | -0.66 | 0.06 | 0.99 | 0.92 | excellent |
| 2100158 | -2.83 | 0.15 | 0.92 | 0.62 | productive (review) |
| 2100159 | -3.83 | 0.23 | 0.91 | 0.68 | productive (review) |
| 2100160 | -2.78 | 0.14 | 0.94 | 0.85 | excellent |
| 2100161 | 1.64 | 0.04 | 0.86 | 0.83 | excellent |
| 2100162 | -2.08 | 0.10 | 0.93 | 0.59 | productive (review) |
| 2100163 | -3.38 | 0.19 | 0.92 | 0.81 | excellent |
| 2100164 | 3.35 | 0.04 | 0.97 | 1.04 | excellent |
| 2100165 | 0.97 | 0.04 | 0.90 | 0.82 | excellent |
| 2100166 | 0.01 | 0.05 | 0.99 | 0.96 | excellent |
| 2100167 | 0.29 | 0.04 | 0.90 | 0.82 | excellent |
| 2100170 | 1.45 | 0.04 | 1.03 | 1.03 | excellent |
| 2100171 | 0.10 | 0.04 | 1.09 | 1.18 | excellent |
| 2100172 | 1.29 | 0.04 | 1.23 | 1.28 | acceptable |
| 2100173 | -1.22 | 0.07 | 0.98 | 0.88 | excellent |
| 2100174 | -0.84 | 0.06 | 0.87 | 0.60 | productive (review) |
| 2100175 | 0.15 | 0.04 | 0.91 | 0.73 | acceptable |
| 2100176 | -0.25 | 0.05 | 0.95 | 0.91 | excellent |
| 2100177 | 0.02 | 0.05 | 1.02 | 1.07 | excellent |
| 2100178 | -1.59 | 0.08 | 0.93 | 0.58 | productive (review) |
| 2100179 | 0.07 | 0.05 | 1.01 | 0.96 | excellent |
| 2100180 | -0.00 | 0.05 | 1.08 | 1.21 | acceptable |
| 2100181 | 0.39 | 0.04 | 0.75 | 0.59 | productive (review) |
| 2100182 | 0.42 | 0.04 | 0.94 | 0.87 | excellent |
| 2100183 | -2.27 | 0.11 | 0.91 | 0.54 | productive (review) |
| 2100184 | 1.42 | 0.04 | 1.01 | 1.00 | excellent |
| 2100185 | 0.30 | 0.04 | 0.98 | 0.98 | excellent |
| 2100186 | -1.30 | 0.07 | 1.01 | 1.21 | acceptable |
| 2100187 | 1.02 | 0.04 | 1.10 | 1.12 | excellent |
| 2100188 | 1.62 | 0.04 | 0.97 | 0.96 | excellent |
| 2100189 | 3.27 | 0.04 | 1.15 | 1.57 | misfit |
| 2100191 | 4.12 | 0.05 | 1.13 | 1.70 | misfit |
| 2100192 | 2.52 | 0.04 | 1.17 | 1.26 | acceptable |
| 2100193 | -0.36 | 0.05 | 1.02 | 0.96 | excellent |
| 2100194 | -0.71 | 0.06 | 1.02 | 0.97 | excellent |
| 2100195 | -0.04 | 0.05 | 0.95 | 0.83 | excellent |
| 2100196 | -0.27 | 0.05 | 1.03 | 1.08 | excellent |
| 2100197 | 1.24 | 0.04 | 0.88 | 0.84 | excellent |
| 2100198 | 1.92 | 0.04 | 1.01 | 1.02 | excellent |
| 2100199 | -0.67 | 0.06 | 0.99 | 1.02 | excellent |
| 2100200 | 0.65 | 0.04 | 0.97 | 0.95 | excellent |
Pilot B — 45 items (10 flagged)
| Item | b | SE | Infit | Outfit | Verdict |
|---|---|---|---|---|---|
| 2100201 | -2.77 | 0.17 | 0.93 | 0.83 | excellent |
| 2100202 | -1.72 | 0.11 | 0.98 | 1.04 | excellent |
| 2100203 | -1.08 | 0.08 | 0.97 | 1.13 | excellent |
| 2100204 | -1.42 | 0.10 | 0.96 | 0.90 | excellent |
| 2100205 | -2.01 | 0.12 | 0.97 | 1.23 | acceptable |
| 2100206 | -2.75 | 0.17 | 0.95 | 1.43 | productive (review) |
| 2100207 | -2.89 | 0.18 | 0.92 | 0.77 | acceptable |
| 2100208 | 0.89 | 0.04 | 0.93 | 0.85 | excellent |
| 2100209 | 1.02 | 0.04 | 1.09 | 1.12 | excellent |
| 2100210 | 1.12 | 0.04 | 1.02 | 1.00 | excellent |
| 2100211 | -2.40 | 0.14 | 0.94 | 0.99 | excellent |
| 2100212 | -0.55 | 0.07 | 0.99 | 1.08 | excellent |
| 2100213 | 0.17 | 0.05 | 0.96 | 0.87 | excellent |
| 2100214 | 1.22 | 0.04 | 1.16 | 1.25 | acceptable |
| 2100215 | -0.32 | 0.06 | 0.95 | 0.94 | excellent |
| 2100216 | 1.55 | 0.04 | 0.89 | 0.86 | excellent |
| 2100217 | -0.40 | 0.06 | 0.93 | 0.68 | productive (review) |
| 2100218 | 3.16 | 0.04 | 0.95 | 0.93 | excellent |
| 2100219 | -0.44 | 0.06 | 0.95 | 0.85 | excellent |
| 2100220 | 2.68 | 0.04 | 1.40 | 1.68 | misfit |
| 2100221 | -0.27 | 0.06 | 0.98 | 0.92 | excellent |
| 2100222 | 1.76 | 0.04 | 0.88 | 0.83 | excellent |
| 2100223 | 0.76 | 0.04 | 1.06 | 0.95 | excellent |
| 2100224 | -0.88 | 0.07 | 0.98 | 1.14 | excellent |
| 2100225 | 1.54 | 0.04 | 0.89 | 0.84 | excellent |
| 2100226 | 1.71 | 0.04 | 1.18 | 1.25 | acceptable |
| 2100227 | 0.89 | 0.04 | 1.03 | 0.98 | excellent |
| 2100228 | 2.12 | 0.04 | 1.01 | 1.03 | excellent |
| 2100229 | 2.56 | 0.04 | 1.02 | 1.04 | excellent |
| 2100230 | -0.32 | 0.06 | 0.92 | 0.73 | acceptable |
| 2100231 | -0.89 | 0.07 | 0.92 | 0.62 | productive (review) |
| 2100232 | 1.08 | 0.04 | 0.97 | 0.94 | excellent |
| 2100233 | -1.01 | 0.08 | 0.88 | 0.49 | misfit |
| 2100234 | 1.09 | 0.04 | 1.08 | 1.06 | excellent |
| 2100235 | 1.03 | 0.04 | 0.97 | 0.93 | excellent |
| 2100236 | -1.31 | 0.09 | 0.93 | 0.54 | productive (review) |
| 2100237 | -0.46 | 0.06 | 1.05 | 1.09 | excellent |
| 2100238 | 1.67 | 0.04 | 1.10 | 1.11 | excellent |
| 2100239 | 2.27 | 0.04 | 1.00 | 1.01 | excellent |
| 2100240 | -0.52 | 0.07 | 0.93 | 0.78 | acceptable |
| 2100241 | 0.03 | 0.05 | 0.93 | 0.77 | acceptable |
| 2100242 | -0.47 | 0.06 | 0.95 | 0.65 | productive (review) |
| 2100243 | -0.74 | 0.07 | 0.92 | 0.66 | productive (review) |
| 2100244 | -2.23 | 0.13 | 0.88 | 0.54 | productive (review) |
| 2100245 | -2.47 | 0.15 | 0.90 | 0.34 | misfit |
Pilot C — 45 items (11 flagged)
| Item | b | SE | Infit | Outfit | Verdict |
|---|---|---|---|---|---|
| 2100249 | 0.40 | 0.08 | 0.99 | 0.92 | excellent |
| 2100250 | -1.44 | 0.15 | 0.99 | 0.93 | excellent |
| 2100251 | -1.36 | 0.15 | 0.97 | 0.87 | excellent |
| 2100252 | -1.64 | 0.17 | 0.96 | 0.83 | excellent |
| 2100253 | -1.89 | 0.18 | 0.91 | 0.64 | productive (review) |
| 2100254 | -1.48 | 0.15 | 0.92 | 0.65 | productive (review) |
| 2100255 | -3.25 | 0.34 | 0.90 | 0.47 | misfit |
| 2100256 | -0.82 | 0.12 | 0.97 | 0.97 | excellent |
| 2100257 | 0.20 | 0.08 | 1.05 | 1.13 | excellent |
| 2100258 | -2.16 | 0.20 | 0.92 | 0.78 | acceptable |
| 2100259 | -3.95 | 0.48 | 0.86 | 0.21 | misfit |
| 2100260 | -5.87 | 1.24 | 0.00 | 0.00 | misfit |
| 2100267 | 3.57 | 0.07 | 0.98 | 1.24 | acceptable |
| 2100268 | -0.11 | 0.09 | 1.07 | 1.37 | productive (review) |
| 2100269 | -0.28 | 0.09 | 1.03 | 1.01 | excellent |
| 2100264 | 0.06 | 0.08 | 0.94 | 0.80 | acceptable |
| 2100265 | 0.85 | 0.07 | 1.00 | 0.90 | excellent |
| 2100266 | 2.84 | 0.06 | 0.85 | 0.81 | excellent |
| 2100270 | -1.51 | 0.15 | 0.93 | 0.78 | acceptable |
| 2100271 | 0.29 | 0.08 | 0.90 | 0.74 | acceptable |
| 2100272 | 2.20 | 0.06 | 1.09 | 1.14 | excellent |
| 2100273 | 0.16 | 0.08 | 0.82 | 0.61 | productive (review) |
| 2100274 | 1.94 | 0.06 | 1.07 | 1.10 | excellent |
| 2100275 | 1.77 | 0.06 | 1.21 | 1.33 | productive (review) |
| 2100276 | 0.83 | 0.07 | 0.94 | 0.88 | excellent |
| 2100277 | 1.01 | 0.07 | 0.96 | 0.87 | excellent |
| 2100278 | 2.21 | 0.06 | 1.03 | 1.08 | excellent |
| 2100279 | -0.58 | 0.10 | 0.92 | 0.73 | acceptable |
| 2100280 | 1.28 | 0.06 | 0.92 | 0.88 | excellent |
| 2100281 | 0.27 | 0.08 | 0.97 | 0.95 | excellent |
| 2100282 | -1.82 | 0.17 | 0.91 | 0.58 | productive (review) |
| 2100283 | 1.27 | 0.06 | 1.08 | 1.16 | excellent |
| 2100284 | 0.14 | 0.08 | 1.06 | 1.07 | excellent |
| 2100285 | 1.75 | 0.06 | 0.81 | 0.77 | acceptable |
| 2100286 | 0.35 | 0.08 | 1.06 | 1.04 | excellent |
| 2100287 | 1.06 | 0.07 | 1.28 | 1.54 | misfit |
| 2100288 | -0.67 | 0.11 | 0.89 | 0.60 | productive (review) |
| 2100289 | -0.73 | 0.11 | 1.00 | 1.22 | acceptable |
| 2100290 | 1.86 | 0.06 | 1.02 | 1.03 | excellent |
| 2100291 | 1.34 | 0.06 | 0.97 | 0.95 | excellent |
| 2100292 | 0.04 | 0.08 | 0.96 | 0.90 | excellent |
| 2100293 | -0.23 | 0.09 | 0.98 | 0.87 | excellent |
| 2100294 | 1.00 | 0.07 | 1.05 | 1.04 | excellent |
| 2100295 | 0.11 | 0.08 | 0.97 | 0.88 | excellent |
| 2100296 | 0.97 | 0.07 | 1.16 | 1.27 | acceptable |
Method
- Model — Rasch / one-parameter logistic: P(X = 1) = exp(θ − b) / (1 + exp(θ − b)) (Bond & Fox, 2015, ch. 3).
- Estimation — joint maximum likelihood with Wright bias correction (×(K−1)/K for items, ×(N−1)/N for persons; Bond & Fox, 2015, app. B). Item difficulties centred at 0.
- Extreme scorers — candidates with raw score 0 or full are excluded from θ estimation (their MLE is undefined). Counts shown in the summary table.
- Fit statistics — infit (info-weighted) and outfit (unweighted) mean-square residuals; 0.5–1.5 productive, 0.7–1.3 preferred (Bond & Fox, 2015, ch. 12).
- Reliability — person and item separation reliabilities derived from the variance of estimates and their standard errors (Bond & Fox, 2015, ch. 4).
Continuous-improvement programme
The items below sit on the published improvement roadmap. They describe planned methodological work, not operational deficiencies in the current pass/fail decision.
- Invariance — not tested. Rasch invariance evidence beyond per-form difficulty contrasts (anchored between-form calibration) remains the highest-priority item in the ongoing validation programme (Bond & Fox, 2015, ch. 6).
- Targeting mismatch at the top. Person means sit above item means on every form, so the test information functions show that measurement is least precise exactly where high-stakes decisions about strong candidates would be taken.
- No equating across forms. Each form is calibrated separately and difficulties are centred at 0 per form; without common-item or concurrent equating (Bond & Fox, 2015, ch. 5), b-parameters are not directly comparable across forms.
- 1PL is a modelling choice. The CTT point-biserials show some discrimination variation, so a 2PL would fit marginally better; Rasch is retained for the measurement properties it guarantees (invariance, conjoint additivity), not because it is the best descriptive fit.
Conclusion
At the aggregate level the Rasch model fits: candidates and items sit on the same interval scale, fit indices are within acceptable bounds, person and item separation are acceptable for a screening test, and item difficulties span the right range (Bond & Fox, 2015). That strengthens the "do we score the test accurately?" answer beyond what CTT alone can give (Green, 2013). The questions about whether forms are interchangeable and whether the score predicts radiotelephony performance still need the equating study and the criterion evidence flagged in the gaps section above.
Glossary of terms
- Cronbach's α
- Internal-consistency reliability on 0/1 item scores. Green (2013) cites ≥ 0.80 as acceptable for screening and ≥ 0.90 for high-stakes decisions.
- McDonald's ω
- Reliability estimate that relaxes the tau-equivalence assumption of α.
- SEM
- Standard error of measurement on the raw-score scale. A candidate's true score lies within roughly ±2·SEM of their observed score with 95% confidence.
- Conditional SEM (cSEM)
- SEM as a function of score (CTT) or ability θ (Rasch). Green (2013, p. 78) argues this is the relevant precision metric for decisions: what matters is precision at the cut, not on average.
- Infit / outfit MSQ
- Rasch fit statistics. Bond and Fox (2015, ch. 12) treat 0.5–1.5 as productive for measurement and 0.7–1.3 as the preferred tighter band.
- Person / item separation
- Rasch reliability analogues: how reliably the form distinguishes candidates along the ability continuum, and items along the difficulty continuum.
- Marginal reliability
- Population-level reliability of the Rasch ability estimates, reported alongside α for direct comparison.
- Targeting
- Alignment of person and item distributions on the logit scale (Bond & Fox, 2015, ch. 4). Good targeting means most candidates encounter items near their ability level.
- Decision consistency (p₀, κ)
- Probability that two parallel forms would classify a candidate the same way at the cut score, and Cohen's κ adjusting for chance agreement.
References
- Bond, T. G., & Fox, C. M. (2015). Applying the Rasch model: Fundamental measurement in the human sciences (3rd ed.). Routledge.
- Green, A. (2013). Exploring language assessment and testing. Routledge.