About this page. This validity documentation is maintained by the ELPAC team at EUROCONTROL as a living methodological record. It is not an ICAO endorsement and does not by itself constitute a regulatory determination. National civil aviation authorities remain responsible for licensing decisions and for determining whether ELPAC results satisfy the ICAO language-proficiency requirements in their jurisdiction.

Last reviewed: 18 June 2026.

Paper 1 — Rasch (1PL) item response theory

Classical Test Theory tells us the scores are reliable. Rasch tells us whether candidates and items can be placed on the same measurement scale at all. The diagnostics below say they can, with the limits a screening test is allowed to have — fit indices are within acceptable bounds, person and item separation are acceptable, and item difficulties span the right range (Bond & Fox, 2015; Green, 2013).

Candidates calibrated

16,240

Extreme scorers (all right / all wrong) excluded from θ estimation as required by Rasch.

Person reliability

0.710.77

Rasch analogue of α; typically slightly lower because it accounts for measurement error in θ. Values ≥ 0.70 are acceptable for a screening test.

Mean infit MSQ

0.98

Bond and Fox (2015, p. 270) describe 0.5–1.5 as productive for measurement and 0.7–1.3 as the more demanding band. The aggregate sits comfortably inside the tighter band on every form.

Calibration summary

What this indicates

On every operational form the Rasch model converges with acceptable person and item separation reliabilities and well-behaved fit. The form-by-form numbers below confirm that candidates and items can be placed on a single, common scale.

One row per operational form. Item separation indicates how reliably the form distinguishes items along the difficulty continuum; person separation does the same for candidates.

FormN (kept)Person rel.Marginal rel.Item sep.Mean bb rangeInfitOutfitMisfit items
ATC A2,082 (−21)0.71 · acceptable0.7115.00.00-3.51.90.980.938 / 45 (1↑ / 7↓)
ATC B2,021 (−26)0.77 · acceptable0.7716.50.00-3.31.70.990.974 / 45 (2↑ / 2↓)
ATC C2,069 (−23)0.72 · acceptable0.7215.70.00-4.22.20.990.952 / 45 (0↑ / 2↓)
ATC D1,342 (−21)0.72 · acceptable0.7212.10.00-2.73.80.980.9016 / 45 (4↑ / 12↓)
Pilot A3,628 (−21)0.76 · acceptable0.7622.60.00-3.84.10.980.959 / 45 (2↑ / 7↓)
Pilot B3,552 (−45)0.76 · acceptable0.7619.80.00-2.93.20.980.9310 / 45 (2↑ / 8↓)
Pilot C1,364 (−25)0.74 · acceptable0.747.80.00-5.93.60.960.9011 / 45 (3↑ / 8↓)

Following Bond and Fox (2015, ch. 12), infit/outfit MSQ in 0.5–1.5 is productive for measurement and 0.7–1.3 is the tighter, preferred band; values above the band indicate noisy (under-fitting) items, values below it indicate over-predictable (over-fitting) items that contribute little new information. A small number of margin items per form is normal and expected for an operational bank, and is the routine target of bank maintenance.

Person–item maps (Wright maps)

What this indicates

Candidate ability and item difficulty sit on the same scale, and the bulk of items is targeted at the borderline candidate — exactly what a screening test at the Level 4 floor needs. Precision is strongest where the pass/fail decision is taken.

Candidate ability (θ, left) and item difficulty (b, right) on the same logit axis. Bond and Fox (2015, ch. 4) treat targeting — the alignment of person and item distributions — as the central diagnostic of how well a test measures its intended population. On every form, item difficulties sit below the candidate ability mass: Paper 1 is calibrated to be passable by the target population, but precision is strongest below the typical candidate and weakest at the top — which has direct consequences for any decision taken near the upper end of the scale.

ATC A — person–item map

N = 2,082

Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.

ATC B — person–item map

N = 2,021

Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.

ATC C — person–item map

N = 2,069

Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.

ATC D — person–item map

N = 1,342

Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.

Pilot A — person–item map

N = 3,628

Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.

Pilot B — person–item map

N = 3,552

Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.

Pilot C — person–item map

N = 1,364

Persons (left) and item difficulties (right) on the shared logit scale. A good match means the bulk of candidates and the bulk of items overlap.

Test information functions

What this indicates

Each form measures most precisely in the zone where the pass/fail decision is taken. Above the pass mark precision tapers — a known and accepted limit of a screening test, not a defect in the operational decision.

Each form measures most precisely in the part of the ability range where the pass/fail decision is taken. Precision tapers above the pass mark — Green (2013, p. 78) treats this as acceptable for a screen, not as a defect. The test information function I(θ) is the formal measure of this, with SEM(θ) = 1 / √I(θ) on the logit scale (Bond & Fox, 2015, ch. 4).

ATC A — test information & SEM

45 items

ATC B — test information & SEM

45 items

ATC C — test information & SEM

45 items

ATC D — test information & SEM

45 items

Pilot A — test information & SEM

45 items

Pilot B — test information & SEM

45 items

Pilot C — test information & SEM

45 items

Item characteristic curves

What this indicates

Items behave the way the model expects: easier items produce a higher chance of a correct answer at every ability level, and the three reference items differ only in difficulty, not in shape. The scale behaves consistently across items.

Three reference items per form — easiest, median and hardest — illustrating the Rasch ICC shape. All ICCs share the same logistic slope; they differ only in horizontal position. Bond and Fox (2015, ch. 12) point out that this equal-slope constraint is the price of Rasch's specific objectivity: the CTT point-biserials show some variation across items, so a 2PL would fit marginally better, but only Rasch supports invariant person and item measures on a common scale.

ATC A — item characteristic curves

3 reference items

Easiest (b = -3.491), median (b = 0.128), hardest (b = 1.87). Curves shift right with item difficulty.

ATC B — item characteristic curves

3 reference items

Easiest (b = -3.292), median (b = 0.084), hardest (b = 1.699). Curves shift right with item difficulty.

ATC C — item characteristic curves

3 reference items

Easiest (b = -4.182), median (b = 0.402), hardest (b = 2.183). Curves shift right with item difficulty.

ATC D — item characteristic curves

3 reference items

Easiest (b = -2.736), median (b = -0.106), hardest (b = 3.747). Curves shift right with item difficulty.

Pilot A — item characteristic curves

3 reference items

Easiest (b = -3.834), median (b = 0.011), hardest (b = 4.12). Curves shift right with item difficulty.

Pilot B — item characteristic curves

3 reference items

Easiest (b = -2.887), median (b = -0.315), hardest (b = 3.158). Curves shift right with item difficulty.

Pilot C — item characteristic curves

3 reference items

Easiest (b = -5.867), median (b = 0.161), hardest (b = 3.567). Curves shift right with item difficulty.

Per-item fit (expandable)

What this indicates

The large majority of items fit the model in the preferred band. A small number of margin items is identified per form; these are the routine targets of item-bank maintenance, not defects in the operational decision.

Full Rasch item table per form. Verdict applies the Bond & Fox (2015) fit bands: excellent 0.8–1.2, acceptable 0.7–1.3, productive (review) 0.5–1.5, otherwise misfit.

ATC A45 items (8 flagged)
ItembSEInfitOutfitVerdict
2000380-3.490.240.900.55productive (review)
2000381-2.480.150.910.62productive (review)
20003820.130.060.940.94excellent
20003830.230.060.970.96excellent
2000384-1.900.120.920.75acceptable
2000385-1.350.090.910.69productive (review)
20003001.070.051.031.03excellent
20003011.200.051.201.27acceptable
20003020.890.050.980.96excellent
2000309-0.770.070.940.88excellent
20003101.400.051.121.16excellent
20003110.330.060.930.89excellent
2000197-0.870.080.940.84excellent
2000198-1.990.120.940.91excellent
2000199-1.340.090.850.56productive (review)
20003220.180.061.151.14excellent
20003231.200.051.031.05excellent
2000324-0.330.070.880.75acceptable
2000325-0.600.070.810.57productive (review)
20003260.430.050.910.88excellent
20003420.030.061.121.27acceptable
20003430.170.060.820.72acceptable
20003441.440.051.191.27acceptable
2000345-0.190.061.061.14excellent
2000346-1.210.090.910.75acceptable
20003480.460.050.970.95excellent
20003490.510.050.970.94excellent
2000350-0.680.071.031.16excellent
20003511.660.051.111.15excellent
20003520.040.061.031.02excellent
20003530.150.060.920.86excellent
20003540.680.051.031.02excellent
20003551.500.051.081.10excellent
2000356-1.040.080.880.69productive (review)
20003581.140.050.980.95excellent
20003950.090.061.000.97excellent
20003960.990.050.890.87excellent
20003970.040.061.010.95excellent
20003981.870.051.131.20excellent
20004000.810.051.071.04excellent
2000271-0.040.060.860.76acceptable
2000272-1.210.090.980.91excellent
2000273-0.160.060.820.66productive (review)
2000274-0.060.060.960.90excellent
20002751.070.051.241.31productive (review)
ATC B45 items (4 flagged)
ItembSEInfitOutfitVerdict
2000454-3.000.170.920.86excellent
2000455-3.290.190.910.81excellent
20004561.560.051.071.08excellent
2000457-0.050.061.051.12excellent
2000458-0.620.071.081.20acceptable
2000459-1.500.090.950.92excellent
20003171.030.050.970.96excellent
20003181.020.050.970.95excellent
20003190.650.051.031.00excellent
20003140.210.050.990.96excellent
2000315-0.530.061.000.92excellent
2000316-0.940.070.880.71acceptable
20002011.180.050.990.98excellent
2000202-2.000.100.920.82excellent
2000203-0.090.060.910.81excellent
20004600.810.050.980.95excellent
2000461-0.610.061.001.08excellent
2000462-1.370.080.900.70productive (review)
20004630.170.051.071.06excellent
20004640.780.051.251.32productive (review)
20004650.960.050.920.89excellent
2000466-0.790.070.820.63productive (review)
2000467-1.200.080.900.86excellent
20004681.700.051.131.23acceptable
20004691.070.051.041.06excellent
2000470-0.490.061.241.64misfit
20004710.490.050.991.00excellent
20004721.320.051.011.02excellent
20004730.900.051.051.05excellent
20004740.010.060.990.94excellent
20004750.390.050.920.92excellent
20004761.440.051.151.20excellent
2000477-0.200.061.081.13excellent
2000478-0.480.060.910.84excellent
2000479-1.070.070.900.71acceptable
2000482-0.360.060.920.87excellent
2000483-0.210.060.930.88excellent
20004841.460.051.151.20acceptable
20004851.300.050.960.95excellent
20004860.080.051.111.22acceptable
20004870.150.050.940.87excellent
2000488-0.550.060.920.79acceptable
20004890.990.050.900.87excellent
2000490-0.430.060.840.71acceptable
20004910.090.050.920.86excellent
ATC C45 items (2 flagged)
ItembSEInfitOutfitVerdict
2000493-4.180.300.900.39misfit
2000494-3.150.180.941.11excellent
2000495-0.150.060.991.00excellent
2000496-0.420.060.960.92excellent
2000497-0.160.061.031.04excellent
2000498-0.580.070.920.82excellent
2000499-0.410.060.970.96excellent
20005000.780.051.111.11excellent
2000501-1.760.100.910.77acceptable
20005020.360.051.241.26acceptable
20005031.030.051.091.09excellent
2000504-0.230.061.011.02excellent
2000505-1.570.090.950.85excellent
2000506-2.180.120.850.46misfit
20005070.440.051.081.13excellent
20005081.880.051.011.02excellent
20005090.920.050.970.98excellent
2000510-0.300.060.910.80acceptable
20005111.070.051.111.15excellent
20005120.480.051.041.06excellent
2000513-0.920.071.021.01excellent
2000514-1.580.090.930.82excellent
20005150.400.051.000.96excellent
20005160.410.050.940.94excellent
20005172.180.050.991.01excellent
20005180.760.051.041.05excellent
20005190.640.050.910.88excellent
20005200.380.051.011.02excellent
20005210.820.050.940.92excellent
20005220.950.051.051.06excellent
20005231.140.050.880.86excellent
20005240.610.050.970.96excellent
20005250.480.050.900.89excellent
20005260.040.050.880.85excellent
20005270.720.050.890.87excellent
2000529-1.120.080.920.76acceptable
20005300.200.050.990.98excellent
20005310.630.051.031.05excellent
2000532-1.760.100.920.71acceptable
20005330.220.050.940.93excellent
20005360.600.051.061.10excellent
20005371.670.051.021.03excellent
2000538-0.240.061.051.03excellent
20005390.470.051.151.21acceptable
20005400.410.051.061.12excellent
ATC D45 items (16 flagged)
ItembSEInfitOutfitVerdict
2100009-2.110.210.910.79acceptable
2100010-1.940.200.960.88excellent
2100011-0.860.130.970.89excellent
2100012-1.410.160.950.76acceptable
2100013-2.740.280.920.59productive (review)
2100014-2.260.230.930.47misfit
2100015-1.320.151.011.36productive (review)
21000161.900.060.910.89excellent
2100017-0.020.091.151.38productive (review)
2100018-1.540.160.960.88excellent
2100019-1.480.160.940.68productive (review)
2100020-0.800.121.001.01excellent
2100021-0.510.110.900.73acceptable
2100022-1.410.161.021.08excellent
2100023-0.960.130.900.62productive (review)
21000242.140.060.960.94excellent
2100025-0.070.090.910.69productive (review)
2100026-1.680.170.860.44misfit
2100027-0.560.110.870.50productive (review)
21000281.590.060.950.93excellent
2100077-0.110.100.960.72acceptable
21000783.750.071.021.31productive (review)
21000790.690.081.121.21acceptable
2100080-0.190.100.920.88excellent
21000811.470.071.151.18excellent
2100034-1.890.190.941.04excellent
21000351.580.061.201.27acceptable
21000362.350.061.361.53misfit
2100037-0.600.110.920.69productive (review)
21000380.870.070.880.77acceptable
21000390.110.091.141.13excellent
2100040-0.890.130.880.55productive (review)
2100041-0.420.110.870.59productive (review)
2100042-0.840.120.940.88excellent
21000430.060.090.890.65productive (review)
21000450.730.070.890.79acceptable
21000463.250.061.091.29acceptable
21000470.270.090.940.98excellent
21000480.990.070.950.89excellent
21000490.600.081.091.14excellent
2100052-0.020.090.910.71acceptable
21000531.370.071.071.08excellent
2100054-0.280.100.900.64productive (review)
21000551.160.070.960.89excellent
21000562.020.061.101.17excellent
Pilot A45 items (9 flagged)
ItembSEInfitOutfitVerdict
2100153-0.490.060.910.79acceptable
2100154-0.480.061.031.14excellent
2100155-1.440.080.970.94excellent
2100156-0.740.061.041.13excellent
2100157-0.660.060.990.92excellent
2100158-2.830.150.920.62productive (review)
2100159-3.830.230.910.68productive (review)
2100160-2.780.140.940.85excellent
21001611.640.040.860.83excellent
2100162-2.080.100.930.59productive (review)
2100163-3.380.190.920.81excellent
21001643.350.040.971.04excellent
21001650.970.040.900.82excellent
21001660.010.050.990.96excellent
21001670.290.040.900.82excellent
21001701.450.041.031.03excellent
21001710.100.041.091.18excellent
21001721.290.041.231.28acceptable
2100173-1.220.070.980.88excellent
2100174-0.840.060.870.60productive (review)
21001750.150.040.910.73acceptable
2100176-0.250.050.950.91excellent
21001770.020.051.021.07excellent
2100178-1.590.080.930.58productive (review)
21001790.070.051.010.96excellent
2100180-0.000.051.081.21acceptable
21001810.390.040.750.59productive (review)
21001820.420.040.940.87excellent
2100183-2.270.110.910.54productive (review)
21001841.420.041.011.00excellent
21001850.300.040.980.98excellent
2100186-1.300.071.011.21acceptable
21001871.020.041.101.12excellent
21001881.620.040.970.96excellent
21001893.270.041.151.57misfit
21001914.120.051.131.70misfit
21001922.520.041.171.26acceptable
2100193-0.360.051.020.96excellent
2100194-0.710.061.020.97excellent
2100195-0.040.050.950.83excellent
2100196-0.270.051.031.08excellent
21001971.240.040.880.84excellent
21001981.920.041.011.02excellent
2100199-0.670.060.991.02excellent
21002000.650.040.970.95excellent
Pilot B45 items (10 flagged)
ItembSEInfitOutfitVerdict
2100201-2.770.170.930.83excellent
2100202-1.720.110.981.04excellent
2100203-1.080.080.971.13excellent
2100204-1.420.100.960.90excellent
2100205-2.010.120.971.23acceptable
2100206-2.750.170.951.43productive (review)
2100207-2.890.180.920.77acceptable
21002080.890.040.930.85excellent
21002091.020.041.091.12excellent
21002101.120.041.021.00excellent
2100211-2.400.140.940.99excellent
2100212-0.550.070.991.08excellent
21002130.170.050.960.87excellent
21002141.220.041.161.25acceptable
2100215-0.320.060.950.94excellent
21002161.550.040.890.86excellent
2100217-0.400.060.930.68productive (review)
21002183.160.040.950.93excellent
2100219-0.440.060.950.85excellent
21002202.680.041.401.68misfit
2100221-0.270.060.980.92excellent
21002221.760.040.880.83excellent
21002230.760.041.060.95excellent
2100224-0.880.070.981.14excellent
21002251.540.040.890.84excellent
21002261.710.041.181.25acceptable
21002270.890.041.030.98excellent
21002282.120.041.011.03excellent
21002292.560.041.021.04excellent
2100230-0.320.060.920.73acceptable
2100231-0.890.070.920.62productive (review)
21002321.080.040.970.94excellent
2100233-1.010.080.880.49misfit
21002341.090.041.081.06excellent
21002351.030.040.970.93excellent
2100236-1.310.090.930.54productive (review)
2100237-0.460.061.051.09excellent
21002381.670.041.101.11excellent
21002392.270.041.001.01excellent
2100240-0.520.070.930.78acceptable
21002410.030.050.930.77acceptable
2100242-0.470.060.950.65productive (review)
2100243-0.740.070.920.66productive (review)
2100244-2.230.130.880.54productive (review)
2100245-2.470.150.900.34misfit
Pilot C45 items (11 flagged)
ItembSEInfitOutfitVerdict
21002490.400.080.990.92excellent
2100250-1.440.150.990.93excellent
2100251-1.360.150.970.87excellent
2100252-1.640.170.960.83excellent
2100253-1.890.180.910.64productive (review)
2100254-1.480.150.920.65productive (review)
2100255-3.250.340.900.47misfit
2100256-0.820.120.970.97excellent
21002570.200.081.051.13excellent
2100258-2.160.200.920.78acceptable
2100259-3.950.480.860.21misfit
2100260-5.871.240.000.00misfit
21002673.570.070.981.24acceptable
2100268-0.110.091.071.37productive (review)
2100269-0.280.091.031.01excellent
21002640.060.080.940.80acceptable
21002650.850.071.000.90excellent
21002662.840.060.850.81excellent
2100270-1.510.150.930.78acceptable
21002710.290.080.900.74acceptable
21002722.200.061.091.14excellent
21002730.160.080.820.61productive (review)
21002741.940.061.071.10excellent
21002751.770.061.211.33productive (review)
21002760.830.070.940.88excellent
21002771.010.070.960.87excellent
21002782.210.061.031.08excellent
2100279-0.580.100.920.73acceptable
21002801.280.060.920.88excellent
21002810.270.080.970.95excellent
2100282-1.820.170.910.58productive (review)
21002831.270.061.081.16excellent
21002840.140.081.061.07excellent
21002851.750.060.810.77acceptable
21002860.350.081.061.04excellent
21002871.060.071.281.54misfit
2100288-0.670.110.890.60productive (review)
2100289-0.730.111.001.22acceptable
21002901.860.061.021.03excellent
21002911.340.060.970.95excellent
21002920.040.080.960.90excellent
2100293-0.230.090.980.87excellent
21002941.000.071.051.04excellent
21002950.110.080.970.88excellent
21002960.970.071.161.27acceptable

Method

  • Model — Rasch / one-parameter logistic: P(X = 1) = exp(θ − b) / (1 + exp(θ − b)) (Bond & Fox, 2015, ch. 3).
  • Estimation — joint maximum likelihood with Wright bias correction (×(K−1)/K for items, ×(N−1)/N for persons; Bond & Fox, 2015, app. B). Item difficulties centred at 0.
  • Extreme scorers — candidates with raw score 0 or full are excluded from θ estimation (their MLE is undefined). Counts shown in the summary table.
  • Fit statistics — infit (info-weighted) and outfit (unweighted) mean-square residuals; 0.5–1.5 productive, 0.7–1.3 preferred (Bond & Fox, 2015, ch. 12).
  • Reliability — person and item separation reliabilities derived from the variance of estimates and their standard errors (Bond & Fox, 2015, ch. 4).

Continuous-improvement programme

The items below sit on the published improvement roadmap. They describe planned methodological work, not operational deficiencies in the current pass/fail decision.

  • Invariance — not tested. Rasch invariance evidence beyond per-form difficulty contrasts (anchored between-form calibration) remains the highest-priority item in the ongoing validation programme (Bond & Fox, 2015, ch. 6).
  • Targeting mismatch at the top. Person means sit above item means on every form, so the test information functions show that measurement is least precise exactly where high-stakes decisions about strong candidates would be taken.
  • No equating across forms. Each form is calibrated separately and difficulties are centred at 0 per form; without common-item or concurrent equating (Bond & Fox, 2015, ch. 5), b-parameters are not directly comparable across forms.
  • 1PL is a modelling choice. The CTT point-biserials show some discrimination variation, so a 2PL would fit marginally better; Rasch is retained for the measurement properties it guarantees (invariance, conjoint additivity), not because it is the best descriptive fit.

Conclusion

At the aggregate level the Rasch model fits: candidates and items sit on the same interval scale, fit indices are within acceptable bounds, person and item separation are acceptable for a screening test, and item difficulties span the right range (Bond & Fox, 2015). That strengthens the "do we score the test accurately?" answer beyond what CTT alone can give (Green, 2013). The questions about whether forms are interchangeable and whether the score predicts radiotelephony performance still need the equating study and the criterion evidence flagged in the gaps section above.

Glossary of terms
Cronbach's α
Internal-consistency reliability on 0/1 item scores. Green (2013) cites ≥ 0.80 as acceptable for screening and ≥ 0.90 for high-stakes decisions.
McDonald's ω
Reliability estimate that relaxes the tau-equivalence assumption of α.
SEM
Standard error of measurement on the raw-score scale. A candidate's true score lies within roughly ±2·SEM of their observed score with 95% confidence.
Conditional SEM (cSEM)
SEM as a function of score (CTT) or ability θ (Rasch). Green (2013, p. 78) argues this is the relevant precision metric for decisions: what matters is precision at the cut, not on average.
Infit / outfit MSQ
Rasch fit statistics. Bond and Fox (2015, ch. 12) treat 0.5–1.5 as productive for measurement and 0.7–1.3 as the preferred tighter band.
Person / item separation
Rasch reliability analogues: how reliably the form distinguishes candidates along the ability continuum, and items along the difficulty continuum.
Marginal reliability
Population-level reliability of the Rasch ability estimates, reported alongside α for direct comparison.
Targeting
Alignment of person and item distributions on the logit scale (Bond & Fox, 2015, ch. 4). Good targeting means most candidates encounter items near their ability level.
Decision consistency (p₀, κ)
Probability that two parallel forms would classify a candidate the same way at the cut score, and Cohen's κ adjusting for chance agreement.

References

  • Bond, T. G., & Fox, C. M. (2015). Applying the Rasch model: Fundamental measurement in the human sciences (3rd ed.). Routledge.
  • Green, A. (2013). Exploring language assessment and testing. Routledge.