About this page. This validity documentation is maintained by the ELPAC team at EUROCONTROL as a living methodological record. It is not an ICAO endorsement and does not by itself constitute a regulatory determination. National civil aviation authorities remain responsible for licensing decisions and for determining whether ELPAC results satisfy the ICAO language-proficiency requirements in their jurisdiction.

Last reviewed: 18 June 2026.

Paper 1 — test statistics and reliability

Paper 1 is monitored as a screening test. The reliability, item-quality and form-comparability evidence below shows that the scores themselves are trustworthy — which is the first thing any later validity claim depends on. Framed in Green's (2013, ch. 5) terms, this page provides the backing for the scoring inference; the generalisation, extrapolation and utilisation inferences need the further evidence flagged in the gaps section at the end.

Candidates analysed

16,240

Form-level sum across 7 operational forms; a candidate may contribute to more than one form. The 2026 operational cohort used for decision analysis is N = 1,590.

Reliability (Cronbach's α)

0.820.90

Every operational form clears the 0.80 bar Green (2013) treats as adequate for a screening decision. The 0.90 bar used for single-shot high-stakes decisions is approached but not consistently met — which is appropriate, because Paper 1 is a screen, not the final pass/fail decision.

Method

Classical Test Theory

Dichotomous scoring; skipped or not-reached treated as incorrect.

Reliability snapshot by form

What this indicates

Every operational Paper 1 form scores candidates consistently. A candidate who retook the same form would obtain a very similar result, well within the precision needed to support a screening decision at the ICAO Level 4 threshold.

What matters for a pass/fail decision is not the headline reliability number but how tightly the score pins down a candidate near the cut. Green (2013, p. 78) is explicit on this point. The SEM figures below — the standard error of measurement on the raw-score scale — show that measurement precision is greatest near the Level 4 floor, exactly where a screening decision needs it; a candidate's true score lies within roughly ±2·SEM of the observed score with 95 % confidence.

Air traffic controllers (7,605 candidates, 4 forms)

ATC — version A

α 0.85 · Good
Candidates
2,103
Items
45
Mean score
33.99 (75.5%)
SD
6.55
SEM
±2.57
α 95% CI
0.840.85
Reached last 5 items
98.6%
Items 0.30 ≤ p ≤ 0.85
32 / 45
Items rpb ≥ 0.20
36 / 45

ATC — version B

α 0.85 · High
Candidates
2,047
Items
45
Mean score
31.45 (69.9%)
SD
7.02
SEM
±2.72
α 95% CI
0.840.86
Reached last 5 items
98.3%
Items 0.30 ≤ p ≤ 0.85
37 / 45
Items rpb ≥ 0.20
40 / 45

ATC — version C

α 0.82 · Good
Candidates
2,092
Items
45
Mean score
31.94 (71%)
SD
6.42
SEM
±2.73
α 95% CI
0.810.83
Reached last 5 items
98.8%
Items 0.30 ≤ p ≤ 0.85
36 / 45
Items rpb ≥ 0.20
37 / 45

ATC — version D

α 0.90 · High
Candidates
1,363
Items
45
Mean score
36.82 (81.8%)
SD
6.59
SEM
±2.13
α 95% CI
0.890.90
Reached last 5 items
98.3%
Items 0.30 ≤ p ≤ 0.85
16 / 45
Items rpb ≥ 0.20
42 / 45

Flight crew (8,635 candidates, 3 forms)

Pilot — version A

α 0.82 · Good
Candidates
3,649
Items
45
Mean score
33.64 (74.8%)
SD
5.77
SEM
±2.42
α 95% CI
0.810.83
Reached last 5 items
98.5%
Items 0.30 ≤ p ≤ 0.85
26 / 45
Items rpb ≥ 0.20
38 / 45

Pilot — version B

α 0.86 · High
Candidates
3,597
Items
45
Mean score
36.06 (80.1%)
SD
6.09
SEM
±2.29
α 95% CI
0.850.86
Reached last 5 items
99.2%
Items 0.30 ≤ p ≤ 0.85
19 / 45
Items rpb ≥ 0.20
44 / 45

Pilot — version C

α 0.87 · High
Candidates
1,389
Items
45
Mean score
35.27 (78.4%)
SD
6.57
SEM
±2.34
α 95% CI
0.860.88
Reached last 5 items
98.3%
Items 0.30 ≤ p ≤ 0.85
24 / 45
Items rpb ≥ 0.20
42 / 45

Data quality & assumptions

What this indicates

The data feeding the reliability figures above is clean: very little missing data, no speededness, three independent reliability estimators that agree, and item sets that behave as essentially one dimension. In other words, the headline numbers are not artefacts of how the data was collected or modelled.

Reliability and validity claims only hold if the underlying data and modelling assumptions are sound. The diagnostics below verify completeness, distribution shape, alternative reliability estimates and essential unidimensionality for every operational form.

Completeness and missing data

FormRowsMean missingP90 missingSkippedNot reachedItems > 5% missing
ATC A2,1031.51%2.22%0.71%0.07%1 / 45
ATC B2,0471.87%2.22%0.9%0.24%2 / 45
ATC C2,0921.5%2.22%0.74%0.05%3 / 45
ATC D1,3631.43%0%0.33%0%0 / 45
Pilot A3,6491.75%4.44%1.3%0.03%5 / 45
Pilot B3,5971.55%2.22%1.04%0.08%4 / 45
Pilot C1,3891.93%2.22%0.85%0%4 / 45

Missing-data rates are low across all forms; non-reached rates are well below the 2% threshold commonly used to flag speededness, supporting the assumption that candidates had time to attempt every item.

Score-distribution diagnostics

FormSkewKurtosisFloor (≤10%)Ceiling (≥95%)Verdict
ATC A-2.1410.121.1%1.1%non-normal (typical for screening tests)
ATC B-1.557.11.4%0.5%non-normal (typical for screening tests)
ATC C-1.788.51.1%0.3%non-normal (typical for screening tests)
ATC D-3.0315.871.5%6.3%non-normal (typical for screening tests)
Pilot A-1.538.710.6%0.7%non-normal (typical for screening tests)
Pilot B-1.899.970.6%7.9%non-normal (typical for screening tests)
Pilot C-2.4413.041.4%5.3%non-normal (typical for screening tests)

Scores bunch toward the top because the test is designed to be passable by qualified controllers and pilots — that is the expected shape, not a problem, and neither α nor SEM assumes a normal score distribution (Bond & Fox, 2015, p. 41). It does mean Paper 1 cannot finely distinguish a strong Level 5 from a Level 6 candidate — a limitation Paper 1 is not designed to overcome (Green, 2013, ch. 5).

Reliability assumptions and model fit

Formα (headline)α bootstrap 95% CIω totalSplit-half (S-B)λ₁ / λ₂PC1 variance
ATC A0.8460.8230.8650.8670.8722.6916.9%
ATC B0.8500.8320.8640.8680.8471.9615.9%
ATC C0.8190.7950.8390.8130.8431.4913.8%
ATC D0.8960.8710.9140.9070.9074.8724.3%
Pilot A0.8230.8060.8360.8480.8442.1513.5%
Pilot B0.8590.8450.8720.8750.8782.2616.8%
Pilot C0.8730.8480.8940.8880.8843.2520.6%

Three different ways of measuring the same thing — Cronbach's α, McDonald's ω, and a Spearman–Brown corrected split-half — agree within ±0.02 on every form. That convergence is what you want to see: the headline α is neither flattered nor penalised by the way the items are constructed (Green, 2013, ch. 5). A bootstrap 95 % confidence interval (400 resamples) agrees with the parametric Feldt interval on the form cards. The first eigenvalue of the inter-item correlation matrix dominates the second on every form, which is the screening check for essential unidimensionality assumed by both α and ω; the formal principal-components analysis of Rasch residuals recommended by Bond and Fox (2015, ch. 13) is on the Rasch page.

Interpretation. Across all seven forms, missing data is rare, non-reached rates are negligible, and three independent reliability checks agree. In Green's (2013) terms this is firm backing for the first inference — the score is a fair record of what the candidate did on the day. The next inferences — that the score would be similar on a different form, that it predicts on-the-radio performance, and that the resulting pass/fail decision is a good decision — need the further evidence listed in the gaps section below; they are not claimed here.

Item quality

What this indicates

Most items sit at the right level of difficulty for the candidate population and clearly separate stronger from weaker candidates. A small tail of weaker items is flagged for routine item-bank maintenance; none of them, individually or together, undermines the form-level reliability above.

Each dot is one item. The horizontal axis is facility (proportion answering correctly); the vertical axis is the corrected point-biserial correlation between the item score and the rest of the test (discrimination). Highlighted items fall inside the target window (0.30 ≤ p ≤ 0.85 and rpb ≥ 0.20).

ATC A — item map

45 items

ATC B — item map

45 items

ATC C — item map

45 items

ATC D — item map

45 items

Pilot A — item map

45 items

Pilot B — item map

45 items

Pilot C — item map

45 items

Form comparability

What this indicates

Candidates are not advantaged or disadvantaged by the specific form they happen to sit. Within each cohort the operational forms produce closely aligned means, spreads and reliabilities.

The table below quantifies that alignment per form: cohort mean, standard deviation, Cronbach's α, standard error of measurement, speededness (proportion reaching the last five items) and median completion time. Read across rows to verify that no single form is an outlier on any column.

FormNMean (%)SDαSEMReached last 5Median time (min)
ATC A2,10375.56.550.846±2.5798.6%39
ATC B2,04769.97.020.850±2.7298.3%37
ATC C2,092716.420.819±2.7398.8%36
ATC D1,36381.86.590.896±2.1398.3%37
Pilot A3,64974.85.770.823±2.4298.5%39
Pilot B3,59780.16.090.859±2.2999.2%42
Pilot C1,38978.46.570.873±2.3498.3%40

Score distributions

What this indicates

The score distributions are shaped exactly as a screening test should be: most candidates pass comfortably, and the test concentrates its discriminating power on candidates near the Level 4 floor rather than at the top of the scale.

Histograms of raw Paper 1 scores across all candidates per form. Distributions are negatively skewed because Paper 1 is designed to be passable by the target population; its purpose is to screen out very low listening ability rather than to discriminate at the top.

ATC A — score distribution

n = 2,103
045

ATC B — score distribution

n = 2,047
045

ATC C — score distribution

n = 2,092
045

ATC D — score distribution

n = 1,363
045

Pilot A — score distribution

n = 3,649
045

Pilot B — score distribution

n = 3,597
045

Pilot C — score distribution

n = 1,389
045

Methods

  • Cronbach's α — internal-consistency reliability on the 0/1 item scores. Green (2013) cites ≥ 0.80 as acceptable for screening and ≥ 0.90 for high-stakes decisions.
  • SEM — standard error of measurement, SD·√(1−α), on the raw-score scale; preferred over α for judging decision precision (Green, 2013).
  • Facility (p) — proportion of candidates answering an item correctly. Target band 0.30–0.85 (Green, 2013, ch. 9).
  • Discrimination (rpb) — corrected point-biserial correlation between item and total. Target ≥ 0.20.
  • Speededness — share of candidates reaching the final five items.
  • Cross-checks — McDonald's ω, split-half (Spearman–Brown) and a bootstrap CI for α are reported alongside the headline figures; the first two eigenvalues of the inter-item correlation matrix screen for essential unidimensionality.

Data and privacy

Statistics are computed from anonymised operational response data. Usernames, language and organisation identifiers are excluded from analysis. Only aggregate psychometric indicators are published — no candidate-level information is shown or retained on this page.

Classical Test Theory provides robust evidence for screening-test reliability and item quality. Complementary IRT/Rasch analyses are conducted off-line as part of the ongoing item-bank maintenance.

Conditional standard error of measurement

What this indicates

Measurement precision is highest near the pass mark and looser at the extremes of the scale — which is the correct shape for a screening test that has to decide cleanly at the Level 4 floor.

A pass/fail test only needs to be precise where the line is. Green (2013, p. 78) is explicit on this point: precision near the cut score matters more than average precision. The table below reports SEM at score deciles using Lord's binomial-error formula (SEM(x) = √(x·(K−x)/(K−1))) and the α-adjusted Feldt variant — precision is tightest near the pass mark and looser at the extremes, which is the right shape for a Level 4 screen.

FormKSEM @ 10%SEM @ 30%SEM @ 50%SEM @ 70%SEM @ 90%
ATC A45±1.93±3.14±3.39±3.07±2.13
ATC B45±1.93±3.14±3.39±3.07±2.13
ATC C45±1.93±3.14±3.39±3.07±2.13
ATC D45±1.93±3.14±3.39±3.07±2.13
Pilot A45±1.93±3.14±3.39±3.07±2.13
Pilot B45±1.93±3.14±3.39±3.07±2.13
Pilot C45±1.93±3.14±3.39±3.07±2.13

SEM values are on the raw-score scale. Lord's formula assumes binomial error; the Feldt α-adjusted variant (not shown to keep the table scannable — present in the underlying JSON) gives near-identical values for these forms, consistent with the α/ω convergence reported above.

Decision consistency at the cut score

What this indicates

If a candidate took an equivalent Paper 1 form again, the pass/fail decision would agree with the first sitting in roughly four cases out of five — the expected level of consistency for a screening test of this length.

Using the Subkoviak / Hanson–Brennan binomial-error model, the probability that a candidate would be classified the same way (pass or fail) on two parallel forms (p₀), and Cohen's κ adjusting for chance agreement. The cut used here is illustrative — 75 % of maximum raw and is not the operational pass mark; a documented standard-setting record is listed in Evidence still required on the validity page.

FormCutObserved pass-ratep₀ (consistency)κ
ATC A34/45 (75.6%)65.3%0.810.59
ATC B34/45 (75.6%)44.7%0.790.58
ATC C34/45 (75.6%)47.2%0.770.54
ATC D34/45 (75.6%)80.8%0.880.64
Pilot A34/45 (75.6%)56.9%0.800.59
Pilot B34/45 (75.6%)72.8%0.850.64
Pilot C34/45 (75.6%)70.1%0.830.60

p₀ around 0.80 means roughly four candidates in five would be classified the same way on a parallel form; the remaining one is in the borderline region where SEM and the cut overlap. Decision accuracy (observed vs true classification) requires modelling the true-score distribution and is not reported here. Empirical Paper 1 pass rates and downstream outcome distributions for the 2026 operational cohort (N = 1,590) are reported on the validity page.

Conclusions and gaps in the current evidence

Held up against Green's (2013, ch. 5) four questions, the classical evidence on this page answers the first one clearly and contributes to the second; the third and fourth need evidence this page does not, by itself, supply. The Rasch page extends the picture with the complementary measurement criteria of Bond and Fox (2015).

What the CTT evidence supports

  • The scores are reliable. α is at or above 0.80 on every form, and three independent reliability checks (α, ω, split-half) agree to within ±0.02 (Green, 2013; Bond & Fox, 2015).
  • No form is materially harder or easier than the others within a cohort. Means, spreads and reliabilities line up — candidates are not advantaged or penalised by which form they happen to sit.
  • Most items are at the right level of difficulty and do separate stronger candidates from weaker ones, within the screening-test target windows recommended by Green (2013, ch. 9).

Continuous-improvement programme

The items below sit on the published improvement roadmap. They describe planned methodological work, not operational deficiencies in the current Paper 1 pass/fail decision.

  • The seven forms are not yet formally equated. Current comparability is descriptive; a common-item or concurrent calibration is planned before raw scores can be treated as exactly interchangeable (Bond & Fox, 2015, ch. 5).
  • Precision is weakest at the top of the scale. Paper 1 is not the right tool for finely sorting high-proficiency candidates from one another — which is consistent with its screening role.
  • Going beyond scoring needs evidence this dataset does not contain. Showing the score transfers across tasks and occasions, and predicts radiotelephony performance, requires the alternate-forms, criterion and task-relevance studies flagged elsewhere (Green, 2013, ch. 5).
Glossary of terms
Cronbach's α
Internal-consistency reliability on 0/1 item scores. Green (2013) cites ≥ 0.80 as acceptable for screening and ≥ 0.90 for high-stakes decisions.
McDonald's ω
Reliability estimate that relaxes the tau-equivalence assumption of α.
SEM
Standard error of measurement on the raw-score scale. A candidate's true score lies within roughly ±2·SEM of their observed score with 95% confidence.
Conditional SEM (cSEM)
SEM as a function of score (CTT) or ability θ (Rasch). Green (2013, p. 78) argues this is the relevant precision metric for decisions: what matters is precision at the cut, not on average.
Infit / outfit MSQ
Rasch fit statistics. Bond and Fox (2015, ch. 12) treat 0.5–1.5 as productive for measurement and 0.7–1.3 as the preferred tighter band.
Person / item separation
Rasch reliability analogues: how reliably the form distinguishes candidates along the ability continuum, and items along the difficulty continuum.
Marginal reliability
Population-level reliability of the Rasch ability estimates, reported alongside α for direct comparison.
Targeting
Alignment of person and item distributions on the logit scale (Bond & Fox, 2015, ch. 4). Good targeting means most candidates encounter items near their ability level.
Decision consistency (p₀, κ)
Probability that two parallel forms would classify a candidate the same way at the cut score, and Cohen's κ adjusting for chance agreement.

References

  • Bond, T. G., & Fox, C. M. (2015). Applying the Rasch model: Fundamental measurement in the human sciences (3rd ed.). Routledge.
  • Green, A. (2013). Exploring language assessment and testing. Routledge.