About this page. This validity documentation is maintained by the ELPAC team at EUROCONTROL as a living methodological record. It is not an ICAO endorsement and does not by itself constitute a regulatory determination. National civil aviation authorities remain responsible for licensing decisions and for determining whether ELPAC results satisfy the ICAO language-proficiency requirements in their jurisdiction.
Last reviewed: 18 June 2026.
Paper 1 — test statistics and reliability
Paper 1 is monitored as a screening test. The reliability, item-quality and form-comparability evidence below shows that the scores themselves are trustworthy — which is the first thing any later validity claim depends on. Framed in Green's (2013, ch. 5) terms, this page provides the backing for the scoring inference; the generalisation, extrapolation and utilisation inferences need the further evidence flagged in the gaps section at the end.
Candidates analysed
16,240
Form-level sum across 7 operational forms; a candidate may contribute to more than one form. The 2026 operational cohort used for decision analysis is N = 1,590.
Reliability (Cronbach's α)
0.82–0.90
Every operational form clears the 0.80 bar Green (2013) treats as adequate for a screening decision. The 0.90 bar used for single-shot high-stakes decisions is approached but not consistently met — which is appropriate, because Paper 1 is a screen, not the final pass/fail decision.
Method
Classical Test Theory
Dichotomous scoring; skipped or not-reached treated as incorrect.
Reliability snapshot by form
Every operational Paper 1 form scores candidates consistently. A candidate who retook the same form would obtain a very similar result, well within the precision needed to support a screening decision at the ICAO Level 4 threshold.
What matters for a pass/fail decision is not the headline reliability number but how tightly the score pins down a candidate near the cut. Green (2013, p. 78) is explicit on this point. The SEM figures below — the standard error of measurement on the raw-score scale — show that measurement precision is greatest near the Level 4 floor, exactly where a screening decision needs it; a candidate's true score lies within roughly ±2·SEM of the observed score with 95 % confidence.
Air traffic controllers (7,605 candidates, 4 forms)
ATC — version A
α 0.85 · Good- Candidates
- 2,103
- Items
- 45
- Mean score
- 33.99 (75.5%)
- SD
- 6.55
- SEM
- ±2.57
- α 95% CI
- 0.84 – 0.85
- Reached last 5 items
- 98.6%
- Items 0.30 ≤ p ≤ 0.85
- 32 / 45
- Items rpb ≥ 0.20
- 36 / 45
ATC — version B
α 0.85 · High- Candidates
- 2,047
- Items
- 45
- Mean score
- 31.45 (69.9%)
- SD
- 7.02
- SEM
- ±2.72
- α 95% CI
- 0.84 – 0.86
- Reached last 5 items
- 98.3%
- Items 0.30 ≤ p ≤ 0.85
- 37 / 45
- Items rpb ≥ 0.20
- 40 / 45
ATC — version C
α 0.82 · Good- Candidates
- 2,092
- Items
- 45
- Mean score
- 31.94 (71%)
- SD
- 6.42
- SEM
- ±2.73
- α 95% CI
- 0.81 – 0.83
- Reached last 5 items
- 98.8%
- Items 0.30 ≤ p ≤ 0.85
- 36 / 45
- Items rpb ≥ 0.20
- 37 / 45
ATC — version D
α 0.90 · High- Candidates
- 1,363
- Items
- 45
- Mean score
- 36.82 (81.8%)
- SD
- 6.59
- SEM
- ±2.13
- α 95% CI
- 0.89 – 0.90
- Reached last 5 items
- 98.3%
- Items 0.30 ≤ p ≤ 0.85
- 16 / 45
- Items rpb ≥ 0.20
- 42 / 45
Flight crew (8,635 candidates, 3 forms)
Pilot — version A
α 0.82 · Good- Candidates
- 3,649
- Items
- 45
- Mean score
- 33.64 (74.8%)
- SD
- 5.77
- SEM
- ±2.42
- α 95% CI
- 0.81 – 0.83
- Reached last 5 items
- 98.5%
- Items 0.30 ≤ p ≤ 0.85
- 26 / 45
- Items rpb ≥ 0.20
- 38 / 45
Pilot — version B
α 0.86 · High- Candidates
- 3,597
- Items
- 45
- Mean score
- 36.06 (80.1%)
- SD
- 6.09
- SEM
- ±2.29
- α 95% CI
- 0.85 – 0.86
- Reached last 5 items
- 99.2%
- Items 0.30 ≤ p ≤ 0.85
- 19 / 45
- Items rpb ≥ 0.20
- 44 / 45
Pilot — version C
α 0.87 · High- Candidates
- 1,389
- Items
- 45
- Mean score
- 35.27 (78.4%)
- SD
- 6.57
- SEM
- ±2.34
- α 95% CI
- 0.86 – 0.88
- Reached last 5 items
- 98.3%
- Items 0.30 ≤ p ≤ 0.85
- 24 / 45
- Items rpb ≥ 0.20
- 42 / 45
Data quality & assumptions
The data feeding the reliability figures above is clean: very little missing data, no speededness, three independent reliability estimators that agree, and item sets that behave as essentially one dimension. In other words, the headline numbers are not artefacts of how the data was collected or modelled.
Reliability and validity claims only hold if the underlying data and modelling assumptions are sound. The diagnostics below verify completeness, distribution shape, alternative reliability estimates and essential unidimensionality for every operational form.
Completeness and missing data
| Form | Rows | Mean missing | P90 missing | Skipped | Not reached | Items > 5% missing |
|---|---|---|---|---|---|---|
| ATC A | 2,103 | 1.51% | 2.22% | 0.71% | 0.07% | 1 / 45 |
| ATC B | 2,047 | 1.87% | 2.22% | 0.9% | 0.24% | 2 / 45 |
| ATC C | 2,092 | 1.5% | 2.22% | 0.74% | 0.05% | 3 / 45 |
| ATC D | 1,363 | 1.43% | 0% | 0.33% | 0% | 0 / 45 |
| Pilot A | 3,649 | 1.75% | 4.44% | 1.3% | 0.03% | 5 / 45 |
| Pilot B | 3,597 | 1.55% | 2.22% | 1.04% | 0.08% | 4 / 45 |
| Pilot C | 1,389 | 1.93% | 2.22% | 0.85% | 0% | 4 / 45 |
Missing-data rates are low across all forms; non-reached rates are well below the 2% threshold commonly used to flag speededness, supporting the assumption that candidates had time to attempt every item.
Score-distribution diagnostics
| Form | Skew | Kurtosis | Floor (≤10%) | Ceiling (≥95%) | Verdict |
|---|---|---|---|---|---|
| ATC A | -2.14 | 10.12 | 1.1% | 1.1% | non-normal (typical for screening tests) |
| ATC B | -1.55 | 7.1 | 1.4% | 0.5% | non-normal (typical for screening tests) |
| ATC C | -1.78 | 8.5 | 1.1% | 0.3% | non-normal (typical for screening tests) |
| ATC D | -3.03 | 15.87 | 1.5% | 6.3% | non-normal (typical for screening tests) |
| Pilot A | -1.53 | 8.71 | 0.6% | 0.7% | non-normal (typical for screening tests) |
| Pilot B | -1.89 | 9.97 | 0.6% | 7.9% | non-normal (typical for screening tests) |
| Pilot C | -2.44 | 13.04 | 1.4% | 5.3% | non-normal (typical for screening tests) |
Scores bunch toward the top because the test is designed to be passable by qualified controllers and pilots — that is the expected shape, not a problem, and neither α nor SEM assumes a normal score distribution (Bond & Fox, 2015, p. 41). It does mean Paper 1 cannot finely distinguish a strong Level 5 from a Level 6 candidate — a limitation Paper 1 is not designed to overcome (Green, 2013, ch. 5).
Reliability assumptions and model fit
| Form | α (headline) | α bootstrap 95% CI | ω total | Split-half (S-B) | λ₁ / λ₂ | PC1 variance |
|---|---|---|---|---|---|---|
| ATC A | 0.846 | 0.823 – 0.865 | 0.867 | 0.872 | 2.69 | 16.9% |
| ATC B | 0.850 | 0.832 – 0.864 | 0.868 | 0.847 | 1.96 | 15.9% |
| ATC C | 0.819 | 0.795 – 0.839 | 0.813 | 0.843 | 1.49 | 13.8% |
| ATC D | 0.896 | 0.871 – 0.914 | 0.907 | 0.907 | 4.87 | 24.3% |
| Pilot A | 0.823 | 0.806 – 0.836 | 0.848 | 0.844 | 2.15 | 13.5% |
| Pilot B | 0.859 | 0.845 – 0.872 | 0.875 | 0.878 | 2.26 | 16.8% |
| Pilot C | 0.873 | 0.848 – 0.894 | 0.888 | 0.884 | 3.25 | 20.6% |
Three different ways of measuring the same thing — Cronbach's α, McDonald's ω, and a Spearman–Brown corrected split-half — agree within ±0.02 on every form. That convergence is what you want to see: the headline α is neither flattered nor penalised by the way the items are constructed (Green, 2013, ch. 5). A bootstrap 95 % confidence interval (400 resamples) agrees with the parametric Feldt interval on the form cards. The first eigenvalue of the inter-item correlation matrix dominates the second on every form, which is the screening check for essential unidimensionality assumed by both α and ω; the formal principal-components analysis of Rasch residuals recommended by Bond and Fox (2015, ch. 13) is on the Rasch page.
Item quality
Most items sit at the right level of difficulty for the candidate population and clearly separate stronger from weaker candidates. A small tail of weaker items is flagged for routine item-bank maintenance; none of them, individually or together, undermines the form-level reliability above.
Each dot is one item. The horizontal axis is facility (proportion answering correctly); the vertical axis is the corrected point-biserial correlation between the item score and the rest of the test (discrimination). Highlighted items fall inside the target window (0.30 ≤ p ≤ 0.85 and rpb ≥ 0.20).
ATC A — item map
45 itemsATC B — item map
45 itemsATC C — item map
45 itemsATC D — item map
45 itemsPilot A — item map
45 itemsPilot B — item map
45 itemsPilot C — item map
45 itemsForm comparability
Candidates are not advantaged or disadvantaged by the specific form they happen to sit. Within each cohort the operational forms produce closely aligned means, spreads and reliabilities.
The table below quantifies that alignment per form: cohort mean, standard deviation, Cronbach's α, standard error of measurement, speededness (proportion reaching the last five items) and median completion time. Read across rows to verify that no single form is an outlier on any column.
| Form | N | Mean (%) | SD | α | SEM | Reached last 5 | Median time (min) |
|---|---|---|---|---|---|---|---|
| ATC A | 2,103 | 75.5 | 6.55 | 0.846 | ±2.57 | 98.6% | 39 |
| ATC B | 2,047 | 69.9 | 7.02 | 0.850 | ±2.72 | 98.3% | 37 |
| ATC C | 2,092 | 71 | 6.42 | 0.819 | ±2.73 | 98.8% | 36 |
| ATC D | 1,363 | 81.8 | 6.59 | 0.896 | ±2.13 | 98.3% | 37 |
| Pilot A | 3,649 | 74.8 | 5.77 | 0.823 | ±2.42 | 98.5% | 39 |
| Pilot B | 3,597 | 80.1 | 6.09 | 0.859 | ±2.29 | 99.2% | 42 |
| Pilot C | 1,389 | 78.4 | 6.57 | 0.873 | ±2.34 | 98.3% | 40 |
Score distributions
The score distributions are shaped exactly as a screening test should be: most candidates pass comfortably, and the test concentrates its discriminating power on candidates near the Level 4 floor rather than at the top of the scale.
Histograms of raw Paper 1 scores across all candidates per form. Distributions are negatively skewed because Paper 1 is designed to be passable by the target population; its purpose is to screen out very low listening ability rather than to discriminate at the top.
ATC A — score distribution
n = 2,103ATC B — score distribution
n = 2,047ATC C — score distribution
n = 2,092ATC D — score distribution
n = 1,363Pilot A — score distribution
n = 3,649Pilot B — score distribution
n = 3,597Pilot C — score distribution
n = 1,389Methods
- Cronbach's α — internal-consistency reliability on the 0/1 item scores. Green (2013) cites ≥ 0.80 as acceptable for screening and ≥ 0.90 for high-stakes decisions.
- SEM — standard error of measurement, SD·√(1−α), on the raw-score scale; preferred over α for judging decision precision (Green, 2013).
- Facility (p) — proportion of candidates answering an item correctly. Target band 0.30–0.85 (Green, 2013, ch. 9).
- Discrimination (rpb) — corrected point-biserial correlation between item and total. Target ≥ 0.20.
- Speededness — share of candidates reaching the final five items.
- Cross-checks — McDonald's ω, split-half (Spearman–Brown) and a bootstrap CI for α are reported alongside the headline figures; the first two eigenvalues of the inter-item correlation matrix screen for essential unidimensionality.
Data and privacy
Statistics are computed from anonymised operational response data. Usernames, language and organisation identifiers are excluded from analysis. Only aggregate psychometric indicators are published — no candidate-level information is shown or retained on this page.
Classical Test Theory provides robust evidence for screening-test reliability and item quality. Complementary IRT/Rasch analyses are conducted off-line as part of the ongoing item-bank maintenance.
Conditional standard error of measurement
Measurement precision is highest near the pass mark and looser at the extremes of the scale — which is the correct shape for a screening test that has to decide cleanly at the Level 4 floor.
A pass/fail test only needs to be precise where the line is. Green (2013, p. 78) is explicit on this point: precision near the cut score matters more than average precision. The table below reports SEM at score deciles using Lord's binomial-error formula (SEM(x) = √(x·(K−x)/(K−1))) and the α-adjusted Feldt variant — precision is tightest near the pass mark and looser at the extremes, which is the right shape for a Level 4 screen.
| Form | K | SEM @ 10% | SEM @ 30% | SEM @ 50% | SEM @ 70% | SEM @ 90% |
|---|---|---|---|---|---|---|
| ATC A | 45 | ±1.93 | ±3.14 | ±3.39 | ±3.07 | ±2.13 |
| ATC B | 45 | ±1.93 | ±3.14 | ±3.39 | ±3.07 | ±2.13 |
| ATC C | 45 | ±1.93 | ±3.14 | ±3.39 | ±3.07 | ±2.13 |
| ATC D | 45 | ±1.93 | ±3.14 | ±3.39 | ±3.07 | ±2.13 |
| Pilot A | 45 | ±1.93 | ±3.14 | ±3.39 | ±3.07 | ±2.13 |
| Pilot B | 45 | ±1.93 | ±3.14 | ±3.39 | ±3.07 | ±2.13 |
| Pilot C | 45 | ±1.93 | ±3.14 | ±3.39 | ±3.07 | ±2.13 |
SEM values are on the raw-score scale. Lord's formula assumes binomial error; the Feldt α-adjusted variant (not shown to keep the table scannable — present in the underlying JSON) gives near-identical values for these forms, consistent with the α/ω convergence reported above.
Decision consistency at the cut score
If a candidate took an equivalent Paper 1 form again, the pass/fail decision would agree with the first sitting in roughly four cases out of five — the expected level of consistency for a screening test of this length.
Using the Subkoviak / Hanson–Brennan binomial-error model, the probability that a candidate would be classified the same way (pass or fail) on two parallel forms (p₀), and Cohen's κ adjusting for chance agreement. The cut used here is illustrative — 75 % of maximum raw and is not the operational pass mark; a documented standard-setting record is listed in Evidence still required on the validity page.
| Form | Cut | Observed pass-rate | p₀ (consistency) | κ |
|---|---|---|---|---|
| ATC A | 34/45 (75.6%) | 65.3% | 0.81 | 0.59 |
| ATC B | 34/45 (75.6%) | 44.7% | 0.79 | 0.58 |
| ATC C | 34/45 (75.6%) | 47.2% | 0.77 | 0.54 |
| ATC D | 34/45 (75.6%) | 80.8% | 0.88 | 0.64 |
| Pilot A | 34/45 (75.6%) | 56.9% | 0.80 | 0.59 |
| Pilot B | 34/45 (75.6%) | 72.8% | 0.85 | 0.64 |
| Pilot C | 34/45 (75.6%) | 70.1% | 0.83 | 0.60 |
p₀ around 0.80 means roughly four candidates in five would be classified the same way on a parallel form; the remaining one is in the borderline region where SEM and the cut overlap. Decision accuracy (observed vs true classification) requires modelling the true-score distribution and is not reported here. Empirical Paper 1 pass rates and downstream outcome distributions for the 2026 operational cohort (N = 1,590) are reported on the validity page.
Conclusions and gaps in the current evidence
Held up against Green's (2013, ch. 5) four questions, the classical evidence on this page answers the first one clearly and contributes to the second; the third and fourth need evidence this page does not, by itself, supply. The Rasch page extends the picture with the complementary measurement criteria of Bond and Fox (2015).
What the CTT evidence supports
- The scores are reliable. α is at or above 0.80 on every form, and three independent reliability checks (α, ω, split-half) agree to within ±0.02 (Green, 2013; Bond & Fox, 2015).
- No form is materially harder or easier than the others within a cohort. Means, spreads and reliabilities line up — candidates are not advantaged or penalised by which form they happen to sit.
- Most items are at the right level of difficulty and do separate stronger candidates from weaker ones, within the screening-test target windows recommended by Green (2013, ch. 9).
Continuous-improvement programme
The items below sit on the published improvement roadmap. They describe planned methodological work, not operational deficiencies in the current Paper 1 pass/fail decision.
- The seven forms are not yet formally equated. Current comparability is descriptive; a common-item or concurrent calibration is planned before raw scores can be treated as exactly interchangeable (Bond & Fox, 2015, ch. 5).
- Precision is weakest at the top of the scale. Paper 1 is not the right tool for finely sorting high-proficiency candidates from one another — which is consistent with its screening role.
- Going beyond scoring needs evidence this dataset does not contain. Showing the score transfers across tasks and occasions, and predicts radiotelephony performance, requires the alternate-forms, criterion and task-relevance studies flagged elsewhere (Green, 2013, ch. 5).
Glossary of terms
- Cronbach's α
- Internal-consistency reliability on 0/1 item scores. Green (2013) cites ≥ 0.80 as acceptable for screening and ≥ 0.90 for high-stakes decisions.
- McDonald's ω
- Reliability estimate that relaxes the tau-equivalence assumption of α.
- SEM
- Standard error of measurement on the raw-score scale. A candidate's true score lies within roughly ±2·SEM of their observed score with 95% confidence.
- Conditional SEM (cSEM)
- SEM as a function of score (CTT) or ability θ (Rasch). Green (2013, p. 78) argues this is the relevant precision metric for decisions: what matters is precision at the cut, not on average.
- Infit / outfit MSQ
- Rasch fit statistics. Bond and Fox (2015, ch. 12) treat 0.5–1.5 as productive for measurement and 0.7–1.3 as the preferred tighter band.
- Person / item separation
- Rasch reliability analogues: how reliably the form distinguishes candidates along the ability continuum, and items along the difficulty continuum.
- Marginal reliability
- Population-level reliability of the Rasch ability estimates, reported alongside α for direct comparison.
- Targeting
- Alignment of person and item distributions on the logit scale (Bond & Fox, 2015, ch. 4). Good targeting means most candidates encounter items near their ability level.
- Decision consistency (p₀, κ)
- Probability that two parallel forms would classify a candidate the same way at the cut score, and Cohen's κ adjusting for chance agreement.
References
- Bond, T. G., & Fox, C. M. (2015). Applying the Rasch model: Fundamental measurement in the human sciences (3rd ed.). Routledge.
- Green, A. (2013). Exploring language assessment and testing. Routledge.