About this page. This validity documentation is maintained by the ELPAC team at EUROCONTROL as a living methodological record. It is not an ICAO endorsement and does not by itself constitute a regulatory determination. National civil aviation authorities remain responsible for licensing decisions and for determining whether ELPAC results satisfy the ICAO language-proficiency requirements in their jurisdiction.

Last reviewed: 18 June 2026.

Paper 2 — examiner agreement and Rasch evidence

Reporting window 1 July 2025 – 18 June 2026. Sample N = 894 candidate sessions, each rated live by an English Language Expert (ELE) and an Operational Expert (OPE). The active ELE panel in window comprises 55 examiners with at least 20 ratings; OPEs are drawn from a smaller operational pool. This page reports paired-rater agreement at the descriptor and holistic levels, the ELE–OPE approval gap at the Level 4 operational floor, and a Many-Facet Rasch model of examiner severity, descriptor difficulty and the rating-scale step thresholds.

How Paper 2 is rated

Paper 2 is a face-to-face speaking assessment that applies the ICAO holistic descriptors and ICAO rating scale criteria by means of two accredited examiners. The English Language Expert (ELE) awards an ICAO rating on each of the six descriptors (pronunciation, fluency, comprehension, interaction, structure, vocabulary) and recommends a Paper 2 level. The Operational Expert (OPE)— a serving or former air traffic controller or pilot — independently applies a Pass / Fail judgement at the Level 4 threshold on the three operational descriptors (pronunciation, comprehension, interaction) and confirms or contests the ELE's recommended level. The final Paper 2 score is then determined as the lowest of the six sub-ratings, in accordance with the ICAO rating scale criteria. A third examiner is invoked when the ELE and OPE cannot reconcile.

ELE ↔ OPE descriptor agreement at the Level 4 threshold

What this indicates

At the Level 4 threshold the language expert and the operational expert agree on every descriptor for every paired session in the reporting window. The cohort that reaches Paper 2 is at or above Level 4 on the operationally-rated descriptors.

For each of the three operationally-rated descriptors, the ELE's full ICAO rating is collapsed to "at or above Level 4" vs "below Level 4" and compared with the OPE's Pass / Fail. In this window every OPE descriptor judgement on every paired session is a Pass and every ELE rating is at or above Level 4, so the 2 × 2 table has mass in one cell only. Cohen's κ requires variance in both margins and is therefore undefined ("—"); the dichotomous agreement rate is trivially 100 %. This is a property of the cohort that passed Paper 1 and reached Paper 2 in window, not of the rating procedure. The operationally meaningful signal at the Level 4 floor is the separate operational-approval judgement reported two sections below — there, the ELE and OPE do disagree.

DescriptorN pairsAgreementCohen κMean ELE level when OPE Pass
pronunciation876100.0 %5.16
comprehension875100.0 %5.19
interaction876100.0 %5.19

Holistic decision — confirmation of the recommended level

What this indicates

The operational expert confirms the language expert's recommended level in the substantial majority of paired sessions. Where they cannot reconcile, the prescribed dispute procedure (third examiner) takes over — the OPE does not silently override the ELE.

The OPE's holistic role is to confirm or contest the ELE's recommended Paper 2 level. Overall confirmation in window is 98.3 % across N = 894 paired sessions. Confirmation is unanimous at Levels 3 and 4 and at Level 6 (≥ 99 %); Level 5 is the noisiest tier at 96.9 %, where 13 of 424 recommendations were not confirmed on first pass. The "Mean Calculated Paper 2" column equals the recommended level in every row: the recommended-vs-calculated weighted κ is 1.00 and exact match 100 %. In other words, every disconfirmation in this window was reconciled back to the original recommended level before the score was issued — by the prescribed dispute procedure, not by the OPE silently overriding the ELE.

ELE-recommended levelNOPE confirmation rateMean Calculated Paper 2
Level 311100.0 %3.00
Level 4219100.0 %4.00
Level 542496.9 %5.00
Level 624099.2 %6.00

Third assessor: 0 of 894 sessions (0.00 %) were escalated to a third examiner in the reporting window. A third examiner is invoked when the ELE and OPE cannot reconcile at the agreed level. The figure reported above is an escalation rate, not an inter-rater agreement coefficient.

Operational approval at the Level 4 floor

What this indicates

When the language expert and the operational expert disagree, the disagreement is always in the same direction: operationally stricter. That one-sided gap is exactly the signal the third-examiner mechanism is designed to catch.

Alongside the descriptor and holistic-level judgements, each examiner records a separate approval at the operational Level 4 floor. The ELE approves the language sample as meeting ICAO Level 4 across the descriptors for which the ELE is responsible; the OPE independently approves it as operationally adequate for radiotelephony work. These are two different judgements about two different objects — language-sample adequacy and operational adequacy — so they are not expected to coincide, and the gap between them is informative rather than error.

N paired sessions
878
ELE-approved rate
100.0 %
OPE-approved rate
78.7 %
QuantityValue
Agreement (both approved or both not approved)78.7 %
Cohen κ0.00
ELE-approved & OPE-not-approved187
ELE-not-approved & OPE-approved0

The ELE-approved margin is effectively constant in this window (every reported ELE judgement is an approval), so Cohen κ is 0.00 by construction: there is no variance on the ELE side for κ to measure. The operationally meaningful number is the bottom-right of the table — 187 of 878 sessions (21.3 %) where the ELE judged the language sample adequate at Level 4 but the OPE did not approve it operationally. The reverse asymmetry (OPE approves, ELE does not) does not occur. This one-sided gap is the operational channel through which the third-examiner mechanism activates when it activates, and it is also the input to OPE calibration and training.

Many-Facet Rasch (MFRM) analysis

What this indicates

Examiner severity, descriptor difficulty and candidate ability are placed on a single, common scale. Residual differences between examiners are quantified and monitored; no individual examiner's tendency is allowed to drive a candidate's score on its own.

A Many-Facet Rasch Model fits candidate ability, ELE examiner severity, descriptor difficulty and a shared set of rating-scale step thresholds on a single logit scale. This lets an individual examiner's tendency to award higher or lower ratings be reported on the same metric as candidate ability, even when different examiners rate different candidates. The model is fitted on 5,256 ELE descriptor ratings from 55 active ELE examiners (≥ 20 observations in window). The category set in this window is {4, 5, 6} — Fail and Levels 1–3 do not occur on the ELE side over this period, so the model is fitted on three collapsed categories. Because the active panel in this window is larger and more dispersed than in earlier reporting periods, the severity range and the count of fit-flagged examiners are correspondingly larger; both are discussed in the captions below.

Descriptor difficulty (logits)

Higher logit = harder to score at the upper category on that descriptor, holding candidate and examiner constant. Standard errors are diagonal observed-information estimates.

DescriptorDifficultySEInfit MnSqOutfit MnSq
pronunciation-0.140.120.600.27
fluency0.080.110.520.22
comprehension-0.460.110.530.24
interaction-0.510.110.490.23
structure0.790.110.500.28
vocabulary0.250.110.600.27

Structure is the hardest descriptor to attain a higher rating on (0.79 logits); interaction (-0.51) and comprehension (-0.46) are the easiest. The narrow standard errors reflect the large number of ratings per descriptor in this window (5,256 in total).

Examiner severity — anonymised distribution

Each bar is one ELE examiner with ≥ 20 ratings in window, sorted from most lenient (negative) to most severe (positive). Examiner identities are anonymised; no examiner is named on this page. One logit ≈ the spacing between adjacent ICAO category steps in this dataset.

Severity range
-6.67 to 5.18 logits
N = 55 ELE examiners
Separation reliability
0.79
Separation index 1.91 · mean SE 0.89

A separation reliability of 0.79 indicates that the panel of ELE examiners is distinguishable but with some overlap. Examiner-level severity estimates are corrections to be read against the panel mean — they are not, on their own, certifications of any individual examiner.

How to read severity spread. A non-zero range is an expected property of the measurement model and does not by itself indicate a fairness problem: every candidate's score combines an ELE and an OPE rating and escalates to a third examiner where they disagree, so individual examiner severity is corrected for in the operational decision.

Rating-scale step thresholds (τ)

Thresholds between adjacent categories on the collapsed {4, 5, 6} scale. Monotonically ordered thresholds indicate a well-functioning rating scale.

ThresholdValue (logits)
τ1 (45)-3.46
τ2 (56)3.46

Model fit

Candidate separation reliability
0.68
Descriptor separation reliability
0.94
Examiners flagged for fit
51 of 55
|MnSq − 1| > 0.3

The model converged in 268 iterations. Fit flags identify examiners whose rating pattern is either too predictable (MnSq well below 1, often a halo or compressed-range pattern in which the examiner gives the same level across several descriptors within a session) or too erratic (MnSq well above 1) relative to the model. In this window the count is inflated by structural features of the data rather than by erratic rating: the collapsed three-category {4, 5, 6} scale plus within-session halo make many low-N examiners' patterns near-deterministic for the model, which produces MnSq values close to zero. The flag is therefore an input to examiner support and calibration review, not a verdict on any individual examiner, and should be read together with the examiner's observation count and standard error.

Data-source note — OPE fluency

OPE fluency is excluded from the descriptor-level statistics on this page. The OPE assesses fluency by means of two underlying assessment criteria, but the current capture format records only a single OPE-fluency value, which does not carry the intended signal. The ELPAC programme is addressing this data-source issue, and OPE fluency will be reintroduced into the published agreement statistics once the capture format separates the two sub-criteria. ELE fluency is recorded on the full ICAO scale and is included in every analysis on this page, including the Many-Facet Rasch model.

Scope and limitations

  • Reporting window 1 July 2025 – 18 June 2026, sample N = 894 candidate sessions. Active ELE panel in the reporting window: 55 examiners with ≥ 20 ratings.
  • OPE fluency is excluded from all paired-rater statistics on this page (see data-source note above). ELE fluency is included in every analysis, including the MFRM.
  • Operational approval is reported alongside descriptor and holistic agreement and is the channel in which the ELE and OPE actually disagree in this window. ELE approval and OPE approval are different judgements about different objects; the gap between them is informative, not error, and κ is uninformative because the ELE margin is constant.
  • MFRM categories are collapsed to {4, 5, 6} because Fail and Levels 1–3 do not occur on the ELE side in this window. Estimates are interpretable on the operational range only. The compressed scale plus within-session halo cause fit-flag inflation for low-N examiners — flags in this window are sensitivity indicators, not quality verdicts.
  • No examiner, candidate, test centre, organisation or State identifier appears in any table or chart on this page. Examiner severities are reported as an anonymised distribution.
  • This page reports examiner agreement and a Rasch model. It does not report cross-form equating or a criterion-validity study against operational radiotelephony performance — those are listed under the validity-argument tab as planned work.

Methods

Descriptor agreement. For each operationally-rated descriptor, the ELE rating is collapsed to "≥ Level 4" vs "< Level 4" and compared with the OPE Pass / Fail. The agreement rate is the proportion of paired ratings on which the two examiners agree on Pass / Fail at the L4 threshold. Cohen's κ uses the dichotomous 2×2 table; it is undefined when one margin is constant.

Holistic confirmation. "OPE confirmation rate" is the proportion of sessions for which "Paper 2 confirmed by OPE" is recorded as Yes. "Mean Calculated Paper 2" is the mean of the final score field; equality with the recommended level in this window reflects that every disconfirmation was reconciled back to the originally recommended level under the prescribed dispute procedure, not that no disconfirmation occurred.

Operational approval. Computed from the session-level fields "ELE approved" and "OPE approved" (both Pass / Fail at the Level 4 operational floor). The 2 × 2 table counts paired sessions in each cell; agreement is the fraction of paired sessions on which both examiners record the same verdict; Cohen κ uses that 2 × 2 table and is uninformative when one margin is constant.

Many-Facet Rasch model. Rating-scale model: P(Xnij = k) ∝ exp(k·(θn − βi − δj) − Σm≤k τm) where θn is candidate ability, βi ELE examiner severity, δj descriptor difficulty and τm shared step thresholds. Identification: Σβ = Σδ = Στ = 0. Estimation: full-information joint maximum likelihood via L-BFGS-B with weak Gaussian priors (σ = 4 on θ, σ = 2 on β, δ, τ) to stabilise low-N examiners. Standard errors are diagonal observed-information estimates.

Public anonymisation. Examiner names are removed before any data reaches the website. The public JSON consumed by this page contains only opaque labels (Rater A, B, …) and distributional statistics; no examiner-identifying covariate (test centre, language background, seniority, country) is published. The mapping from real examiner to opaque label is regenerated for each reporting window and held only by the ELPAC team, so labels are not stable across windows. Low-N examiners are pooled. Examiners who would prefer their row to be suppressed in a future window can contact the ELPAC team via the contact page.

References

  • Linacre, J. M. (1989/2014). Many-Facet Rasch Measurement. MESA Press / Winsteps.com.
  • Bond, T. G., & Fox, C. M. (2015). Applying the Rasch Model: Fundamental Measurement in the Human Sciences (3rd ed.). Routledge.
  • Knoch, U. (2009). Diagnostic assessment of writing: A comparison of two rating scales. Language Testing, 26(2), 275–304.
  • ICAO Doc 9835 — Manual on the Implementation of ICAO Language Proficiency Requirements (2nd ed., 2010), §4.6 and Attachment A.

See also the validity argument for the interpretation/use framework these statistics support, and the Paper 1 Rasch analysis for the corresponding listening-test evidence.