Skip to main content
EUROCONTROL
ELPAC

About this page. This validity documentation is maintained by the ELPAC team at EUROCONTROL as a living methodological record. It is not an ICAO endorsement and does not by itself constitute a regulatory determination. The ICAO Doc 9835 and Doc 10197 pages in this section are ELPAC's own self-assessments and have not been reviewed or endorsed by ICAO. National civil aviation authorities remain responsible for licensing decisions and for determining whether ELPAC results satisfy the ICAO language-proficiency requirements in their jurisdiction.

Last reviewed: 7 September 2026.

Paper 2 — assessor agreement and Rasch evidence

Reporting window 1 July 2025 – 18 June 2026. Sample N = 894 test taker sessions, each rated live by an English Language Expert (ELE) and an Operational Expert (OPE). The active ELE panel in window comprises 55 assessors with at least 20 ratings; OPEs are drawn from a smaller operational pool. This page reports paired-rater agreement at the descriptor and holistic levels, the ELE–OPE approval gap at the level 4 operational floor, and a Many-Facet Rasch model of assessor severity, descriptor difficulty and the rating-scale step thresholds.

How Paper 2 is rated

Paper 2 is a live speaking test that applies the ELPAC assessment criteria, which derive from the ICAO rating scale and the holistic descriptors, by means of two accredited assessors. The English Language Expert (ELE) awards an ICAO rating on each of the six descriptors (pronunciation, fluency, comprehension, interaction, structure, vocabulary) and recommends a Paper 2 level. The Operational Expert (OPE), an operational or former air traffic controller or pilot, independently applies a Pass / Fail judgement at the level 4 threshold on the three operational descriptors (pronunciation, comprehension, interaction) and confirms or contests the ELE's recommended level. The final Paper 2 score is then determined as the lowest of the six criteria, with no compensatory marking. Where the two assessors disagree on the outcome, a third assessor decides.

ELE ↔ OPE descriptor agreement at the Level 4 threshold

What this indicates

At the level 4 threshold the language expert and the operational expert agree on every descriptor for every paired session in the reporting window. The cohort that reaches Paper 2 is at or above level 4 on the operationally-rated descriptors.

For each of the three operationally-rated descriptors, the ELE's full ICAO rating is collapsed to "at or above level 4" vs "below level 4" and compared with the OPE's Pass / Fail. In this window every OPE descriptor judgement on every paired session is a Pass and every ELE rating is at or above level 4, so the 2 × 2 table has mass in one cell only. Cohen's κ requires variance in both margins and is therefore undefined ("—"); the dichotomous agreement rate is trivially 100 %. This is a property of the cohort that passed Paper 1 and reached Paper 2 in window, not of the rating procedure. The operationally meaningful signal at the level 4 floor is the separate operational-approval judgement reported two sections below, where the ELE and OPE do disagree.

DescriptorN pairsAgreementCohen κMean ELE level when OPE Pass
pronunciation876100.0 %5.16
comprehension875100.0 %5.19
interaction876100.0 %5.19

Holistic decision — confirmation of the recommended level

What this indicates

The operational expert confirms the language expert's recommended level in the substantial majority of paired sessions. Where the two assessors disagree on the outcome, the prescribed dispute procedure (third assessor) takes over; the OPE does not silently override the ELE.

The OPE's holistic role is to confirm or contest the ELE's recommended Paper 2 level. Overall confirmation in window is 98.3 % across N = 894 paired sessions. Confirmation is unanimous at levels 3 and 4 and at level 6 (≥ 99 %); level 5 is the noisiest tier at , where 0 of recommendations were not confirmed on first pass. The "Mean Calculated Paper 2" column equals the recommended level in every row: the recommended-vs-calculated weighted κ is 1.00 and exact match 100 %. In other words, every disconfirmation in this window was reconciled back to the original recommended level before the score was issued, by the prescribed dispute procedure, not by the OPE silently overriding the ELE.

ELE-recommended levelNOPE confirmation rateMean Calculated Paper 2
Level 311100.0 %3.00
Level 4219100.0 %4.00
Level 542496.9 %5.00
Level 624099.2 %6.00

Third assessor: 0 of 894 sessions (0.00 %) were escalated to a third assessor in the reporting window. A third assessor is invoked when the ELE and OPE disagree on the outcome at the agreed level. The figure reported above is an escalation rate, not an inter-rater agreement coefficient.

Operational approval at the Level 4 floor

What this indicates

When the language expert and the operational expert disagree, the disagreement is always in the same direction: operationally stricter. That one-sided gap is exactly the signal the third-assessor mechanism is designed to catch.

Alongside the descriptor and holistic-level judgements, each assessor records a separate approval at the operational level 4 floor. The ELE approves the language sample as meeting ICAO level 4 across the descriptors for which the ELE is responsible; the OPE independently approves it as operationally adequate for radiotelephony work. These are two different judgements about two different objects, language-sample adequacy and operational adequacy, so they are not expected to coincide, and the gap between them is informative rather than error.

N paired sessions
878
ELE-approved rate
100.0 %
OPE-approved rate
78.7 %
QuantityValue
Agreement (both approved or both not approved)78.7 %
Cohen κ0.00
ELE-approved & OPE-not-approved187
ELE-not-approved & OPE-approved0

The ELE-approved margin is effectively constant in this window (every reported ELE judgement is an approval), so Cohen κ is 0.00 by construction: there is no variance on the ELE side for κ to measure. The operationally meaningful number is the bottom-right of the table: 187 of 878 sessions (21.3 %) where the ELE judged the language sample adequate at level 4 but the OPE did not approve it operationally. The reverse asymmetry (OPE approves, ELE does not) does not occur. This one-sided gap is the operational channel through which the third-assessor mechanism activates when it activates, and it is also the input to OPE calibration and training.

Many-Facet Rasch (MFRM) analysis

What this indicates

Assessor severity, descriptor difficulty and test taker ability are placed on a single, common scale. Residual differences between assessors are quantified and monitored; no individual assessor's tendency is allowed to drive a test taker's score on its own.

A Many-Facet Rasch Model fits test taker ability, ELE assessor severity, descriptor difficulty and a shared set of rating-scale step thresholds on a single logit scale. This lets an individual assessor's tendency to award higher or lower ratings be reported on the same metric as test taker ability, even when different assessors rate different test takers. The model is fitted on 5,256 ELE descriptor ratings from 55 active ELE assessors (≥ 20 observations in window). The category set in this window is {4, 5, 6}: Fail and levels 1–3 do not occur on the ELE side over this period, so the model is fitted on three collapsed categories. Because the active panel in this window is larger and more dispersed than in earlier reporting periods, the severity range and the count of fit-flagged assessors are correspondingly larger; both are discussed in the captions below.

Descriptor difficulty (logits)

Higher logit = harder to score at the upper category on that descriptor, holding test taker and assessor constant. Standard errors are diagonal observed-information estimates.

DescriptorDifficultySEInfit MnSqOutfit MnSq
pronunciation-0.140.120.600.27
fluency0.080.110.520.22
comprehension-0.460.110.530.24
interaction-0.510.110.490.23
structure0.790.110.500.28
vocabulary0.250.110.600.27

Structure is the hardest descriptor to attain a higher rating on (0.79 logits); interaction (-0.51) and comprehension (-0.46) are the easiest. The narrow standard errors reflect the large number of ratings per descriptor in this window (5,256 in total).

Assessor severity — anonymised distribution

Each bar is one ELE assessor with ≥ 20 ratings in window, sorted from most lenient (negative) to most severe (positive). Assessor identities are anonymised; no assessor is named on this page. One logit ≈ the spacing between adjacent ICAO category steps in this dataset.

Severity range
-6.67 to 5.18 logits
N = 55 ELE assessors
Separation reliability
0.79
Separation index 1.91 · mean SE 0.89

A separation reliability of 0.79 indicates that the panel of ELE assessors is distinguishable but with some overlap. Assessor-level severity estimates are corrections to be read against the panel mean; they are not, on their own, certifications of any individual assessor.

How to read severity spread. A non-zero range is an expected property of the measurement model and does not by itself indicate a fairness problem: every test taker's score combines an ELE and an OPE rating and escalates to a third assessor where they disagree, so individual assessor severity is corrected for in the operational decision.

Rating-scale step thresholds (τ)

Thresholds between adjacent categories on the collapsed {4, 5, 6} scale. Monotonically ordered thresholds indicate a well-functioning rating scale.

ThresholdValue (logits)
τ1 (45)-3.46
τ2 (56)3.46

Model fit

Test taker separation reliability
0.68
Descriptor separation reliability
0.94
Assessors flagged for fit
51 of 55
|MnSq − 1| > 0.3

The model converged in 268 iterations. Fit flags identify assessors whose rating pattern is either too predictable (MnSq well below 1, often a halo or compressed-range pattern in which the assessor gives the same level across several descriptors within a session) or too erratic (MnSq well above 1) relative to the model. In this window the count is inflated by structural features of the data rather than by erratic rating: the collapsed three-category {4, 5, 6} scale plus within-session halo make many low-N assessors' patterns near-deterministic for the model, which produces MnSq values close to zero. The flag is therefore an input to assessor support and calibration review, not a verdict on any individual assessor, and should be read together with the assessor's observation count and standard error.

Data-source note — OPE fluency

OPE fluency is excluded from the descriptor-level statistics on this page. The OPE assesses fluency by means of two underlying assessment criteria, but the current capture format records only a single OPE-fluency value, which does not carry the intended signal. The ELPAC programme is addressing this data-source issue, and OPE fluency will be reintroduced into the published agreement statistics once the capture format separates the two sub-criteria. ELE fluency is recorded on the full ICAO scale and is included in every analysis on this page, including the Many-Facet Rasch model.

Scope and limitations

  • Reporting window 1 July 2025 – 18 June 2026, sample N = 894 test taker sessions. Active ELE panel in the reporting window: 55 assessors with ≥ 20 ratings.
  • OPE fluency is excluded from all paired-rater statistics on this page (see data-source note above). ELE fluency is included in every analysis, including the MFRM.
  • Operational approval is reported alongside descriptor and holistic agreement and is the channel in which the ELE and OPE actually disagree in this window. ELE approval and OPE approval are different judgements about different objects; the gap between them is informative, not error, and κ is uninformative because the ELE margin is constant.
  • MFRM categories are collapsed to {4, 5, 6} because Fail and levels 1–3 do not occur on the ELE side in this window. Estimates are interpretable on the operational range only. The compressed scale plus within-session halo cause fit-flag inflation for low-N assessors, so flags in this window are sensitivity indicators, not quality verdicts.
  • No assessor, test taker, test centre, organisation or State identifier appears in any table or chart on this page. Assessor severities are reported as an anonymised distribution.
  • This page reports assessor agreement and a Rasch model. It does not report cross-form equating or a criterion-validity study against operational radiotelephony performance — those are listed under the validity-argument tab as planned work.

Methods

Descriptor agreement. For each operationally-rated descriptor, the ELE rating is collapsed to "≥ level 4" vs "< level 4" and compared with the OPE Pass / Fail. The agreement rate is the proportion of paired ratings on which the two assessors agree on Pass / Fail at the L4 threshold. Cohen's κ uses the dichotomous 2×2 table; it is undefined when one margin is constant.

Holistic confirmation. "OPE confirmation rate" is the proportion of sessions for which "Paper 2 confirmed by OPE" is recorded as Yes. "Mean Calculated Paper 2" is the mean of the final score field; equality with the recommended level in this window reflects that every disconfirmation was reconciled back to the originally recommended level under the prescribed dispute procedure, not that no disconfirmation occurred.

Operational approval. Computed from the session-level fields "ELE approved" and "OPE approved" (both Pass / Fail at the level 4 operational floor). The 2 × 2 table counts paired sessions in each cell; agreement is the fraction of paired sessions on which both assessors record the same verdict; Cohen κ uses that 2 × 2 table and is uninformative when one margin is constant.

Many-Facet Rasch model. Rating-scale model: P(Xnij = k) ∝ exp(k·(θn − βi − δj) − Σm≤k τm) where θn is test taker ability, βi ELE assessor severity, δj descriptor difficulty and τm shared step thresholds. Identification: Σβ = Σδ = Στ = 0. Estimation: full-information joint maximum likelihood via L-BFGS-B with weak Gaussian priors (σ = 4 on θ, σ = 2 on β, δ, τ) to stabilise low-N assessors. Standard errors are diagonal observed-information estimates.

Public anonymisation. Assessor names are removed before any data reaches the website. The public JSON consumed by this page contains only opaque labels (Rater A, B, …) and distributional statistics; no assessor-identifying covariate (test centre, language background, seniority, country) is published. The mapping from real assessor to opaque label is regenerated for each reporting window and held only by the ELPAC team, so labels are not stable across windows. Low-N assessors are pooled. Assessors who would prefer their row to be suppressed in a future window can contact the ELPAC team via the contact page.

References

  • Linacre, J. M. (1989/2014). Many-Facet Rasch Measurement. MESA Press / Winsteps.com.
  • Bond, T. G., & Fox, C. M. (2015). Applying the Rasch Model: Fundamental Measurement in the Human Sciences (3rd ed.). Routledge.
  • Knoch, U. (2009). Diagnostic assessment of writing: A comparison of two rating scales. Language Testing, 26(2), 275–304.
  • ICAO Doc 9835 — Manual on the Implementation of ICAO Language Proficiency Requirements (2nd ed., 2010), §4.6 and Attachment A.

See also the validity argument for the assessment use argument these statistics support, and the Paper 1 Rasch analysis for the corresponding listening-test evidence.