Skip to main content
EUROCONTROL
ELPAC

About this page. This validity documentation is maintained by the ELPAC team at EUROCONTROL as a living methodological record. It is not an ICAO endorsement and does not by itself constitute a regulatory determination. The ICAO Doc 9835 and Doc 10197 pages in this section are ELPAC's own self-assessments and have not been reviewed or endorsed by ICAO. National civil aviation authorities remain responsible for licensing decisions and for determining whether ELPAC results satisfy the ICAO language-proficiency requirements in their jurisdiction.

Last reviewed: 28 September 2026.

Paper 2 — rating results

This page describes how the digital Paper 2 has been rated under ELPAC Assessment Scheme v1.5. Between 2 February 2026 and 23 September 2026, 2,114 test takers from 39 test centres took Paper 2. Each test taker was rated by two assessors: an English Language Expert (ELE) and an Operational Expert (OPE). In total, 115 ELE assessors and 160 OPE assessors took part. These figures replace those published under the previous assessment scheme.

How Paper 2 is rated

Paper 2 is a live speaking test with two parts: Task 1 (role play and debriefing) and Tasks 2 and 3 (picture description and interview). Two accredited assessors take part. The ELE rates the six criteria (Pronunciation, Structure, Vocabulary, Fluency, Comprehension, Interaction) at level 4, 5 or 6, and decides whether the test taker reaches level 4. The OPE, an operational or former air traffic controller or pilot, judges at the level 4 threshold whether the test taker's Pronunciation, Fluency (two sub-criteria), Comprehension and Interaction are sufficient for safe and efficient work, and decides whether the test taker's English exceeds the demands of the job (the level 6 question).

Each assessor rates their assigned criteria independently, using the ELPAC assessment criteria, which derive from the ICAO rating scale and the holistic descriptors. The ICAO level awarded is determined by the lowest rating across the six criteria from both assessors; there is no compensatory marking. Where the two assessors disagree on the outcome, a third assessor decides.

Who took Paper 2

What this indicates

ELPAC test takers are not one homogeneous group. They differ in operational role and in proficiency level, and the mix differs strongly between test centres. Figures for all test takers together therefore describe this particular mix, not a fixed property of the test.

Operational roleTest takersShare
Pilot83339.4 %
ATC TWR50724.0 %
ATC ENR49023.2 %
ATC APP25712.2 %
ATC (other)271.3 %
Paper 2 resultTest takersShare
Level 31255.9 %
Level 489942.5 %
Level 571333.7 %
Level 636417.2 %
Referred to third assessor130.6 %

Differences between test centre populations

What this indicates

Test centres differ markedly in the level profile of their test takers. The largest group of centres accounts for 46 % of all test takers, so it has a large effect on figures for all test takers together. Results are therefore also shown by centre profile. Centres are not named.

Each centre with at least 20 test takers is assigned to a profile according to its most frequent Paper 2 result. Centres with fewer than 20 test takers are grouped together. The profiles describe populations, not the quality of a centre.

Centre profileCentresNPilotsL3L4L5L6
Mostly level 642540 %0 %7 %23 %70 %
Mostly level 51273230 %1 %24 %56 %19 %
Mostly level 4 or below396257 %11 %67 %20 %2 %
Smaller centres (fewer than 20 test takers each)2016643 %7 %37 %31 %17 %

OPE and ELE decisions at the level 4 threshold

What this indicates

The two assessors reached the same level 4 decision for every test taker (51 test takers below level 4). This result is reported for completeness only. It is not yet used as evidence of inter-assessor agreement, because it is still being confirmed that the two level 4 decisions are recorded fully independently of each other in the digital platform.

Reaches level 4OPE yesOPE no
ELE yes2,0630
ELE no051

OPE and ELE decisions at the level 6 threshold

What this indicates

The two assessors answer different questions at this point. The OPE decides whether the test taker's English exceeds the demands of the job. The ELE decides whether all six criteria reach level 6. The OPE says yes much more often (45 % of test takers received an OPE yes without an ELE level 6); the reverse occurred once. Because the final level is the lowest rating from both assessors, level 6 is awarded only when both agree. The pattern shows that the linguistic level 6 standard is stricter than operational sufficiency, which is the intended role of the ELE.

Level 6OPE yesOPE no
ELE yes3681
ELE no961784

Overall agreement is 54.5 % (Cohen's κ = 0.22). A low κ is expected here: κ measures whether two judges apply the same standard, and at this threshold they deliberately do not. The table below shows how agreement depends on the centre profile. Where most test takers are at level 6, or where almost none are, κ is low even when the assessors agree in most cases, because κ depends on how evenly the outcomes are spread.

Centre profileNAgreementκThird assessor
Mostly level 625473.2 %0.160.0 %
Mostly level 573235.8 %0.090.1 %
Mostly level 4 or below96263.4 %0.050.0 %
Smaller centres (fewer than 20 test takers each)16656.6 %0.257.2 %

OPE criteria at the level 4 threshold

What this indicates

When the OPE finds a test taker below level 4, this is most often on Comprehension (3.5 % of ratings). The two fluency sub-criteria, now recorded separately, give the same result in 99.1 % of sessions. OPE fluency is therefore included in the statistics on this page.

OPE criterionRatingsBelow level 4Rate
Pronunciation2,110140.7 %
Fluency 12,109251.2 %
Fluency 22,109291.4 %
Comprehension2,110753.5 %
Interaction2,110562.6 %

ELE criteria and the Many-Facet Rasch model

What this indicates

Structure is the most demanding criterion and Comprehension and Interaction the least demanding. The criteria are clearly distinguished from each other (separation reliability 0.98). Differences in severity between individual ELE assessors cannot be estimated reliably from this data set, for the reason explained below.

ELE criterionL4L5L6MeanDifficulty (logits)
Pronunciation36 %41 %23 %4.87-0.22 ± 0.09
Fluency40 %37 %22 %4.820.30 ± 0.09
Comprehension33 %43 %24 %4.91-0.71 ± 0.08
Interaction34 %42 %24 %4.90-0.63 ± 0.08
Structure44 %36 %21 %4.770.96 ± 0.09
Vocabulary40 %38 %22 %4.820.30 ± 0.09

Criteria from least to most demanding: Comprehension, Interaction, Pronunciation, Vocabulary, Fluency, Structure. The model uses 8,664 ELE ratings from 26 assessors with at least 20 test takers each. The two step thresholds (level 4 to 5, level 5 to 6) are -3.15 and 3.15 logits, so the three levels are clearly ordered.

Why assessor severity is not reported. In this reporting period, every ELE assessor rated test takers from one test centre only. An assessor who gives lower ratings may therefore be stricter, or may be rating a population with a lower proficiency level; the data cannot tell these two explanations apart. A severity ranking of assessors is therefore not published for this period. To make this separation possible, a sample of recordings will be rated by assessors from other centres. Assessor quality is currently also monitored through accreditation, refresher training and the third-assessor procedure.

Third-assessor referrals

What this indicates

13 of 2,114 test takers (0.61 %) were referred to a third assessor. In all other cases the two assessors' ratings led to a result without referral.

This is a referral rate. It is not an inter-rater agreement coefficient.

Limitations

  • The results reflect the test takers of this period. Because centre populations differ, figures for all test takers together will change when the mix of centres changes.
  • The assessment scheme changed with the introduction of the digital Paper 2. Figures from the previous scheme are not directly comparable and are no longer shown.
  • ICAO levels are ordinal. Logit values are on the Rasch scale and are relative to the mean of all criteria.
  • This page does not report a study against operational radiotelephony performance. That work is listed under the validity argument as planned work.

Methods

Data. One record per test taker session, from the digital platform. Demonstration records were removed. All names of test takers, assessors and test centres were removed before any figure was produced.

Decisions. Level 4 decision: recommended level of 4 or higher. Level 6 decision: recommended level of 6. Agreement is the share of test takers for whom both assessors made the same decision; Cohen's κ uses the 2 × 2 table.

Centre profiles. Centres with at least 20 test takers are grouped by their most frequent Paper 2 result; smaller centres form one group.

Many-Facet Rasch model. Rating-scale model with facets test taker, ELE assessor and criterion, on the categories level 4, 5 and 6: P(Xnij = k) ∝ exp(k·(θn − βi − δj) − Σm≤k τm). Joint maximum likelihood with weak priors; sum-to-zero identification. Assessors with fewer than 20 test takers are excluded. Test takers are nested in centres, so centre cannot be estimated as a separate facet; the design check reports how many assessors rated in more than one centre.

References

  • Linacre, J. M. (1989/2014). Many-Facet Rasch Measurement. MESA Press / Winsteps.com.
  • Eckes, T. (2015). Introduction to Many-Facet Rasch Measurement (2nd ed.). Peter Lang.
  • Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543–549.
  • ICAO Doc 9835 — Manual on the Implementation of ICAO Language Proficiency Requirements (2nd ed., 2010).

See also the validity argument for the assessment use argument these statistics support, and the Paper 1 Rasch analysis for the listening test evidence.