Statement on the fitness of ELPAC for licensing use
1. Purpose and scope. The English Language Proficiency for Aeronautical Communication (ELPAC) test is the language proficiency assessment administered by EUROCONTROL for the purpose of determining whether air traffic controllers and flight crew members meet the language proficiency requirements set out in ICAO Annex 1 (Personnel Licensing), §1.2.9 and Appendix 1, as elaborated in ICAO Doc 9835 (Manual on the Implementation of ICAO Language Proficiency Requirements). The assessment is administered for the purpose of supporting licensing decisions; the licensing decision itself rests with the competent national civil aviation authority.
2. Construct and design alignment. ELPAC comprises two components, each aligned with the ICAO construct. Paper 1, a computer-delivered listening comprehension component, operates as a screening instrument targeted at the ICAO Operational Level 4 threshold. Paper 2 is a face-to-face speaking assessment rated live by two accredited examiners — an English Language Expert (ELE) and an Operational Expert (OPE), a serving or former air traffic controller or pilot. Each examiner independently applies the ICAO holistic descriptors and ICAO rating scale criteria. The overall ICAO level awarded is the lowest of the six sub-ratings. When the two examiners cannot reconcile, the case is resolved through a documented dispute-resolution mechanism (a third examiner).
3. Evidence supporting use. On the operational dataset published on this site, Paper 1 forms exhibit reliability and item-difficulty targeting consistent with their intended screening function across the seven operational forms. Paper 2 inter-examiner outcomes demonstrate agreement in the substantial majority of paired sessions, with residual differences in examiner severity quantified on a common scale through a Many-Facet Rasch model and with a low rate of escalation to a third examiner. The quantitative evidence is integrated into a qualitative validity argument framed in the terms of Kane (2013) and Chapelle, Enright & Jamieson (2008), and is documented in full on the linked operational-validity pages.
4. Limitations and ongoing work. The published evidence is intentionally incomplete, and one limitation is structural rather than methodological. ELPAC is administered independently of the candidate's employer and operational career: the test programme has no routine access to a candidate's subsequent radiotelephony performance on the line, and individual test takers cannot be systematically followed into live ATC or flight operations. A conventional predictive criterion study against operational radiotelephony performance is therefore difficult to mount and cannot, on its own, be relied upon to close the extrapolation inference. The programme accordingly pursues a range of complementary feedback mechanisms — including structured stakeholder and user feedback from competent authorities and approved training organisations, periodic expert review of test materials against the evolving target language use domain, and, where lawfully available and de-identified, signals from communication- related occurrence reporting and downstream training outcomes — together with formal IRT equating across Paper 1 forms, the reintroduction of OPE fluency once the data-capture format is revised, and complementary κ-style examiner-agreement coefficients at rating level. These activities are presented as ongoing validation work; no claim is made that the current evidence base directly demonstrates operational performance in service.
5. Conclusion. Subject to the limitations set out in paragraph 4, the evidence presented supports the use of ELPAC by competent authorities as one element of the technical basis for determining compliance with the ICAO language proficiency requirements for the licensing of air traffic controllers and flight crew members. No claim is made that ELPAC, on its own, establishes a candidate's operational radiotelephony performance in service. The determination of compliance, and the weight to be attached to ELPAC alongside other evidence available to the authority, remains with the national authority in accordance with the applicable national regulations.
ELPAC is fit for its intended purpose
On the operational dataset published on this site (N = 1,590 candidates over the reporting period 2026-01-05 to 2026-08-31), the Paper 1 listening comprehension screen reaches a Cronbach α reliability of 0.82–0.90 across the seven operational forms, the overall Paper 1 pass rate is 88.9 %, and the Paper 2 two-rater protocol (N = 894 sessions) escalates to a third examiner in only 1.57 % of cases. Taken together with the qualitative validity argument, this evidence is intended to support national civil aviation authorities in determining compliance with the ICAO language-proficiency requirements for air traffic controllers and pilots.
How the design and the evidence fit together
The ELPAC programme builds its case for use in two layers. The qualitative validity argument (Kane 2013; Chapelle, Enright & Jamieson 2008) chains six warranted inferences from the test performance to the licensing decision, anchored in the ICAO language-proficiency requirements and the published test design. The quantitative operational validity monitoring layer provides backing for those inferences: Classical Test Theory results show that Paper 1 forms are reliable and consistently targeted; a Rasch (IRT) calibration shows that item difficulty is well-aligned with candidate ability across forms; and a Many-Facet Rasch model for Paper 2 quantifies examiner severity (range -6.67 to 5.18 logits) on the same scale as candidate ability, with high separation reliability. None of these pieces of evidence are sufficient on their own; together they form a coherent case that an ELPAC score means what the licensing rule needs it to mean.
What we plan to add next
The published evidence is intentionally incomplete. The following items are recognised gaps in the validity argument and are the subject of planned work; none of them, on the current evidence, displaces the headline finding above.
- Predictive evidence against operational radiotelephony. This is a recognised and structurally difficult gap: because ELPAC is independent of the candidate's operational employer, individual test takers cannot be routinely followed into line operations, and a conventional predictive criterion study against radiotelephony performance is hard to mount. The programme is therefore seeking alternative forms of feedback on the operational meaningfulness of ELPAC scores — including structured stakeholder and user feedback, periodic expert review of materials against the target language use domain, and, where lawfully available and de-identified, signals from communication-related occurrence reporting and downstream training outcomes — rather than relying on a single criterion study.
- Cross-form equating. Formal IRT equating across the seven Paper 1 forms, to confirm that scores from different forms are interchangeable at the Level 4 cut-score.
- OPE fluency descriptor. The current Paper 2 data-capture format records a single OPE-fluency value although the OPE assesses fluency through two underlying sub-criteria. The capture format is being revised; OPE fluency will be reintroduced into the published agreement statistics once the two sub-criteria are recorded separately.
- Examiner-level inter-rater coefficients at rating level. The Many-Facet Rasch model already quantifies examiner severity in this window; future work will publish complementary κ-style coefficients at rating level once the OPE-fluency capture change is in place.
What this means in plain language
ELPAC is the English language proficiency assessment used by EUROCONTROL to determine whether air traffic controllers and pilots meet the language requirements set out in ICAO Annex 1 (Personnel Licensing) and ICAO Doc 9835 (Manual on the Implementation of ICAO Language Proficiency Requirements). It is administered for the purpose of supporting licensing decisions by the competent authority.
The assessment consists of two components. Paper 1 is a computer-delivered listening comprehension component applied as a screening instrument for candidates whose performance is demonstrably below ICAO Operational Level 4. The evidence demonstrates that the component produces consistent results across the operational forms and that item difficulty is appropriately targeted to the candidate population. Paper 2 is a face-to-face speaking assessment conducted by two accredited examiners — an English Language Expert (ELE) and an Operational Expert (OPE), a serving or former pilot or air traffic controller — applying the ICAO holistic descriptors and ICAO rating scale criteria. The evidence demonstrates that inter-examiner agreement is achieved in the substantial majority of assessments, and that residual differences in examiner severity are quantified, monitored, and reported.
One important limit should be stated: ELPAC is run independently of the airlines and air-navigation service providers that employ the candidates, so the programme cannot routinely monitor candidates' performance in operational service and cannot, by itself, prove that a given ELPAC score translates into a given level of performance on the radio. Closing that gap with a classical predictive study is genuinely difficult, and the programme is instead working to gather other forms of feedback — from competent authorities, training organisations and, where available, occurrence data — to keep checking that ELPAC scores remain meaningful in operational use. On the evidence presented, ELPAC is intended to contribute to the technical basis on which competent authorities determine compliance with the ICAO language proficiency requirements for licensing purposes; the determination itself, and the weight given to ELPAC alongside other evidence, rests with the national authority.
For data, methodology or reuse questions, please use the contact page.