Skip to main content
EUROCONTROL
ELPAC

About this page. This validity documentation is maintained by the ELPAC team at EUROCONTROL as a living methodological record. It is not an ICAO endorsement and does not by itself constitute a regulatory determination. The ICAO Doc 9835 and Doc 10197 pages in this section are ELPAC's own self-assessments and have not been reviewed or endorsed by ICAO. National civil aviation authorities remain responsible for licensing decisions and for determining whether ELPAC results satisfy the ICAO language-proficiency requirements in their jurisdiction.

Last reviewed: 7 September 2026.

Summary of the ELPAC validity evidence

A short, plain-language overview of what the evidence on this site shows about the ELPAC test, why it is suitable for its intended use, and what further work is planned.

Regulatory statement

Statement on the fitness of ELPAC for licensing use

1. Purpose and scope. The English Language Proficiency for Aeronautical Communication (ELPAC) test is the language proficiency assessment administered by EUROCONTROL for the purpose of determining whether air traffic controllers and flight crew members meet the language proficiency requirements set out in ICAO Annex 1 (Personnel Licensing), §1.2.9 and Appendix 1, as elaborated in ICAO Doc 9835 (Manual on the Implementation of ICAO Language Proficiency Requirements). The assessment is administered for the purpose of supporting licensing decisions; the licensing decision itself rests with the competent national civil aviation authority.

2. Construct and design alignment. ELPAC comprises two components, each aligned with the construct of aeronautical radio communication. Paper 1, a computer-delivered listening comprehension component, operates as a screening instrument targeted at the ICAO Operational Level 4 threshold. Paper 2 is a live speaking test rated by two accredited assessors: an English Language Expert (ELE) and an Operational Expert (OPE), an operational or former air traffic controller or pilot. Each assessor independently applies the ELPAC assessment criteria, which derive from the ICAO rating scale and the holistic descriptors. The ICAO level awarded is the lowest of the six criteria; there is no compensatory marking. Where the two assessors disagree on the outcome, a third assessor decides.

3. Evidence supporting use. On the operational dataset published on this site, Paper 1 forms exhibit reliability and item-difficulty targeting consistent with their intended screening function across the seven operational forms. Paper 2 inter-assessor outcomes demonstrate agreement in the substantial majority of paired sessions, with residual differences in assessor severity quantified on a common scale through a Many-Facet Rasch model and with a low rate of escalation to a third assessor. The quantitative evidence is integrated into a qualitative assessment use argument framed in the terms of Kane (2013), Chapelle (2020) and Chapelle & Voss (2021), and is documented in full on the linked operational-validity pages.

4. Limitations and ongoing work. The published evidence is intentionally incomplete, and one limitation is structural rather than methodological. ELPAC is administered independently of the test taker's employer and operational career: the test programme has no routine access to a test taker's subsequent radiotelephony performance on the line, and individual test takers cannot be systematically followed into live ATC or flight operations. A conventional predictive criterion study against operational radiotelephony performance is therefore difficult to mount and cannot, on its own, be relied upon to close the extrapolation inference. The programme accordingly pursues a range of complementary feedback mechanisms, including structured stakeholder and user feedback from competent authorities and approved training organisations, periodic expert review of test materials against the evolving target language use domain, and, where lawfully available and de-identified, signals from communication- related occurrence reporting and downstream training outcomes, together with formal IRT equating across Paper 1 forms, the reintroduction of OPE fluency once the data-capture format is revised, and complementary κ-style assessor-agreement coefficients at rating level. These activities are presented as ongoing validation work; no claim is made that the current evidence base directly demonstrates operational performance in service.

5. Conclusion. Subject to the limitations set out in paragraph 4, the evidence presented supports the use of ELPAC by competent authorities as one element of the technical basis for determining compliance with the ICAO language proficiency requirements for the licensing of air traffic controllers and flight crew members. No claim is made that ELPAC, on its own, establishes a test taker's operational radiotelephony performance in service. The determination of compliance, and the weight to be attached to ELPAC alongside other evidence available to the authority, remains with the national authority in accordance with the applicable national regulations.

Headline finding

ELPAC is fit for its intended purpose

On the operational dataset published on this site (N = 1,590 test takers over the reporting period 2026-01-05 to 2026-08-31), the Paper 1 listening comprehension screen reaches a Cronbach α reliability of 0.820.90 across the seven operational forms, the overall Paper 1 pass rate is 88.9 %, and the Paper 2 two-rater protocol (N = 894 sessions) escalates to a third assessor in only 1.57 % of cases. Taken together with the qualitative validity argument, this evidence is intended to support national civil aviation authorities in determining compliance with the ICAO language-proficiency requirements for air traffic controllers and pilots.

  • Evidenced on this site: Paper 1 reliability and Rasch calibration, and Paper 2 assessor-severity and agreement statistics.
  • Evidenced by design and documented practice: the construct of aeronautical radio communication, operationalised through the ELPAC assessment criteria which derive from the ICAO rating scale and the holistic descriptors, independent double rating with a defined reconciliation route, and centralised quality assurance.
  • Not yet evidenced: a predictive study against in-service radiotelephony performance, for the structural reasons set out below.
Why it is fit for purpose

How the design and the evidence fit together

The ELPAC programme builds its case for use in two layers. The qualitative assessment use argument (Kane 2013; Chapelle 2020; Chapelle & Voss 2021) chains six warranted inferences from the test performance to the licensing decision, based on the ICAO language-proficiency requirements and the published test design. The quantitative operational validity monitoring layer provides backing for those inferences: Classical Test Theory results show that Paper 1 forms are reliable and consistently targeted; a Rasch (IRT) calibration shows that item difficulty is well-aligned with test taker ability across forms; and a Many-Facet Rasch model for Paper 2 quantifies assessor severity (range -6.67 to 5.18 logits) on the same scale as test taker ability, with high separation reliability. None of these pieces of evidence are sufficient on their own; together they form a coherent case that an ELPAC score means what the licensing rule needs it to mean.

Identified areas for future research

What we plan to add next

The published evidence is intentionally incomplete. The following items are recognised gaps in the validity argument and are the subject of planned work; none of them, on the current evidence, displaces the headline finding above.

  • Predictive evidence against operational radiotelephony. This is a recognised and structurally difficult gap: because ELPAC is independent of the test taker's operational employer, individual test takers cannot be routinely followed into line operations, and a conventional predictive criterion study against radiotelephony performance is hard to mount. The programme is therefore seeking alternative forms of feedback on the operational meaningfulness of ELPAC scores, including structured stakeholder and user feedback, periodic expert review of materials against the target language use domain, and, where lawfully available and de-identified, signals from communication-related occurrence reporting and downstream training outcomes, rather than relying on a single criterion study.
  • Cross-form equating. Formal IRT equating across the seven Paper 1 forms, to confirm that scores from different forms are interchangeable at the Level 4 cut-score.
  • OPE fluency descriptor. The current Paper 2 data-capture format records a single OPE-fluency value although the OPE assesses fluency through two underlying sub-criteria. The capture format is being revised; OPE fluency will be reintroduced into the published agreement statistics once the two sub-criteria are recorded separately.
  • Assessor-level inter-rater coefficients at rating level. The Many-Facet Rasch model already quantifies assessor severity in this window; future work will publish complementary κ-style coefficients at rating level once the OPE-fluency capture change is in place.
In plain language

What this means in plain language

ELPAC is the English language proficiency assessment used by EUROCONTROL to determine whether air traffic controllers and pilots meet the language requirements set out in ICAO Annex 1 (Personnel Licensing) and ICAO Doc 9835 (Manual on the Implementation of ICAO Language Proficiency Requirements). It is administered for the purpose of supporting licensing decisions by the competent authority.

The assessment consists of two components. Paper 1 is a computer-delivered listening comprehension component applied as a screening instrument for test takers whose performance is demonstrably below ICAO Operational Level 4. The evidence demonstrates that the component produces consistent results across the operational forms and that item difficulty is appropriately targeted to the test taker population. Paper 2 is a live speaking test conducted by two accredited assessors: an English Language Expert (ELE) and an Operational Expert (OPE), an operational or former air traffic controller or pilot, applying the ELPAC assessment criteria, which derive from the ICAO rating scale and the holistic descriptors. The evidence demonstrates that inter-assessor agreement is achieved in the substantial majority of assessments, and that residual differences in assessor severity are quantified, monitored, and reported.

One important limit should be stated: ELPAC is run independently of the airlines and air-navigation service providers that employ the test takers, so the programme cannot routinely monitor test takers' performance in operational service and cannot, by itself, prove that a given ELPAC score translates into a given level of performance on the radio. Closing that gap with a classical predictive study is genuinely difficult, and the programme is instead working to gather other forms of feedback from competent authorities, training organisations and, where available, occurrence data, to keep checking that ELPAC scores remain meaningful in operational use. On the evidence presented, ELPAC is intended to contribute to the technical basis on which competent authorities determine compliance with the ICAO language proficiency requirements for licensing purposes; the determination itself, and the weight given to ELPAC alongside other evidence, rests with the national authority.