This document records the interpretation and use argument (IUA) supporting the ELPAC assessment, structured in accordance with the framework of Kane (1992, 2006, 2013) and the application to second-language testing set out in Chapelle, Enright and Jamieson (2008). The argument is decomposed into six inferences: domain description, evaluation, generalisation, explanation, extrapolation and utilisation.
The argument applies to the use of ELPAC scores by competent authorities and approved organisations for the purpose of issuing, endorsing or renewing the ICAO language proficiency endorsement under ICAO Annex 1, Commission Regulation (EU) 2015/340 (ATCO.B.030) and Commission Regulation (EU) No 1178/2011 (FCL.055).
The qualitative argument set out in the sections that follow is supported by quantitative operational validity monitoring evidence published on this site, comprising: Classical Test Theory and Rasch (IRT) psychometric evidence for Paper 1 across seven operational forms; a Many-Facet Rasch examiner-severity model and rater-agreement statistics for Paper 2; and operational outcome statistics for the 2026 reporting cohort (N = 1,590; reporting period 2026-01-05 to 2026-08-31). Further studies identified in the programme of continuing validation are planned and do not affect the determination set out below.
The ICAO construct and the ELPAC test construct
A validity argument is anchored in an explicit statement of what the test is intended to measure. This section distinguishes two related but non-identical constructs: the regulatory construct defined by ICAO, and the test construct that ELPAC operationalises. ELPAC does not define its own construct — it inherits ICAO's and renders it assessable.
1. The ICAO construct (regulatory)
ICAO Annex 1 (Personnel Licensing), §1.2.9 and Appendix 1, together with ICAO Doc 9835 (Manual on the Implementation of ICAO Language Proficiency Requirements), define the regulated construct as the operational language proficiency required for safe radiotelephony communication by licensed air traffic controllers and flight crew. ICAO defines this construct in two complementary ways:
- Holistic descriptors (Doc 9835 §4.6 / Attachment A) — the qualitative definition of what an operationally proficient speaker can do: communicate in voice-only and face-to-face situations; handle work-related topics with accuracy and clarity; use communicative strategies to recognise and resolve misunderstanding; handle a complication or unexpected turn of events; and remain intelligible to the aeronautical community.
- ICAO Rating Scale — six descriptors (Pronunciation, Structure, Vocabulary, Fluency, Comprehension, Interactions) across six levels, with Operational Level 4 as the licensing threshold.
Two properties of the ICAO construct are decisive for what follows. It is communicative and operational, not general English; and it is bounded to the ICAO radiotelephony target language use (TLU) domain — it does not extend to cabin-crew, ground-handling or engineering English, and it does not evaluate the technical or procedural correctness of the underlying operational decision.
2. The ELPAC test construct
The construct ELPAC sets out to measure is operational communicative competence in aviation radiotelephony: the sub-construct of the ICAO regulatory construct that is realised in a candidate's observable language behaviour under operationally relevant conditions. "Communicative competence" is taken in the tradition that runs from Hymes (1972) and Canale & Swain (1980) through Bachman (1990) and Bachman & Palmer (2010), restricted to the ICAO TLU domain. Substantively, the construct is defined by the ICAO holistic descriptors; gradably, it is operationalised through the six ICAO Rating Scale descriptors. Throughout the argument below, "language proficiency" is to be read in this operational, communicative sense.
- Pronunciation, Structure, Vocabulary, Fluency, Comprehension, Interactions — as features of communicative behaviour
- Handling of routine and non-routine work-related communication
- Communicative strategies to check, confirm, clarify and resolve misunderstanding
- Intelligibility to the aeronautical community
- General English proficiency beyond the ICAO TLU domain
- Technical or procedural correctness of ATC or piloting decisions
- Personality, confidence or interpersonal style
- Accent conformity beyond intelligibility to the aeronautical community
3. Construct alignment
The crosswalk below records, element by element, how the ICAO construct is operationalised in ELPAC and where in the argument each element is discharged.
| ICAO construct element | ELPAC operationalisation | Inference |
|---|---|---|
| Holistic descriptors — voice-only + face-to-face; routine + non-routine; communicative strategies | Paper 1 authentic radiotelephony audio; Paper 2 three-task progression (routine → complication → extended interaction) | 1, 2 |
| Rating Scale — six descriptors, six levels, Level 4 threshold | Independent double rating by ELE + OPE against the six descriptors; final award is the lowest sub-rating | 2, 4 |
| Operational relevance — intelligibility to the aeronautical community | OPE (serving or former ATCO or pilot) on every Paper 2 panel; separate ATC and Pilot versions | 1, 2, 4 |
| Level 4 as the regulated licensing cut | Paper 1 as a Level-4-targeted screen; Paper 2 lowest-of-six composite award | 3, 6 |
What this argument concludes
- ELPAC measures the construct ICAO regulates — operational communicative competence in aviation radiotelephony — and not general English.
- The two papers together elicit rateable performance on all six ICAO Rating Scale descriptors, and Paper 2 is double-rated by an English Language Expert and an Operational Expert with a defined third-examiner route where they cannot reconcile.
- On the qualitative argument below and the published psychometric evidence, ELPAC is determined to provide an appropriate technical basis for ICAO language proficiency licensing decisions.
- The two weakest links are extrapolation to in-service radiotelephony performance and the consequences of the licensing decision. Both rest on test and procedural design rather than a criterion study, for the structural reasons set out in Inference 5 and the continuing validation programme.
The argument in one table
The six inferences below form the spine of the argument. Each row links to the full section. Strong means published quantitative backing on this site; partial means design and procedural backing with a planned quantitative study; design-only means the warrant is supported by task and construct design rather than by a single criterion study, for the structural reasons set out in the relevant section.
| Inference | One-line warrant | Backing |
|---|---|---|
| 1. Domain description | The TLU domain is ICAO radiotelephony, as specified in Doc 9835. | strong |
| 2. Evaluation | Paper 1 and Paper 2 jointly elicit rateable performance on the six ICAO descriptors. | strong |
| 3. Generalisation | The awarded level generalises across forms, sessions, examiners and centres. | partial |
| 4. Explanation | Scoring procedures constrain examiner and operational-expertise variance. | strong |
| 5. Extrapolation | The score supports an inference to operational radiotelephony performance. | design-only |
| 6. Utilisation | Licensing decisions made from the score are appropriate and their consequences are bounded. | design-only |
Crosswalk to the four plain questions
Green (2013, ch. 5) reduces the full Kane chain to four plain questions a regulator can put to a test. Those questions map onto the six inferences set out on this page as follows; no separate evidence base is maintained for the Green view.
| Green (2013) question | Kane inference |
|---|---|
| Scoring — is the score accurate? | 4. Explanation |
| Generalisation — would a different form/day/examiner give a similar score? | 3. Generalisation |
| Extrapolation — does doing well on the test mean doing well on the radio? | 5. Extrapolation |
| Utilisation — are the resulting decisions better than the alternatives? | 6. Utilisation |
Domain description
Backing: strongClaim
Show the evidence on ELPACHide the evidence on ELPAC
The test is anchored in ICAO Doc 9835 (Manual on the Implementation of ICAO Language Proficiency Requirements) and in the ICAO Language Proficiency Rating Scale. The Rating Scale characterises operational language along six descriptors — Pronunciation, Structure, Vocabulary, Fluency, Comprehension and Interactions — across six levels. Separate ATC and Pilot versions are designed so that the domain sampled in each version matches the candidate's role, covering standard phraseology and plain aviation English in routine and non-routine situations.
Alongside the Rating Scale, ICAO Doc 9835 (§4.6 and Attachment A) sets out holistic descriptors that define, qualitatively, what an operationally proficient speaker can do. ELPAC takes the holistic descriptors as the substantive description of the target domain. They straddle domain and construct: they describe the TLU domain here (Inference 1) and inform what is rated in Paper 2 (Inference 2). An operationally proficient speaker is one who can:
- communicate effectively in voice-only (telephone / radiotelephony) and in face-to-face situations;
- communicate on common, concrete and work-related topics with accuracy and clarity;
- use appropriate communicative strategies to exchange messages and to recognise and resolve misunderstandings (e.g. to check, confirm or clarify information) in a general or work-related context;
- handle successfully and with relative ease the linguistic challenges presented by a complication or unexpected turn of events that occurs within the context of a routine work situation or communicative task with which they are otherwise familiar; and
- use a dialect or accent that is intelligible to the aeronautical community.
Together, the holistic descriptors and the six-descriptor Rating Scale specify the operational language use domain: the Rating Scale supplies the gradable criteria, and the holistic descriptors anchor those criteria in operationally meaningful behaviour.
Limitations and rebuttals
Evaluation (construct representation)
Backing: strongClaim
Show the evidence on ELPACHide the evidence on ELPAC
Limitations and rebuttals
Generalisation
Backing: partialClaim
Show the evidence on ELPACHide the evidence on ELPAC
Limitations and rebuttals
Generalisation is conditional on the continued operation of the quality-assurance processes described above. Local conditions (microphone quality, ambient noise, examiner fatigue) introduce residual variance. At cut-scores between bands, coarse categorisation increases the consequences of small measurement differences for candidates whose true level lies near a boundary.
The minimum-across-sub-ratings convention — mandated by the ICAO rating scale criteria and discussed in detail under Inference 6 — means that the reliability of the composite final level is, in general, lower than that of any single sub-rating (Knoch 2009). Procedural quality assurance constrains rater and form variance at the sub-rating level; the reliability of the minimum-based composite itself is monitored separately, through the Many-Facet Rasch examiner-severity model published in the Paper 2 operational validity monitoring section.
Explanation (scoring and response process)
Backing: strongClaim
Show the evidence on ELPACHide the evidence on ELPAC
Limitations and rebuttals
Extrapolation to operational performance
Backing: design-onlyClaim
Show the evidence on ELPACHide the evidence on ELPAC
Limitations and rebuttals
Utilisation (decision and consequences)
Backing: design-onlyClaim
Show the evidence on ELPACHide the evidence on ELPAC
Limitations and rebuttals
Overall determination
On the basis of the qualitative interpretation and use argument set out above and the quantitative operational validity monitoring evidence published on this site, the ELPAC assessment is determined to provide an appropriate technical basis for licensing decisions made by competent authorities under ICAO Annex 1, Commission Regulation (EU) 2015/340 and Commission Regulation (EU) No 1178/2011. The strongest warrants are the anchoring of the construct in ICAO Doc 9835, the double-rated independent assessment of Paper 2 with defined reconciliation, and the centralised quality assurance applied to forms, examiners and accredited test centres. Reliability of Paper 1 across operational forms, examiner severity for Paper 2, and the third-assessor escalation rate for the 2026 reporting cohort are within the ranges expected for a high-stakes language proficiency assessment of this design.
Operational outcomes, 2026 reporting period (N = 1,590)
The figures below are derived from the de-identified operational test register for the period stated. They are reported as observed counts and rates only, in support of the consequence and generalisation inferences set out above. These values are descriptive only; they do not constitute predictive-validity evidence against operational radiotelephony performance.
Final ICAO level, candidates with a Paper 1 pass
Distribution of the final score — the lowest of the six sub-ratings, in accordance with the ICAO rating scale criteria — for candidates progressing past Paper 1 in the reporting period.
| Final result | Candidates | Share |
|---|---|---|
| Level 3 | 48 | 3.4 % |
| ICAO Level 4 | 460 | 32.6 % |
| ICAO Level 5 | 519 | 36.7 % |
| ICAO Level 6 | 259 | 18.3 % |
| Fail | 7 | 0.5 % |
| Pending / not recorded | 120 | 8.5 % |
Paper 1 result × Final score (observed counts)
Descriptive cross-tabulation of the Paper 1 screening decision against the final ICAO level recorded for the same candidate in the reporting period. The table records what occurred; it is not a misclassification matrix and does not characterise the accuracy of the screening decision.
| Paper 1 | Level 3 | ICAO Level 4 | ICAO Level 5 | ICAO Level 6 | Fail | Not recorded |
|---|---|---|---|---|---|---|
| Fail | 0 | 0 | 0 | 0 | 177 | 0 |
| Pass | 48 | 460 | 519 | 259 | 7 | 120 |
Paper 2 third-assessor escalation rate
A third assessor was not invoked in any of the 894 Paper 2 sessions (0.00 %) in the reporting window 2025-07-01 to 2026-06-18 — the two-rater protocol produced a concordant award without escalation in every case. This figure records the escalation rate only and shall not be substituted for a formal inter-rater agreement coefficient; full paired-rater statistics and a Many-Facet Rasch examiner-severity model for Paper 2 are published on the Paper 2 tab.
- Counts are observed values over the reporting window; no inferential or causal interpretation is applied.
- Paper 1 fails that show a recorded Paper 2 level reflect the operational test register as supplied and are not interpreted as misclassification.
- The third-assessor figure is an escalation rate and is not a substitute for an inter-rater agreement coefficient (e.g., Cohen's kappa, many-facet Rasch).
Programme of continuing validation
The ELPAC programme operates a continuing validation programme alongside routine operational validity monitoring. Planned activities include formal cross-form equating of Paper 1, differential item functioning analyses once the necessary demographic frame is in place, and the reintroduction of OPE-fluency into the published Paper 2 agreement statistics following the ongoing capture-format revision. A classical predictive criterion study linking ELPAC scores to operational radiotelephony performance is structurally difficult, as the test is administered independently of the candidate's employer and the programme has no routine access to in-service performance data; the programme therefore additionally pursues complementary feedback mechanisms — structured stakeholder and user feedback, periodic expert review of materials, and, where lawfully available and de-identified, signals from communication- related occurrence reporting and downstream training outcomes — in support of the extrapolation inference. These activities are part of continuous improvement and do not, on the current evidence, displace the determination of fitness for purpose set out above.
References
Scholarly literature
- Alderson, J. C. (2009). Air safety, language assessment policy, and policy implementation: Responsibilities and verifiability. Annual Review of Applied Linguistics, 29, 168–187.
- Alderson, J. C. (2010). A survey of aviation English tests. Language Testing, 27(1), 51–72.
- Alderson, J. C., & Wall, D. (1993). Does washback exist? Applied Linguistics, 14(2), 115–129.
- Bachman, L. F. (1990). Fundamental Considerations in Language Testing. Oxford University Press.
- Bachman, L. F., & Palmer, A. S. (1996). Language Testing in Practice. Oxford University Press.
- Bachman, L. F., & Palmer, A. S. (2010). Language Assessment in Practice. Oxford University Press.
- Bailey, K. M. (1996). Working for washback: A review of the washback concept in language testing. Language Testing, 13(3), 257–279.
- Bond, T. G., & Fox, C. M. (2015). Applying the Rasch model: Fundamental measurement in the human sciences (3rd ed.). Routledge.
- Canale, M., & Swain, M. (1980). Theoretical bases of communicative approaches to second language teaching and testing. Applied Linguistics, 1(1), 1–47.
- Chapelle, C. A., Enright, M. K., & Jamieson, J. M. (Eds.) (2008). Building a Validity Argument for the Test of English as a Foreign Language. Routledge.
- Emery, H. J. (2014). Developments in LSP testing 30 years on? The case of aviation English. Language Assessment Quarterly, 11(2), 198–215.
- Estival, D., Farris, C., & Molesworth, B. (2016). Aviation English: A lingua franca for pilots and air traffic controllers. Routledge.
- Green, A. (2013). Exploring language assessment and testing. Routledge.
- Hughes, A. (2003). Testing for language teachers (2nd ed.). Cambridge University Press.
- Hymes, D. H. (1972). On communicative competence. In J. B. Pride & J. Holmes (Eds.), Sociolinguistics (pp. 269–293). Penguin.
- Kane, M. T. (1992). An argument-based approach to validity. Psychological Bulletin, 112(3), 527–535.
- Kane, M. T. (2006). Validation. In R. L. Brennan (Ed.), Educational Measurement (4th ed., pp. 17–64). American Council on Education / Praeger.
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.
- Kim, H., & Elder, C. (2009). Understanding aviation English as a lingua franca. Australian Review of Applied Linguistics, 32(3), 23.1–23.17.
- Knoch, U. (2009). Diagnostic assessment of writing: A comparison of two rating scales. Language Testing, 26(2), 275–304.
- Knoch, U. (2014). Using subject specialists to validate an ESP rating scale: The case of the ICAO rating scale. English for Specific Purposes, 33, 77–86.
- Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational Measurement (3rd ed., pp. 13–103). American Council on Education / Macmillan.
- Messick, S. (1996). Validity and washback in language testing. Language Testing, 13(3), 241–256.
- Mislevy, R. J., Steinberg, L. S., & Almond, R. G. (2003). Focus article: On the structure of educational assessments. Measurement: Interdisciplinary Research and Perspectives, 1(1), 3–62.
- Pill, J., & McNamara, T. (2016). How much is enough? Involving occupational experts in setting standards on a specific-purpose language test for health professionals. Language Testing, 33(2), 217–234.
- Read, J., & Knoch, U. (2009). Clearing the air: Applied linguistic perspectives on aviation communication. Australian Review of Applied Linguistics, 32(3), 21.1–21.11.
Standards and regulations
- ICAO Doc 9835 — Manual on the Implementation of ICAO Language Proficiency Requirements (2nd ed., 2010); in particular §4.6 and Attachment A (Holistic Descriptors and ICAO Language Proficiency Rating Scale).
- ICAO Annex 1 — Personnel Licensing.
- ICAO Annex 10, Volume II — Aeronautical Telecommunications (Communication Procedures).
- Commission Regulation (EU) 2015/340 — Technical requirements and administrative procedures relating to air traffic controllers' licences and certificates (ATCO.B.030).
- Commission Regulation (EU) No 1178/2011, Part-FCL (FCL.055) — Flight crew licensing, language proficiency.