This document records the assessment use argument supporting the ELPAC test. It sets out the chain of claims and evidence that links test performance to the licensing decision, following the argument-based approach of Kane (1992, 2006, 2013) as developed for language assessment by Chapelle (2020) and Chapelle and Voss (2021).
The argument applies to the use of ELPAC scores by competent authorities and approved organisations for the purpose of issuing, endorsing or renewing the ICAO language proficiency endorsement under ICAO Annex 1, Commission Regulation (EU) 2015/340 (ATCO.B.030) and Commission Regulation (EU) No 1178/2011 (FCL.055).
The qualitative argument set out in the sections that follow is supported by quantitative operational validity monitoring evidence published on this site, comprising: Classical Test Theory and Rasch (IRT) psychometric evidence for Paper 1 across seven operational forms; a Many-Facet Rasch assessor-severity model and rater-agreement statistics for Paper 2; and operational outcome statistics for the 2026 reporting cohort (N = 1,590; reporting period 2026-01-05 to 2026-08-31). Further studies identified in the programme of continuing validation are planned and do not affect the claim and evidence status set out below.
The ICAO language proficiency requirements and the ELPAC test construct
An assessment use argument is based on an explicit statement of what the test is intended to measure. Two distinct things are involved. The ICAO language proficiency requirements are a given regulatory standard: they are set by ICAO and are not open to reinterpretation by a test provider. The ELPAC test construct is the single measurement construct of this test. It derives from the ICAO language proficiency requirements and operationalises them in the test.
1. The ICAO language proficiency requirements (given standard)
ICAO Annex 1 (Personnel Licensing), §1.2.9 and Appendix 1, together with ICAO Annex 10 Volume II and ICAO Doc 9835 (Manual on the Implementation of ICAO Language Proficiency Requirements), require licensed air traffic controllers and flight crew to demonstrate the operational language proficiency needed for safe radiotelephony communication. The requirement has two components:
- Holistic descriptors (Doc 9835 §4.6 / Attachment A): the qualitative statement of what an operationally proficient speaker can do: communicate in voice-only and face-to-face situations; handle work-related topics with accuracy and clarity; use communicative strategies to recognise and resolve misunderstanding; handle a complication or unexpected turn of events; and remain intelligible to the aeronautical community.
- ICAO Rating Scale: six descriptors (Pronunciation, Structure, Vocabulary, Fluency, Comprehension, Interactions) across six levels, with Operational Level 4 as the licensing threshold.
Two properties of the requirements are decisive for what follows. They are operational and communicative rather than concerned with general English; and they are bounded to the ICAO radiotelephony target language use (TLU) domain: they do not extend to cabin-crew, ground-handling or engineering English, and they are silent on the technical or procedural correctness of the underlying operational decision.
2. The ELPAC test construct
The ELPAC test construct derives from the ICAO language proficiency requirements and operationalises them in the test. It is operational communicative competence in aviation radiotelephony: the language behaviour a test taker displays, observably, under operationally relevant conditions. Substantively it is defined by the ICAO holistic descriptors, gradably it is operationalised through the six ICAO Rating Scale descriptors, and its scope is delimited by the ICAO radiotelephony target language use domain. "Communicative competence" is taken in the tradition that runs from Hymes (1972) and Canale & Swain (1980) through Bachman (1990) and Bachman & Palmer (2010), restricted to that domain. Throughout the argument below, "language proficiency" is to be read in this operational, communicative sense. The ELPAC test construct is also based on research into the target language use domain (Agius, 2024).
- Pronunciation, Structure, Vocabulary, Fluency, Comprehension, Interactions — as features of communicative behaviour
- Handling of routine and non-routine work-related communication
- Communicative strategies to check, confirm, clarify and resolve misunderstanding
- Intelligibility to the aeronautical community
- Mutual intelligibility and communicative success
- General English proficiency beyond the ICAO radiotelephony target language use domain
- Technical or procedural correctness of ATC or piloting decisions
- Personality, confidence or interpersonal style
- Accent conformity beyond intelligibility to the aeronautical community
3. Alignment with the ICAO language proficiency requirements
The crosswalk below records, element by element, how each requirement is operationalised in the ELPAC test construct and where in the argument it is discharged.
| ICAO requirement element | ELPAC operationalisation | Inference |
|---|---|---|
| Holistic descriptors — voice-only + face-to-face; routine + non-routine; communicative strategies | Paper 1 authentic radiotelephony audio; Paper 2 three-task progression (routine → complication → extended interaction) | 1, 2 |
| Rating Scale — six descriptors, six levels, Level 4 threshold | Independent double rating by ELE + OPE against the six criteria; final award is the lowest of the six | 2, 4 |
| Operational relevance — intelligibility to the aeronautical community | OPE (an operational or former ATCO or pilot) on every Paper 2 panel; separate ATC and Pilot versions | 1, 2, 4 |
| Level 4 as the regulated licensing cut | Paper 1 as a Level-4-targeted screen; Paper 2 lowest-of-six composite award | 3, 6 |
How the ELPAC Assessment Scheme relates to the ICAO Rating Scale
Read the explanatory noteHide the explanatory note
The ICAO Language Proficiency Rating Scale is the licensing reference: it defines six proficiency levels across six descriptors and sets Operational level 4 as the minimum for licensing. ICAO designed and owns that scale.
The Rating Scale describes proficiency in general terms. In a live test, however, not every descriptor is equally observable from a single sample of performance. To help accredited ELPAC assessors — the English Language Expert (ELE) and the Operational Expert (OPE) — place observed performance consistently on the ICAO scale, ELPAC uses the ELPAC Assessment Scheme. The scheme provides observable, operationally grounded criteria that translate the Rating Scale descriptors into test-observable behaviour. It is an operational interpretation to support reliable grading; it does not replace, revise or supersede the ICAO Rating Scale.
The ICAO Language Proficiency Rating Scale is like a set of diagnostic categories. A clinician does not treat the category itself; they look for observable signs and symptoms to decide which category best describes the patient's condition. The ELPAC Assessment Scheme supplies those observable criteria for language performance, so that the assessor can map observed spoken language onto the ICAO rating scale consistently and reliably.
The Assessment Scheme is published in the Documents library and is used in ELPAC assessor training and standardisation.
What this argument concludes
- This page sets out the assessment use argument for the ELPAC test: a structured, evidence-based justification for using ELPAC scores as a basis for ICAO language proficiency licensing decisions, organised around the six inferences of domain description, evaluation, generalisation, explanation, extrapolation and utilisation.
- The ELPAC test construct derives from the ICAO language proficiency requirements and operationalises them in the test. It is operational communicative competence in aviation radiotelephony — observable language behaviour under operationally relevant conditions — covering pronunciation, structure, vocabulary, fluency, comprehension, interaction, routine and non-routine communication, communicative strategies, intelligibility, and mutual intelligibility and communicative success.
- The two papers together elicit rateable performance on all six ICAO Rating Scale descriptors. Paper 2 is double-rated by an English Language Expert and an Operational Expert, with a defined third-assessor route where the two assessors disagree on the outcome.
- On the qualitative argument below and the published psychometric evidence, ELPAC provides an appropriate technical basis for ICAO language proficiency licensing decisions; the current status of the supporting evidence is set out in the claim and evidence status section.
- The two weakest links are extrapolation to in-service radiotelephony performance and the consequences of the licensing decision. Both rest on test and procedural design rather than a criterion study, for the structural reasons set out in Inference 5 and the continuing validation programme.
- A separate criterion-by-criterion overview sets out how ELPAC relates to the eight test design criteria in ICAO Doc 10197.
- An item-by-item overview covers the checklist for aviation language testing in ICAO Doc 9835, Appendix C.
The argument in one table
The six inferences below form the spine of the argument. Each row links to the full section. Strong means published quantitative backing on this site; partial means design and procedural backing with a planned quantitative study; design-only means the warrant is supported by task and construct design rather than by a single criterion study, for the structural reasons set out in the relevant section. Two kinds of support are distinguished throughout: empirical evidence (data published on this site) and documented practice (published procedures, design decisions and quality assurance).
| Inference | One-line warrant | Backing |
|---|---|---|
| 1. Domain description | The TLU domain is ICAO radiotelephony, as specified in Doc 9835. | strong |
| 2. Evaluation | Paper 1 and Paper 2 jointly elicit rateable performance on the six ICAO descriptors. | strong |
| 3. Generalisation | The awarded level generalises across forms, sessions, assessors and centres. | partial |
| 4. Explanation | Scoring procedures constrain assessor and operational-expertise variance. | strong |
| 5. Extrapolation | The score supports an inference to operational radiotelephony performance. | design-only |
| 6. Utilisation | Licensing decisions made from the score are appropriate and their consequences are bounded. | design-only |
Crosswalk to the four plain questions
Green (2013, ch. 5) reduces the full six-inference chain to four plain questions a regulator can put to a test. Those questions map onto the six inferences of the assessment use argument set out on this page as follows; no separate evidence base is maintained for the Green view.
| Green (2013) question | Assessment use argument inference |
|---|---|
| Scoring — is the score accurate? | 4. Explanation |
| Generalisation — would a different form/day/assessor give a similar score? | 3. Generalisation |
| Extrapolation — does doing well on the test mean doing well on the radio? | 5. Extrapolation |
| Utilisation — are the resulting decisions better than the alternatives? | 6. Utilisation |
Domain description
Backing: strongClaim
Show the evidence on ELPACHide the evidence on ELPAC
The test is based on ICAO Doc 9835 (Manual on the Implementation of ICAO Language Proficiency Requirements) and on the ICAO Language Proficiency Rating Scale. The Rating Scale characterises operational language along six descriptors (Pronunciation, Structure, Vocabulary, Fluency, Comprehension and Interactions) across six levels. Separate ATC and Pilot versions are designed so that the domain sampled in each version matches the test taker's role, covering standard phraseology and plain aviation English in routine and non-routine situations.
Alongside the Rating Scale, ICAO Doc 9835 (§4.6 and Attachment A) sets out holistic descriptors that define, qualitatively, what an operationally proficient speaker can do. ELPAC takes the holistic descriptors as the substantive description of the target domain. They straddle domain and construct: they describe the TLU domain here (Inference 1) and inform what is rated in Paper 2 (Inference 2). An operationally proficient speaker is one who can:
- communicate effectively in voice-only (telephone / radiotelephony) and in face-to-face situations;
- communicate on common, concrete and work-related topics with accuracy and clarity;
- use appropriate communicative strategies to exchange messages and to recognise and resolve misunderstandings (e.g. to check, confirm or clarify information) in a general or work-related context;
- handle successfully and with relative ease the linguistic challenges presented by a complication or unexpected turn of events that occurs within the context of a routine work situation or communicative task with which they are otherwise familiar; and
- use a dialect or accent that is intelligible to the aeronautical community.
Together, the holistic descriptors and the six-descriptor Rating Scale specify the operational language use domain: the Rating Scale supplies the gradable criteria, and the holistic descriptors provide the operationally meaningful basis for those criteria. The target language use domain assessed in the ELPAC test was derived from existing research by Agius (2018, 2024), Knoch (2014), Kim (2018), Kim & Elder (2009, 2015) and Monteiro (2019).
Rebuttal: What could weaken this claim
Evaluation (construct representation)
Backing: strongClaim
Show the evidence on ELPACHide the evidence on ELPAC
Rebuttal: What could weaken this claim
Generalisation
Backing: partialClaim
Show the evidence on ELPACHide the evidence on ELPAC
Rebuttal: What could weaken this claim
Limitations
Generalisability is conditional on the continued statistical evaluation of the ELPAC test's performance. A variety of local factors may introduce residual construct-irrelevant variance. At cut-scores between bands, coarse categorisation increases the consequences of small measurement differences for test takers whose true level lies near a boundary.
The non-compensatory scoring convention mandated by the ICAO language proficiency requirements, and discussed in detail under Inference 6 (Utilisation), means that the reliability of the composite final level is, in general, lower than that of any single criterion (Knoch 2009). To mitigate this concern, the scoring mechanism of the ELPAC test has been adapted so that the OPE and ELE no longer need to come to a consensual decision about a test taker's performance. The reliability of test scores is monitored using statistical analysis instruments, published in the Paper 2 operational validity monitoring section.
Explanation (scoring and response process)
Backing: strongClaim
Show the evidence on ELPACHide the evidence on ELPAC
Paper 1 is computer-marked, which minimises the risk of rater variance in the scoring of the listening component; item-level and form-level sampling error remain. Paper 2 is rated by two assessors independently: an Operational Expert (OPE), who is an operational or former air traffic controller or pilot, and an English Language Expert (ELE). The two assessors take complementary perspectives: the OPE brings operational judgement suitable to the radiotelephony context, whereas the ELE brings linguistic judgement to the assessment. Each assessor rates the test taker independently against the ELPAC assessment criteria that are assigned to them according to their role as OPE or ELE.
The two independent ratings are then aggregated, so that the final ICAO level awarded to the test taker is the lowest score they achieve in any one of the six criteria in the ELPAC assessment scheme, in accordance with the ICAO rating scale criteria and the holistic descriptors. In cases where the two assessors disagree on the outcome of the test, the case is resolved through a documented dispute-resolution mechanism (a third assessor), rather than by on-the-spot consensus. Sessions are recorded, which supports adjudication, audit and post-hoc re-rating as part of routine quality assurance.
Rebuttal: What could weaken this claim
Limitations
Extrapolation to operational performance
Backing: design-onlyClaim
Show the evidence on ELPACHide the evidence on ELPAC
Rebuttal: What could weaken this claim
Limitations
Utilisation (decision and consequences)
Backing: design-onlyClaim
Show the evidence on ELPACHide the evidence on ELPAC
Rebuttal: What could weaken this claim
Limitations
Claim and evidence status
The claim made on this site is that the ELPAC test provides an appropriate technical basis for licensing decisions made by competent authorities under ICAO Annex 1, Commission Regulation (EU) 2015/340 and Commission Regulation (EU) No 1178/2011. For transparency towards expert readers, the status of the supporting evidence is stated in three groups.
- Evidenced on this site (empirical). Paper 1 reliability and item-difficulty targeting across the seven operational forms; Rasch (IRT) calibration of the Paper 1 item bank; Paper 2 assessor severity quantified on a common scale through a Many-Facet Rasch model, with paired-rater agreement statistics and the third-assessor escalation rate for the 2026 reporting cohort.
- Evidenced by design and documented practice. The construct of aeronautical radio communication set out in the construct section, operationalised through the ELPAC assessment criteria which derive from the ICAO rating scale and the holistic descriptors of ICAO Doc 9835; the independent double-rating of Paper 2 with a defined reconciliation route; and the ongoing statistical evaluation of the test, the assessors and the accredited test centres. These are the strongest procedural warrants in the argument; they are documented controls, not the results of a single empirical study.
- Not yet evidenced. A predictive criterion study linking ELPAC scores to in-service radiotelephony performance, which is structurally difficult for the reasons set out under Inference 5, and the studies listed in the programme of continuing validation below. The claim above is made subject to this status and does not assert that ELPAC, on its own, demonstrates operational performance in service.
Operational outcomes, 2026 reporting period (N = 1,590)
The figures below are derived from the de-identified operational test register for the period stated. They are reported as observed counts and rates only, in support of the consequence and generalisation inferences set out above. These values are descriptive only; they do not constitute predictive-validity evidence against operational radiotelephony performance.
Final ICAO level, test takers with a Paper 1 pass
Distribution of the final score (the lowest of the six criteria, with no compensatory marking) for test takers progressing past Paper 1 in the reporting period.
| Final result | Test takers | Share |
|---|---|---|
| Level 3 | 48 | 3.4 % |
| ICAO Level 4 | 460 | 32.6 % |
| ICAO Level 5 | 519 | 36.7 % |
| ICAO Level 6 | 259 | 18.3 % |
| Fail | 7 | 0.5 % |
| Pending / not recorded | 120 | 8.5 % |
Paper 1 result × Final score (observed counts)
Descriptive cross-tabulation of the Paper 1 screening decision against the final ICAO level recorded for the same test taker in the reporting period. The table records what occurred; it is not a misclassification matrix and does not characterise the accuracy of the screening decision.
| Paper 1 | Level 3 | ICAO Level 4 | ICAO Level 5 | ICAO Level 6 | Fail | Not recorded |
|---|---|---|---|---|---|---|
| Fail | 0 | 0 | 0 | 0 | 177 | 0 |
| Pass | 48 | 460 | 519 | 259 | 7 | 120 |
Paper 2 third-assessor escalation rate
A third assessor was not required in any of the 894 Paper 2 sessions (0.00 %) in the reporting window 2025-07-01 to 2026-06-18. The two-rater protocol produced a concordant award without escalation in every case. This figure records the escalation rate only and shall not be substituted for a formal inter-rater agreement coefficient; full paired-rater statistics and a Many-Facet Rasch assessor-severity model for Paper 2 are published on the Paper 2 tab.
- Counts are observed values over the reporting window; no inferential or causal interpretation is applied.
- Paper 1 fails that show a recorded Paper 2 level reflect the operational test register as supplied and are not interpreted as misclassification.
- The third-assessor figure is an escalation rate and is not a substitute for an inter-rater agreement coefficient (e.g., Cohen's kappa, many-facet Rasch).
Programme of continuing validation
The ELPAC test operates a continuing validation programme alongside routine operational validity monitoring. Planned activities include formal cross-version equating of Paper 1, differential item functioning analyses once the necessary demographic frame is in place, and the reintroduction of OPE-fluency into the published Paper 2 agreement statistics following the ongoing capture-format revision. A classical predictive criterion study linking ELPAC scores to operational radiotelephony performance is structurally difficult, as the test has no routine access to in-service performance data; the programme therefore additionally pursues complementary feedback mechanisms: structured user feedback and periodic expert review of materials, in support of the extrapolation inference. These activities are part of continuous improvement and do not, on the current evidence, displace the claim and evidence status set out above.
References
Scholarly literature
- Agius, W. (2024). Exploring the features of language performance indicative of air traffic controllers’ workplace language socialisation (Doctoral dissertation). Lancaster University.
- Alderson, J. C. (2009). Air safety, language assessment policy, and policy implementation: Responsibilities and verifiability. Annual Review of Applied Linguistics, 29, 168–187.
- Alderson, J. C. (2010). A survey of aviation English tests. Language Testing, 27(1), 51–72.
- Alderson, J. C., & Wall, D. (1993). Does washback exist? Applied Linguistics, 14(2), 115–129.
- Bachman, L. F. (1990). Fundamental Considerations in Language Testing. Oxford University Press.
- Bachman, L. F., & Palmer, A. S. (1996). Language Testing in Practice. Oxford University Press.
- Bachman, L. F., & Palmer, A. S. (2010). Language Assessment in Practice. Oxford University Press.
- Bailey, K. M. (1996). Working for washback: A review of the washback concept in language testing. Language Testing, 13(3), 257–279.
- Bond, T. G., & Fox, C. M. (2015). Applying the Rasch model: Fundamental measurement in the human sciences (3rd ed.). Routledge.
- Canale, M., & Swain, M. (1980). Theoretical bases of communicative approaches to second language teaching and testing. Applied Linguistics, 1(1), 1–47.
- Chapelle, C. A. (2020). Argument-Based Validation in Testing and Assessment. SAGE.
- Chapelle, C. A., & Voss, E. (Eds.) (2021). Validity Argument in Language Testing: Case Studies of Validation Research. Cambridge University Press.
- Emery, H. J. (2014). Developments in LSP testing 30 years on? The case of aviation English. Language Assessment Quarterly, 11(2), 198–215.
- Estival, D., Farris, C., & Molesworth, B. (2016). Aviation English: A lingua franca for pilots and air traffic controllers. Routledge.
- Green, A. (2013). Exploring language assessment and testing. Routledge.
- Hughes, A. (2003). Testing for language teachers (2nd ed.). Cambridge University Press.
- Hymes, D. H. (1972). On communicative competence. In J. B. Pride & J. Holmes (Eds.), Sociolinguistics (pp. 269–293). Penguin.
- Kane, M. T. (1992). An argument-based approach to validity. Psychological Bulletin, 112(3), 527–535.
- Kane, M. T. (2006). Validation. In R. L. Brennan (Ed.), Educational Measurement (4th ed., pp. 17–64). American Council on Education / Praeger.
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.
- Kim, H. (2018). What constitutes professional communication in aviation: Is language proficiency enough for testing purposes? Language Testing, 35(3), 403–426.
- Kim, H., & Elder, C. (2009). Understanding aviation English as a lingua franca. Australian Review of Applied Linguistics, 32(3), 23.1–23.17.
- Kim, H., & Elder, C. (2015). Interrogating the construct of aviation English: Feedback from test takers in Korea. Language Testing, 32(2), 129–149.
- Knoch, U. (2009). Diagnostic assessment of writing: A comparison of two rating scales. Language Testing, 26(2), 275–304.
- Knoch, U. (2014). Using subject specialists to validate an ESP rating scale: The case of the ICAO rating scale. English for Specific Purposes, 33, 77–86.
- Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational Measurement (3rd ed., pp. 13–103). American Council on Education / Macmillan.
- Messick, S. (1996). Validity and washback in language testing. Language Testing, 13(3), 241–256.
- Mislevy, R. J., Steinberg, L. S., & Almond, R. G. (2003). Focus article: On the structure of educational assessments. Measurement: Interdisciplinary Research and Perspectives, 1(1), 3–62.
- Monteiro, A. L. T. (2019). Reconsidering the measurement of proficiency in pilot and air traffic controller radiotelephony communication: From construct definition to task design (Doctoral dissertation). Carleton University.
- Pill, J., & McNamara, T. (2016). How much is enough? Involving occupational experts in setting standards on a specific-purpose language test for health professionals. Language Testing, 33(2), 217–234.
- Read, J., & Knoch, U. (2009). Clearing the air: Applied linguistic perspectives on aviation communication. Australian Review of Applied Linguistics, 32(3), 21.1–21.11.
Standards and regulations
- ICAO Doc 9835 — Manual on the Implementation of ICAO Language Proficiency Requirements (2nd ed., 2010); in particular §4.6 and Attachment A (Holistic Descriptors and ICAO Language Proficiency Rating Scale).
- ICAO Annex 1 — Personnel Licensing.
- ICAO Annex 10, Volume II — Aeronautical Telecommunications (Communication Procedures).
- Commission Regulation (EU) 2015/340 — Technical requirements and administrative procedures relating to air traffic controllers' licences and certificates (ATCO.B.030).
- Commission Regulation (EU) No 1178/2011, Part-FCL (FCL.055) — Flight crew licensing, language proficiency.