About this page. This validity documentation is maintained by the ELPAC team at EUROCONTROL as a living methodological record. It is not an ICAO endorsement and does not by itself constitute a regulatory determination. National civil aviation authorities remain responsible for licensing decisions and for determining whether ELPAC results satisfy the ICAO language-proficiency requirements in their jurisdiction.

Last reviewed: 18 June 2026.

Validity argument

Validity argument for the ELPAC assessment

This document sets out the interpretation and use argument supporting the ELPAC assessment as a basis for ICAO language proficiency licensing decisions, in accordance with ICAO Annex 1, ICAO Doc 9835, Commission Regulation (EU) 2015/340 and Commission Regulation (EU) No 1178/2011. The qualitative argument is presented below; quantitative psychometric backing (Classical Test Theory, Rasch and Many-Facet Rasch analyses) is published in the operational validity monitoring section.

Purpose

This document records the interpretation and use argument (IUA) supporting the ELPAC assessment, structured in accordance with the framework of Kane (1992, 2006, 2013) and the application to second-language testing set out in Chapelle, Enright and Jamieson (2008). The argument is decomposed into six inferences: domain description, evaluation, generalisation, explanation, extrapolation and utilisation.

Scope

The argument applies to the use of ELPAC scores by competent authorities and approved organisations for the purpose of issuing, endorsing or renewing the ICAO language proficiency endorsement under ICAO Annex 1, Commission Regulation (EU) 2015/340 (ATCO.B.030) and Commission Regulation (EU) No 1178/2011 (FCL.055).

Status of evidence

The qualitative argument set out in the sections that follow is supported by quantitative operational validity monitoring evidence published on this site, comprising: Classical Test Theory and Rasch (IRT) psychometric evidence for Paper 1 across seven operational forms; a Many-Facet Rasch examiner-severity model and rater-agreement statistics for Paper 2; and operational outcome statistics for the 2026 reporting cohort (N = 1,590; reporting period 2026-01-05 to 2026-08-31). Further studies identified in the programme of continuing validation are planned and do not affect the determination set out below.

Construct

The ICAO construct and the ELPAC test construct

A validity argument is anchored in an explicit statement of what the test is intended to measure. This section distinguishes two related but non-identical constructs: the regulatory construct defined by ICAO, and the test construct that ELPAC operationalises. ELPAC does not define its own construct — it inherits ICAO's and renders it assessable.

1. The ICAO construct (regulatory)

ICAO Annex 1 (Personnel Licensing), §1.2.9 and Appendix 1, together with ICAO Doc 9835 (Manual on the Implementation of ICAO Language Proficiency Requirements), define the regulated construct as the operational language proficiency required for safe radiotelephony communication by licensed air traffic controllers and flight crew. ICAO defines this construct in two complementary ways:

  • Holistic descriptors (Doc 9835 §4.6 / Attachment A) — the qualitative definition of what an operationally proficient speaker can do: communicate in voice-only and face-to-face situations; handle work-related topics with accuracy and clarity; use communicative strategies to recognise and resolve misunderstanding; handle a complication or unexpected turn of events; and remain intelligible to the aeronautical community.
  • ICAO Rating Scale — six descriptors (Pronunciation, Structure, Vocabulary, Fluency, Comprehension, Interactions) across six levels, with Operational Level 4 as the licensing threshold.

Two properties of the ICAO construct are decisive for what follows. It is communicative and operational, not general English; and it is bounded to the ICAO radiotelephony target language use (TLU) domain — it does not extend to cabin-crew, ground-handling or engineering English, and it does not evaluate the technical or procedural correctness of the underlying operational decision.

2. The ELPAC test construct

The construct ELPAC sets out to measure is operational communicative competence in aviation radiotelephony: the sub-construct of the ICAO regulatory construct that is realised in a candidate's observable language behaviour under operationally relevant conditions. "Communicative competence" is taken in the tradition that runs from Hymes (1972) and Canale & Swain (1980) through Bachman (1990) and Bachman & Palmer (2010), restricted to the ICAO TLU domain. Substantively, the construct is defined by the ICAO holistic descriptors; gradably, it is operationalised through the six ICAO Rating Scale descriptors. Throughout the argument below, "language proficiency" is to be read in this operational, communicative sense.

In construct
  • Pronunciation, Structure, Vocabulary, Fluency, Comprehension, Interactions — as features of communicative behaviour
  • Handling of routine and non-routine work-related communication
  • Communicative strategies to check, confirm, clarify and resolve misunderstanding
  • Intelligibility to the aeronautical community
Out of construct
  • General English proficiency beyond the ICAO TLU domain
  • Technical or procedural correctness of ATC or piloting decisions
  • Personality, confidence or interpersonal style
  • Accent conformity beyond intelligibility to the aeronautical community

3. Construct alignment

The crosswalk below records, element by element, how the ICAO construct is operationalised in ELPAC and where in the argument each element is discharged.

Crosswalk from each ICAO construct element to its ELPAC operationalisation and the inference that discharges it.
ICAO construct elementELPAC operationalisationInference
Holistic descriptors — voice-only + face-to-face; routine + non-routine; communicative strategiesPaper 1 authentic radiotelephony audio; Paper 2 three-task progression (routine → complication → extended interaction)1, 2
Rating Scale — six descriptors, six levels, Level 4 thresholdIndependent double rating by ELE + OPE against the six descriptors; final award is the lowest sub-rating2, 4
Operational relevance — intelligibility to the aeronautical communityOPE (serving or former ATCO or pilot) on every Paper 2 panel; separate ATC and Pilot versions1, 2, 4
Level 4 as the regulated licensing cutPaper 1 as a Level-4-targeted screen; Paper 2 lowest-of-six composite award3, 6
In short

What this argument concludes

  • ELPAC measures the construct ICAO regulates — operational communicative competence in aviation radiotelephony — and not general English.
  • The two papers together elicit rateable performance on all six ICAO Rating Scale descriptors, and Paper 2 is double-rated by an English Language Expert and an Operational Expert with a defined third-examiner route where they cannot reconcile.
  • On the qualitative argument below and the published psychometric evidence, ELPAC is determined to provide an appropriate technical basis for ICAO language proficiency licensing decisions.
  • The two weakest links are extrapolation to in-service radiotelephony performance and the consequences of the licensing decision. Both rest on test and procedural design rather than a criterion study, for the structural reasons set out in Inference 5 and the continuing validation programme.
At a glance

The argument in one table

The six inferences below form the spine of the argument. Each row links to the full section. Strong means published quantitative backing on this site; partial means design and procedural backing with a planned quantitative study; design-only means the warrant is supported by task and construct design rather than by a single criterion study, for the structural reasons set out in the relevant section.

The six inferences of the ELPAC validity argument, with a one-line warrant and the strength of backing for each.
InferenceOne-line warrantBacking
1. Domain descriptionThe TLU domain is ICAO radiotelephony, as specified in Doc 9835.strong
2. EvaluationPaper 1 and Paper 2 jointly elicit rateable performance on the six ICAO descriptors.strong
3. GeneralisationThe awarded level generalises across forms, sessions, examiners and centres.partial
4. ExplanationScoring procedures constrain examiner and operational-expertise variance.strong
5. ExtrapolationThe score supports an inference to operational radiotelephony performance.design-only
6. UtilisationLicensing decisions made from the score are appropriate and their consequences are bounded.design-only
Complementary view — Green (2013)

Crosswalk to the four plain questions

Green (2013, ch. 5) reduces the full Kane chain to four plain questions a regulator can put to a test. Those questions map onto the six inferences set out on this page as follows; no separate evidence base is maintained for the Green view.

Mapping of the four Green (2013) validity questions onto the six Kane inferences.
Green (2013) questionKane inference
Scoring — is the score accurate?4. Explanation
Generalisation — would a different form/day/examiner give a similar score?3. Generalisation
Extrapolation — does doing well on the test mean doing well on the radio?5. Extrapolation
Utilisation — are the resulting decisions better than the alternatives?6. Utilisation
Inference 1

Domain description

Backing: strong

Claim

ELPAC defines its target language use (TLU) domain — as distinct from the construct set out in the Construct section above — as the operational English of air traffic radiotelephony and flight-deck communication, as specified by ICAO.
Show the evidence on ELPAC

The test is anchored in ICAO Doc 9835 (Manual on the Implementation of ICAO Language Proficiency Requirements) and in the ICAO Language Proficiency Rating Scale. The Rating Scale characterises operational language along six descriptors — Pronunciation, Structure, Vocabulary, Fluency, Comprehension and Interactions — across six levels. Separate ATC and Pilot versions are designed so that the domain sampled in each version matches the candidate's role, covering standard phraseology and plain aviation English in routine and non-routine situations.

Alongside the Rating Scale, ICAO Doc 9835 (§4.6 and Attachment A) sets out holistic descriptors that define, qualitatively, what an operationally proficient speaker can do. ELPAC takes the holistic descriptors as the substantive description of the target domain. They straddle domain and construct: they describe the TLU domain here (Inference 1) and inform what is rated in Paper 2 (Inference 2). An operationally proficient speaker is one who can:

  • communicate effectively in voice-only (telephone / radiotelephony) and in face-to-face situations;
  • communicate on common, concrete and work-related topics with accuracy and clarity;
  • use appropriate communicative strategies to exchange messages and to recognise and resolve misunderstandings (e.g. to check, confirm or clarify information) in a general or work-related context;
  • handle successfully and with relative ease the linguistic challenges presented by a complication or unexpected turn of events that occurs within the context of a routine work situation or communicative task with which they are otherwise familiar; and
  • use a dialect or accent that is intelligible to the aeronautical community.

Together, the holistic descriptors and the six-descriptor Rating Scale specify the operational language use domain: the Rating Scale supplies the gradable criteria, and the holistic descriptors anchor those criteria in operationally meaningful behaviour.

Limitations and rebuttals

The ICAO construct is acknowledged in the applied linguistics literature to under-specify certain aspects of the target language use domain — notably the boundary between standard phraseology and plain English, and the treatment of listener factors (Alderson 2009; Emery 2014). These are properties of the ICAO framework rather than of the ELPAC assessment. The domain scope is also bounded by ICAO Annex 1 and Annex 10 Volume II, and accordingly does not extend to cabin-crew, ground-handling or engineering English.
Inference 2

Evaluation (construct representation)

Backing: strong

Claim

The tasks in Paper 1 and Paper 2 jointly elicit rateable performance on the ICAO construct of operational communicative competence, as set out in the Construct section above, under operationally relevant conditions. The construct is operationalised through the six ICAO Rating Scale descriptors, read against the ICAO holistic descriptors.
Show the evidence on ELPAC
Paper 1 (Listening) uses authentic radiotelephony recordings and longer aviation audio. It samples Comprehension across accents, registers and routine / non-routine content, and so addresses the holistic expectation of intelligibility across the aeronautical community. Paper 2 (Speaking) is structured around three tasks that progressively shift the candidate from routine description, to handling an unusual situation, to extended interaction including hypothesis and problem-solving. This progression maps directly onto the holistic descriptors: routine work-related communication in the early tasks; handling a complication or unexpected turn of events; and use of communicative strategies to recognise and resolve misunderstandings under processing load. Pronunciation, Structure, Vocabulary, Fluency and Interactions are accordingly observed as features of communicative behaviour, not as isolated linguistic traits on rehearsed material.

Limitations and rebuttals

Construct representation is operationalised through the six ICAO descriptors, read against the holistic descriptors. Construct-representation considerations of the kind identified by Messick (1989) and Read & Knoch (2009) apply to any descriptor-based scale of operational communicative competence. The Paper 2 speaking sample, of approximately 20–25 minutes, is bounded by design; rare high-stakes operational events are represented by analogous tasks. The ICAO scale assesses communicative effectiveness and does not score the technical correctness of operational decisions.
Inference 3

Generalisation

Backing: partial

Claim

An awarded ELPAC level generalises across forms, sessions, examiners and test centres: different test occasions would produce a comparable level for the same candidate.
Show the evidence on ELPAC
Items are drawn from calibrated banks, and equivalent forms are constructed to comparable difficulty. Examiners are centrally accredited and periodically re-standardised. Test centres operate under a common technical and procedural specification overseen by EUROCONTROL and its test development partner ZHAW (Zurich University of Applied Sciences), and the same digital test system is used at every accredited centre, which standardises audio quality, timing and item presentation. The reported outcome is a discrete ordinal level (in practice 3, 4, 5 or 6); within each band, coarse categorisation tends to improve classification consistency, because small differences in underlying performance do not translate into different awarded levels.

Limitations and rebuttals

Generalisation is conditional on the continued operation of the quality-assurance processes described above. Local conditions (microphone quality, ambient noise, examiner fatigue) introduce residual variance. At cut-scores between bands, coarse categorisation increases the consequences of small measurement differences for candidates whose true level lies near a boundary.

The minimum-across-sub-ratings convention — mandated by the ICAO rating scale criteria and discussed in detail under Inference 6 — means that the reliability of the composite final level is, in general, lower than that of any single sub-rating (Knoch 2009). Procedural quality assurance constrains rater and form variance at the sub-rating level; the reliability of the minimum-based composite itself is monitored separately, through the Many-Facet Rasch examiner-severity model published in the Paper 2 operational validity monitoring section.

Inference 4

Explanation (scoring and response process)

Backing: strong

Claim

ELPAC's scoring procedures are designed to constrain examiner and operational-expertise variance, so that the awarded level reflects features of the candidate's language behaviour rather than features of who happened to rate the session.
Show the evidence on ELPAC
Paper 1 is computer-marked, which removes rater variance from the listening component; item-level and form-level sampling error remain. Paper 2 is rated live by two accredited examiners — an Operational Expert (OPE), who is a serving or former air traffic controller or pilot, and an English Language Expert (ELE). The two examiners hold complementary perspectives: the OPE brings operational judgement on comprehension and interaction in a radiotelephony context, and the ELE brings linguistic judgement on pronunciation, structure, vocabulary and fluency. Each examiner rates independently against the six-descriptor ICAO Rating Scale, read in light of the holistic descriptors. The two independent ratings are then compared, and the overall ICAO level awarded is the lowest of the six sub-ratings, in accordance with the ICAO rating scale criteria. When the two examiners cannot reconcile, the case is resolved through a documented dispute-resolution mechanism (a third examiner), rather than by on-the-spot consensus. Independent double rating with reconciliation is methodologically stronger than consensus scoring, because it preserves the evidence of rater disagreement and forces it through an explicit resolution step. Sessions are recorded, which supports adjudication, audit and post-hoc re-rating as part of routine quality assurance.

Limitations and rebuttals

Independent double rating with reconciliation is more robust than single rating, but it does not eliminate residual examiner, halo or severity effects. Those effects are quantified, and monitored on a continuing basis, by the Many-Facet Rasch examiner-severity model and the rater-agreement statistics published in the Paper 2 operational validity monitoring section.
Inference 5

Extrapolation to operational performance

Backing: design-only

Claim

ELPAC scores support an inference about the candidate's English-language performance in operational air traffic control and flight-deck communication. The warrant rests on two foundations: the operational provenance of the test materials, and direct alignment with the ICAO Language Proficiency Rating Scale used for licensing.
Show the evidence on ELPAC
ELPAC reports against the same ICAO Rating Scale applied by competent authorities for the language proficiency endorsement, so test results map onto the operational scale without further translation. Test materials are drawn from the target language use domain — authentic radiotelephony recordings, non-routine situations, and interaction under time pressure — rather than from general English content. In the terms of Bachman & Palmer (1996, 2010), this task authenticity supports the extrapolation inference. A criterion study against operational radiotelephony performance is part of the programme of continuing validation.

Limitations and rebuttals

ELPAC is administered independently of the candidate's operational employer. In consequence, the programme has no routine access to a candidate's subsequent radiotelephony performance on the line, and individual test takers cannot be systematically followed into live ATC or flight operations. A conventional predictive criterion study against operational radiotelephony performance is therefore structurally difficult to mount; the absence of such studies is a recognised feature of the aviation-English testing field (Alderson 2010; Knoch 2014). The extrapolation inference is accordingly supported indirectly, rather than by a single criterion study, through task authenticity, construct alignment with the ICAO Rating Scale, and complementary feedback mechanisms: structured stakeholder and user feedback; periodic expert review of materials; and, where lawfully available and de-identified, signals from communication-related occurrence reporting and downstream training outcomes. Operational communication also includes workload, fatigue and team factors that fall outside the scope of any language test. ELPAC accordingly complements, and does not substitute for, ongoing operational supervision and unit competence assessment.
Inference 6

Utilisation (decision and consequences)

Backing: design-only

Claim

The decisions made from ELPAC scores — issuing, endorsing or renewing a language proficiency rating — are appropriate decisions, and the foreseeable negative consequences of misclassification are bounded by the surrounding regulatory regime.
Show the evidence on ELPAC
Decisions are made by national authorities under ICAO Annex 1 and Annex 10 Volume II, and, in EASA Member States, additionally under Regulation (EU) 2015/340 (ATCO.B.030) for air traffic controllers and Commission Regulation (EU) No 1178/2011 (Part-FCL, FCL.055) for flight crew. Under EU 2015/340, the language proficiency endorsement is valid for 3 years at Level 4, 6 years at Level 5 and 9 years at Level 6. For flight crew under Part-FCL (FCL.055), Level 4 is valid for 4 years and Level 5 for 6 years from the date of assessment; Level 6 is, in principle, not subject to re-assessment, although some Member States impose periodic checks. The overall ICAO level is determined by the lowest of the six skill ratings, in accordance with the ICAO rating scale criteria set out in ICAO Doc 9835. This minimum-across-sub-ratings convention is a deliberate standard-setting choice that prioritises safety over score economy. Following Alderson (2009) and Pill & McNamara (2016) on standard-setting in high-stakes LSP testing, it is best understood as a policy decision about consequences, rather than as a property of the measurement model. National authorities operate re-sit and appeal procedures, the detail of which varies by State, and Paper 2 recordings support those processes. See the FAQ for the validity table by role and level.

Limitations and rebuttals

The minimum-across-sub-ratings convention, and the re-assessment intervals set in EU 2015/340 and Part-FCL FCL.055 are deliberate policy-level provisions of the regulatory framework, rather than properties of the measurement model. Their effect is to bias classification outcomes towards safety. High-stakes single decisions, in particular the award of Level 6, are subject to the independent double-rating regime described under Inference 4.

Overall determination

On the basis of the qualitative interpretation and use argument set out above and the quantitative operational validity monitoring evidence published on this site, the ELPAC assessment is determined to provide an appropriate technical basis for licensing decisions made by competent authorities under ICAO Annex 1, Commission Regulation (EU) 2015/340 and Commission Regulation (EU) No 1178/2011. The strongest warrants are the anchoring of the construct in ICAO Doc 9835, the double-rated independent assessment of Paper 2 with defined reconciliation, and the centralised quality assurance applied to forms, examiners and accredited test centres. Reliability of Paper 1 across operational forms, examiner severity for Paper 2, and the third-assessor escalation rate for the 2026 reporting cohort are within the ranges expected for a high-stakes language proficiency assessment of this design.

Operational evidence — reporting period 2026-01-05 to 2026-08-31

Operational outcomes, 2026 reporting period (N = 1,590)

The figures below are derived from the de-identified operational test register for the period stated. They are reported as observed counts and rates only, in support of the consequence and generalisation inferences set out above. These values are descriptive only; they do not constitute predictive-validity evidence against operational radiotelephony performance.

Paper 1 — overall
88.9 %
1,413 pass / 177 fail of 1,590
Paper 1 — ATC stream
84.0 %
653 pass / 124 fail of 777
Paper 1 — Pilot stream
93.5 %
760 pass / 53 fail of 813

Final ICAO level, candidates with a Paper 1 pass

Distribution of the final score — the lowest of the six sub-ratings, in accordance with the ICAO rating scale criteria — for candidates progressing past Paper 1 in the reporting period.

Distribution of the final ICAO level among candidates who passed Paper 1.
Final resultCandidatesShare
Level 3483.4 %
ICAO Level 446032.6 %
ICAO Level 551936.7 %
ICAO Level 625918.3 %
Fail70.5 %
Pending / not recorded1208.5 %

Paper 1 result × Final score (observed counts)

Descriptive cross-tabulation of the Paper 1 screening decision against the final ICAO level recorded for the same candidate in the reporting period. The table records what occurred; it is not a misclassification matrix and does not characterise the accuracy of the screening decision.

Observed counts of Paper 1 result against final ICAO level.
Paper 1Level 3ICAO Level 4ICAO Level 5ICAO Level 6FailNot recorded
Fail00001770
Pass484605192597120

Paper 2 third-assessor escalation rate

A third assessor was not invoked in any of the 894 Paper 2 sessions (0.00 %) in the reporting window 2025-07-01 to 2026-06-18 — the two-rater protocol produced a concordant award without escalation in every case. This figure records the escalation rate only and shall not be substituted for a formal inter-rater agreement coefficient; full paired-rater statistics and a Many-Facet Rasch examiner-severity model for Paper 2 are published on the Paper 2 tab.

Scope and limitations of this dataset
  • Counts are observed values over the reporting window; no inferential or causal interpretation is applied.
  • Paper 1 fails that show a recorded Paper 2 level reflect the operational test register as supplied and are not interpreted as misclassification.
  • The third-assessor figure is an escalation rate and is not a substitute for an inter-rater agreement coefficient (e.g., Cohen's kappa, many-facet Rasch).

Programme of continuing validation

The ELPAC programme operates a continuing validation programme alongside routine operational validity monitoring. Planned activities include formal cross-form equating of Paper 1, differential item functioning analyses once the necessary demographic frame is in place, and the reintroduction of OPE-fluency into the published Paper 2 agreement statistics following the ongoing capture-format revision. A classical predictive criterion study linking ELPAC scores to operational radiotelephony performance is structurally difficult, as the test is administered independently of the candidate's employer and the programme has no routine access to in-service performance data; the programme therefore additionally pursues complementary feedback mechanisms — structured stakeholder and user feedback, periodic expert review of materials, and, where lawfully available and de-identified, signals from communication- related occurrence reporting and downstream training outcomes — in support of the extrapolation inference. These activities are part of continuous improvement and do not, on the current evidence, displace the determination of fitness for purpose set out above.

References

Scholarly literature

  • Alderson, J. C. (2009). Air safety, language assessment policy, and policy implementation: Responsibilities and verifiability. Annual Review of Applied Linguistics, 29, 168–187.
  • Alderson, J. C. (2010). A survey of aviation English tests. Language Testing, 27(1), 51–72.
  • Alderson, J. C., & Wall, D. (1993). Does washback exist? Applied Linguistics, 14(2), 115–129.
  • Bachman, L. F. (1990). Fundamental Considerations in Language Testing. Oxford University Press.
  • Bachman, L. F., & Palmer, A. S. (1996). Language Testing in Practice. Oxford University Press.
  • Bachman, L. F., & Palmer, A. S. (2010). Language Assessment in Practice. Oxford University Press.
  • Bailey, K. M. (1996). Working for washback: A review of the washback concept in language testing. Language Testing, 13(3), 257–279.
  • Bond, T. G., & Fox, C. M. (2015). Applying the Rasch model: Fundamental measurement in the human sciences (3rd ed.). Routledge.
  • Canale, M., & Swain, M. (1980). Theoretical bases of communicative approaches to second language teaching and testing. Applied Linguistics, 1(1), 1–47.
  • Chapelle, C. A., Enright, M. K., & Jamieson, J. M. (Eds.) (2008). Building a Validity Argument for the Test of English as a Foreign Language. Routledge.
  • Emery, H. J. (2014). Developments in LSP testing 30 years on? The case of aviation English. Language Assessment Quarterly, 11(2), 198–215.
  • Estival, D., Farris, C., & Molesworth, B. (2016). Aviation English: A lingua franca for pilots and air traffic controllers. Routledge.
  • Green, A. (2013). Exploring language assessment and testing. Routledge.
  • Hughes, A. (2003). Testing for language teachers (2nd ed.). Cambridge University Press.
  • Hymes, D. H. (1972). On communicative competence. In J. B. Pride & J. Holmes (Eds.), Sociolinguistics (pp. 269–293). Penguin.
  • Kane, M. T. (1992). An argument-based approach to validity. Psychological Bulletin, 112(3), 527–535.
  • Kane, M. T. (2006). Validation. In R. L. Brennan (Ed.), Educational Measurement (4th ed., pp. 17–64). American Council on Education / Praeger.
  • Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.
  • Kim, H., & Elder, C. (2009). Understanding aviation English as a lingua franca. Australian Review of Applied Linguistics, 32(3), 23.1–23.17.
  • Knoch, U. (2009). Diagnostic assessment of writing: A comparison of two rating scales. Language Testing, 26(2), 275–304.
  • Knoch, U. (2014). Using subject specialists to validate an ESP rating scale: The case of the ICAO rating scale. English for Specific Purposes, 33, 77–86.
  • Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational Measurement (3rd ed., pp. 13–103). American Council on Education / Macmillan.
  • Messick, S. (1996). Validity and washback in language testing. Language Testing, 13(3), 241–256.
  • Mislevy, R. J., Steinberg, L. S., & Almond, R. G. (2003). Focus article: On the structure of educational assessments. Measurement: Interdisciplinary Research and Perspectives, 1(1), 3–62.
  • Pill, J., & McNamara, T. (2016). How much is enough? Involving occupational experts in setting standards on a specific-purpose language test for health professionals. Language Testing, 33(2), 217–234.
  • Read, J., & Knoch, U. (2009). Clearing the air: Applied linguistic perspectives on aviation communication. Australian Review of Applied Linguistics, 32(3), 21.1–21.11.

Standards and regulations

  • ICAO Doc 9835 — Manual on the Implementation of ICAO Language Proficiency Requirements (2nd ed., 2010); in particular §4.6 and Attachment A (Holistic Descriptors and ICAO Language Proficiency Rating Scale).
  • ICAO Annex 1 — Personnel Licensing.
  • ICAO Annex 10, Volume II — Aeronautical Telecommunications (Communication Procedures).
  • Commission Regulation (EU) 2015/340 — Technical requirements and administrative procedures relating to air traffic controllers' licences and certificates (ATCO.B.030).
  • Commission Regulation (EU) No 1178/2011, Part-FCL (FCL.055) — Flight crew licensing, language proficiency.