Skip to main content

About this page. This validity documentation is maintained by the ELPAC team at EUROCONTROL as a living methodological record. It is not an ICAO endorsement and does not by itself constitute a regulatory determination. National civil aviation authorities remain responsible for licensing decisions and for determining whether ELPAC results satisfy the ICAO language-proficiency requirements in their jurisdiction.

Last reviewed: 18 June 2026.

Test design criteria

ELPAC and ICAO Doc 10197

ICAO Doc 10197, Test Design Guidelines: Handbook on the Design of Tests for the ICAO Language Proficiency Requirements (First Edition, 2024), sets out eight criteria that licensing authorities can use when considering a language proficiency test for approval, and that test developers can use when reviewing their own instrument. It is guidance on test design; it is not a certification scheme and conformity with it is not awarded by ICAO.

The overview below is the ELPAC team's own assessment of the ELPAC test instrument against those eight criteria. It is provided so that authorities and user organisations can see, criterion by criterion, what in ELPAC's design addresses each point and where the supporting evidence on this site can be found. It does not represent a determination by ICAO.

Doc 10197 applies the criteria in a fixed order. Criteria 1, 2 and 3 are treated as essential for licensing purposes: a test that fails any of them may not be suitable for assessing controllers or pilots for licensing at all. Later criteria are only reached once the earlier ones are satisfied, and a failure at one point also counts against the criteria that follow it. The sections below are therefore presented in ICAO's own order.

Before the criteria: evaluation requirements

Doc 10197's evaluation checklist opens with two prerequisites: a full set of test specifications, and a full sample test including interlocutor scripts, candidate instructions, input material and audio.

ELPAC's specifications, task functions, delivery format and rating arrangements are set out in the ELPAC Handbook and the ELPAC Assessment Scheme, both available under documents. A complete sample test is published for each role, with the real interface, instructions and audio. Paper 2 interlocutor scripts and other live test material are not published, because they are live test content; they are handled through the ELPAC team under the licensing arrangements.

At a glance

The eight criteria in one table

Met by design means the criterion is addressed by a structural feature of the instrument. Largely met means the design addresses it and quantitative monitoring supports it, with the qualification noted in the relevant section.

The eight ICAO Doc 10197 test design criteria, with a one-line summary of how ELPAC addresses each and the assessed status.
CriterionHow ELPAC addresses itStatus
1. Radiotelephony communication contextsBoth papers are built on radiotelephony material, combining ICAO phraseology with plain language.Met by design
2. Separate tests for pilots and controllersELPAC ATC and ELPAC Pilot are separate instruments with role-specific audio, tasks and examiner pairing.Met by design
3. Listening comprehension tasks separate from speaking tasksPaper 1 is a dedicated, separately scored listening comprehension paper of 45 items.Met by design
4. Distinct sections and a range of task typesTwo papers, multiple parts within Paper 1 and three differentiated Paper 2 tasks.Met by design
5. Extended and interactive communicationPaper 2 is a live face-to-face assessment with co-constructed dialogue and unscripted follow-up questions.Met by design
6. Differentiation between ICAO language proficiency levelsTasks are graded in demand across the whole ICAO range, and level differentiation is monitored quantitatively.Largely met
7. Real-world contextsThe speaking paper is built on operational role play, with a debriefing task in a recognisable workplace context.Met by design
8. Test bankMultiple live versions per role, built to common specifications and monitored for comparability.Largely met
Criterion 1

Radiotelephony communication contexts

Met by design

The criterion, as published in Doc 10197

Test instruments need to include appropriate tasks that directly assess how test-takers use language in radiotelephony communication contexts.

What the ICAO checklist asks

  • Does the test instrument include tasks that directly assess how test-takers communicate in real-world radio communication situations?
  • Do the tasks require test-takers to communicate in non-routine or unexpected situations using radiotelephony?
  • Does the test instrument include separate tasks to assess both listening and speaking skills in radiotelephony communication contexts?
  • Are the tasks specific to the test-taker's professional role with regard to how they communicate over the radio as controllers or pilots?
Show how ELPAC addresses this

Paper 1 presents recordings of radiotelephony exchanges and longer operational audio (briefings, ATIS-style messages, incident reports), with comprehension items keyed to prescribed correct responses. Paper 2 Task 1a places the candidate in their own operational role while the Operational Expert plays the counterpart from a script, mixing routine messages that call for standard phraseology with developing non-routine events that require plain English.

Listening and speaking in radiotelephony contexts are assessed in separate tasks, in separate papers, and the material for each role is drawn from that role's own radio work.

Criterion 2

Separate tests for pilots and controllers

Met by design

The criterion, as published in Doc 10197

Separate test instruments need to be designed for pilots and controllers.

What the ICAO checklist asks

  • Are the tasks specific to the communication needs of pilots or controllers?
  • Do the tasks assess the specific language pilot or controller test-takers use, without assessing the specific speaking or listening needs of the other role?
  • Is the context and content of the input matched to the pilot or ATCO test-taker needs?
Show how ELPAC addresses this

ELPAC ATC and ELPAC Pilot are two separate instruments. They assess the same construct, but the audio, charts, scenarios and role-play counterparts are drawn from the candidate's own environment: air traffic control radiotelephony for controllers, flight deck communication for pilots. Paper 1 is delivered in four parts for the ATC version and five parts for the Pilot version. A controller is not asked to handle flight deck tasks, and a pilot is not asked to work traffic.

Role matching extends to the rating panel: the Operational Expert is an operational or former air traffic controller or pilot, matched to the candidate's role.

Criterion 3

Listening comprehension tasks separate from speaking tasks

Met by design

The criterion, as published in Doc 10197

Test instruments need to contain tasks dedicated to assessing listening comprehension, separate from tasks designed to assess speaking performance.

What the ICAO checklist asks

  • Does the test instrument contain tasks that are designed to assess only listening comprehension skills and not speaking skills?
  • Is there a sufficient quantity and variety of listening tasks, and a variety of recordings and task types?
  • Are listening tasks included that assess comprehension skills in radiotelephony contexts, and tasks that cater specifically to the listening contexts of either pilots or controllers?
  • Do the listening comprehension tasks include clear instructions and a presentation of items and audio such that results reflect listening comprehension rather than reading, memory or multi-tasking skills?
Show how ELPAC addresses this

Paper 1 assesses listening comprehension on its own, with 45 computer-marked items across four (ATC) or five (Pilot) parts and a documented item-level psychometric record. No speaking is required, so comprehension is measured directly rather than being read off the interaction in Paper 2.

The recordings vary in type and length: short radiotelephony transmissions in one set of parts, and longer operational audio such as briefings, ATIS-style messages and incident reports in others. Comprehension observed inside the Paper 2 interaction functions as additional supporting evidence within the Comprehension sub-rating, not as the primary measure.

Qualification

On ICAO's last checklist point, Paper 1 is designed to keep non-listening demands low: each part is introduced before its audio begins, items are short, each recording plays once and the response format is multiple-choice. ELPAC has not published a dedicated study isolating reading load or memory load as sources of item difficulty; that is listed below as further evidence.

Criterion 4

Distinct sections and a range of task types

Met by design

The criterion, as published in Doc 10197

Test instruments need to comprise distinct sections with a range of appropriate task types.

What the ICAO checklist asks

  • Does each section of both the speaking and listening test comprise different parts that assess a range of language and communication skills?
  • Does each part of the test make use of different types of content (radiotelephony in routine and non-routine situations, plain language in work-related aviation contexts, and so on)?
  • Are appropriate task types used throughout the test?
  • Does each part of the speaking test use different task types to assess different skills in different contexts?
Show how ELPAC addresses this

The instrument has two clearly separated papers. Paper 1 moves from recognition items on short radiotelephony transmissions to multiple-choice comprehension on longer operational audio, across four (ATC) or five (Pilot) parts. Paper 2 has three tasks with different demands: an operational role play with an informal debriefing to a supervisor (Task 1a and 1b), a task on unusual or non-routine situations (Task 2), and extended interaction involving hypothesising and problem-solving (Task 3).

The content types differ across parts as well as the task types: standard phraseology in routine traffic, plain language under a developing complication, and plain language in a work-related but non-radiotelephony context (the debriefing). A candidate cannot reach a high overall level on the strength of one narrow skill.

Criterion 5

Extended and interactive communication

Met by design

The criterion, as published in Doc 10197

Test instruments need to include tasks that allow test-takers to engage and participate in interactive and extended co-constructed dialogues.

What the ICAO checklist asks

  • Does at least one part of the test allow the test-taker to participate in an interactive exchange with an interlocutor?
  • Do interlocutor frames allow the interlocutor opportunities to jointly participate in a dialogue with the test-taker?
  • Are role-play tasks constructed so that the test-taker participates in a jointly constructed exchange that evolves throughout the role-play?
  • Does the instrument contain a combination of interactive tasks that collectively produce co-constructed dialogue of sufficient length and elicit a sufficient range of language for rating?
Show how ELPAC addresses this

Paper 2 is conducted face to face, not through recorded prompts. The Operational Expert works from an interlocutor script that is written to be answered rather than merely delivered, and the traffic situation develops in response to what the candidate does. The supervisor debriefing in Task 1b asks targeted follow-up questions whenever detail is missing, so the exchange builds on what the candidate has already said.

Task 3 extends the interaction into hypothesising and problem-solving, which produces longer stretches of speech than a question-and-answer format would. Across the three tasks the dialogue is long enough and varied enough to rate all six sub-ratings, which is what makes the Interaction sub-rating of the ICAO rating scale directly observable.

Criterion 6

Differentiation between ICAO language proficiency levels

Largely met

The criterion, as published in Doc 10197

Test instruments need to include tasks and items that allow the assessment to differentiate between ICAO language proficiency levels.

What the ICAO checklist asks

  • Are there a range of tasks for different levels matching the ICAO levels the test aims to assess?
  • Do some tasks clearly assess higher levels of performance, with a sufficient amount of more difficult content in both the listening and speaking parts?
  • Is the complexity of the language the speaking tasks elicit a feature of the task content rather than of the rubric, prompts or interlocutor questions?
  • Is the difficulty of the listening tasks and items based on the complexity of the recordings rather than on interpreting long or complex instructions, questions or options?
Show how ELPAC addresses this

Differentiation is built into the task design first. Paper 1 progresses from short transmissions, where comprehension of a single instruction or request is enough, to longer and denser operational audio that carries more information than a Level 4 listener will reliably hold. In Paper 2, routine traffic in Task 1a is where Level 4 performance becomes visible, while Task 2 (unusual and non-routine situations) and Task 3 (hypothesising and problem-solving) require the flexibility, sensitivity to nuance and handling of the unexpected that separate Levels 5 and 6. Both papers therefore contain content deliberately pitched above the licensing threshold.

That design intention is then checked empirically rather than assumed. Paper 2 is rated live by two accredited examiners, an English Language Expert and an Operational Expert, who independently apply the ICAO holistic descriptors and the six ICAO rating scale criteria; the overall level awarded is the lowest of the six sub-ratings. Rating scale functioning and examiner severity are monitored with many-facet Rasch measurement on the Paper 2 evidence page. For Paper 1, item facility and discrimination, reliability and standard error of measurement are published per form under classical test statistics and Rasch analysis, which show how precisely the paper separates candidates around the decision points that matter.

On ICAO's point about the source of difficulty: instructions and rubrics are short and are given before the audio or the task begins, and the Paper 2 interlocutor script is written in operational language, so difficulty comes from the operational content rather than from working out what is being asked.

Qualification

Empirical discrimination is strongest around the Level 4 decision, which is where the licensing consequence sits. Precision at the boundaries furthest from the centre of the distribution rests on smaller numbers of candidates and is reported with its measurement error rather than presented as exact.

Criterion 7

Real-world contexts

Met by design

The criterion, as published in Doc 10197

Test instruments need to contain appropriate tasks that assess test-takers' abilities to understand and communicate in real-world contexts.

What the ICAO checklist asks

  • Does the instrument include tasks that directly assess how test-takers communicate in real-world job-related situations, in radiotelephony contexts and in other appropriate work-related situations?
  • Does the instrument include separate tasks to assess both listening and speaking skills in real-world job-related contexts?
  • Are the tasks specific to the needs of pilots or controllers in how each type of test-taker needs to communicate in real-world situations?
Show how ELPAC addresses this

Paper 2 is not an interview about aviation. The candidate works a traffic situation from a chart in their own operational role, handles routine traffic and then a developing complication, and afterwards reports the event to a supervisor. That covers both halves of the criterion: the radiotelephony context, and a work-related situation away from the radio.

Paper 1 uses genuine radiotelephony-style audio and operational material such as briefings and incident reports rather than general-English listening passages, so listening as well as speaking is assessed in contexts the candidate recognises from work, and in the contexts specific to their own role.

Doc 10197's guidance chapter on this criterion identifies role play as the speaking task type that maximises authenticity, and names debriefings and reporting on events as appropriate work-related tasks outside radiotelephony. Both are central to Paper 2 rather than peripheral to it.

Criterion 8

Test bank

Largely met

The criterion, as published in Doc 10197

Test instruments need to have a sufficient number of equivalent versions, with each version of the test representing the test instrument in the same way.

What the ICAO checklist asks

  • Is a test bank available, and does it contain sufficient versions to ensure the test remains secure, so the results are valid?
  • Does the test bank contain enough versions for the populations of pilots and of controllers who need to be assessed?
  • Does each version of the test represent the test instrument (test construct) in the same way?
  • Are all test versions equivalent in how different parts of the test assess specific levels of difficulty?
Show how ELPAC addresses this

Several operational Paper 1 versions are in circulation for each role, all built to the same specification: identical structure, number of parts and 45 items, so each version represents the construct in the same way. Comparability between versions is not assumed but checked: per-version reliability, mean score, standard error of measurement and item quality are published side by side under classical test statistics, and item difficulties are placed on a common scale in the Rasch analysis. Paper 2 uses several scripted scenarios per role, written to a common task specification; scenario allocation and rotation are controlled under the ELPAC Handbook and Assessment Scheme.

Reference: ICAO Doc 10197, Test Design Guidelines: Handbook on the Design of Tests for the ICAO Language Proficiency Requirements, First Edition, 2024. Criterion statements and checklist questions are quoted or condensed from that document for the purpose of this self-assessment; the document itself is published by ICAO and is not reproduced here.

Related pages

Continue reading