SLP speaking is rated on what your speech does, not on how it sounds
Updated 2026-08-18 · Independent educational resource
Speaking is a performance, not a pronunciation sample
Independent resource. SLP Command is not affiliated with NATO, BILC, any Ministry of Defence, or any official examining body. National tests implement STANAG 6001 descriptors differently. Always confirm administration details with the authority that runs your sitting.
Most candidates prepare for speaking as if the examiner were listening for mistakes. That is the wrong model, and it produces a recognisable failure: a careful, error-light answer that never attempts what the task asked for.
A speaking rating asks whether your speech did the job — described, narrated, compared, justified, hedged, recommended — at the level's standard of precision. Accuracy is one input to that judgement. It is not the judgement.
The four factors behind a rating
Proficiency ratings in the STANAG/ILR family are usually read across four factors rather than as a single impression:
| Factor | The question it answers | Typical way it is lost |
|---|---|---|
| Content | What subject matter could you actually handle? | Comfortable only on personal and routine topics when the level asks for abstract ones |
| Tasks | What did your speech do — describe, narrate, argue, qualify? | Answering a "justify and recommend" prompt with a description |
| Accuracy | Was it precise enough to be understood without effort? | Errors that force the listener to reinterpret, not occasional slips |
| Text produced | What shape of speech came out — a phrase, a paragraph, a sustained argument? | Level 3 reasoning delivered as disconnected sentences |
This four-factor reading is standard testing practice and how SLP Command structures its own evaluation. It is an interpretive lens, not a sentence quoted from STANAG 6001.
The weakest factor caps the rating
These four are not averaged. A response with Level 3 content and Level 2 accuracy is not credited somewhere in between — the limiting factor decides.
That single fact explains most results that feel unfair:
- The fluent speaker capped by precision, because errors keep costing the listener effort.
- The precise speaker capped by tasks, because they never attempted the reasoning the prompt required.
- The well-prepared speaker capped by content, fluent on their own unit and lost on an abstract policy question.
- The speaker capped by text produced, who has the argument but delivers it as fragments that never build.
It also tells you what to train: not "speaking" in general, but the factor that is holding you.
A worked contrast
Prompt (illustrative, not from a live official paper): Your unit has been offered additional training hours that must be taken from either maintenance or physical training. Recommend which, and justify it.
A confident answer that is capped: a fluent, accurate description of what maintenance involves and why physical training matters. Nothing wrong with the language. It described when it was asked to recommend — the task was not performed.
An answer that reaches the level: names the recommendation early, gives the reason that actually decides it, concedes the cost on the other side, and qualifies the conditions under which the answer would change.
The second answer can contain more errors and still be the stronger performance, because the factor it is strong on is the one the prompt was testing.
How to train it this week
- Take a prompt that requires a position, not a description — "recommend", "justify", "compare and decide".
- Record yourself answering under a clock, in one take. No restarts; restarts train a skill the sitting will not let you use.
- Before listening back, write down which of the four factors you think was weakest.
- Listen back once and check. Most people are wrong about which factor limited them — that is the point of the exercise.
- Train that factor specifically for a week. Precision drills will not fix a task problem, and task drills will not fix precision.
Speaking to yourself without recording feels productive and teaches very little, because the factor you are weakest on is exactly the one you cannot hear while you are producing it.
How SLP Command evaluates speaking
Speaking evaluation returns each of the four factors as met or not met, with the evidence it used, and names the limiting factor when a task was not credited. A single task does not receive a decimal profile — one performance is not a rating.
Sending audio for evaluation is a separate, explicit, revocable choice, and never a condition of using the rest of the product. The Responsible AI policy states what the model receives.
Questions
Will my accent lower my score?
An accent is not itself a failing. What matters is whether it costs the listener effort — intelligibility is assessed, a particular accent is not the target.
I speak fluently. Why was I not credited at Level 3?
Fluency is one factor among several. A confident, fast answer that never attempts the reasoning the task called for can be credited below a slower answer that does.
Is it better to say less and be accurate, or say more and risk errors?
Neither strategy wins on its own, because the weakest factor caps the rating. Saying very little protects accuracy while failing on the tasks attempted; overreaching does the reverse.