ESL listening activities can now be built end to end with AI: a graded script, a synthetic recording, comprehension questions and a transcript in about ten minutes. The gain is enormous for teachers with no suitable audio. The catch is that generated speech is too clean, too even and too polite to prepare anyone for a real conversation.
Key takeaways
- Listening is the skill most teachers under resource, because suitable audio is hard to find and slow to make.
- AI is strongest at drafting scripts and questions, and weakest at producing speech that sounds like people actually talking.
- Always give the model a level, a context, a number of speakers and a reason for the conversation. Vague prompts produce radio announcements.
- Synthetic audio has no overlap, no false starts and no background noise, so it is easier than real life. Plan for that gap.
- Use generated audio for controlled practice and authentic audio for the messy parts. Both belong in a course.
- Recording your own learners is personal data. Get permission before it goes anywhere near a tool.
Why listening is the skill that gets skipped
Most teachers do not avoid listening because they think it is unimportant. They avoid it because the logistics are miserable. The coursebook audio is about a topic your class finished two weeks ago. The podcast you like is far too fast. The news clip is three minutes of vocabulary nobody needs. Making your own recording means writing a script, finding a second voice, recording it without a dog barking, then writing the questions.
That is the entire reason listening lessons often become reading lessons with the teacher performing the text out loud.
Generated audio changes that arithmetic. You can produce a two minute conversation between two speakers, at roughly the right level, about the thing your class is studying this week, in the time it takes to make tea. Whether it is any good is a separate question, and the honest answer is that it depends on what you ask for and what you do afterwards.
What AI can genuinely produce for a listening lesson
A graded script. This is the strongest use by a distance. Ask for a conversation at a stated level, with a stated number of speakers, a stated setting and a stated outcome, and you will get a workable first draft.
Audio from that script. Text to speech has moved far enough that a two voice dialogue sounds like two people rather than a train station announcement. It still sounds like two very calm people.
Comprehension questions. Gist questions, detail questions and inference questions, generated from the script rather than from the recording, which means they are always accurate to the text.
A transcript. You already have it, because you wrote the script. That sounds obvious until you remember how much class time is lost hunting for the line where the speaker said the thing.
Variations. The same situation at two levels, or the same dialogue with one detail changed, which is the basis of several good tasks further down this article.

The prompt that actually works
Most disappointing listening material comes from a prompt like “write a B1 conversation about travel”. What you get back is two people taking turns to deliver complete, grammatical sentences about how much they enjoy travelling. Nobody has ever spoken like this.
Give the model six things instead: the level, the setting, the relationship between the speakers, the reason the conversation is happening, the length in words, and the language you want to appear. Then add one instruction that changes everything: tell it that the speakers should interrupt each other, that one of them should not understand something the first time, and that the conversation should end without everything being resolved.
Add the target language explicitly. If the lesson is about polite requests, say that polite requests must appear at least four times in natural positions. Models are good at following that instruction and terrible at guessing it.
| Listening task | What AI drafts well | What you check before class | Best suited to |
|---|---|---|---|
| Two speaker dialogue | Script, roles, target language in context | Turn length, whether anyone sounds human | A1 to B1 controlled practice |
| Monologue or voicemail | Clear structure and a task with one answer | Speed, and whether details are guessable | Note taking and number practice |
| Dictogloss text | Short dense paragraphs built on one structure | Length, because generated texts run long | B1 to B2 grammar consolidation |
| Comprehension questions | Gist, detail and inference sets with keys | Questions answerable without listening | Every level |
| Authentic clip support | Glossaries and pre teaching from a transcript | Transcript accuracy on accented speech | B2 and above |
Level bands here are typical starting points rather than fixed rules. Your class will move them.
The problem nobody mentions: it is all too easy
Generated speech is articulate. Speakers wait their turn, finish their sentences, pronounce every word and never say “sorry, what?”. Real English does none of this. People talk over each other, abandon half their sentences, swallow function words and repair what they just said.
A learner who only ever practises on synthetic audio gets very good at understanding synthetic audio. Then they arrive at a service desk in Manchester and understand nothing, which is demoralising and entirely predictable.
The fix is not to abandon generated audio. It is to be deliberate about the split. Use generated material for controlled practice, for the target language, for the things where you need a clean example. Use real speech, with all its mess, for training the skill of coping. Both are necessary, and the second one is where the confidence comes from.
You can also push the model towards mess. Ask for hesitation, self correction, an interruption and one speaker who is distracted. It will not be perfect, but it will be better than the default.
“30hr Advanced grammar teaching course”
“The course content was really accessible and easy to go through […]”
Kate Bygrave · September 2026 · verified review of a course from The TEFL Institute, the accredited training provider behind tefl.ai · Read the full review · All our reviews
Six listening activities worth the preparation
1. Two versions, one situation. Generate the same conversation twice, changing three facts. Half the class hears one version, half hears the other. They then compare in pairs and find the differences, which forces both groups to talk about what they actually heard rather than what they assumed.
2. Predict the ending. Play the first two thirds. Learners write what happens next, then hear it. Generated dialogues are ideal here because you control where the break falls.
3. Dictogloss. A short dense paragraph, read twice at normal speed, reconstructed in groups. Ask the model for a paragraph built around one grammatical structure and keep it to about eighty words, because the first draft will be double that.
4. The unhelpful speaker. Generate a service conversation where one speaker is vague, distracted or mishears. Learners note what information is missing. This is the closest a generated task gets to real listening.
5. Transcript repair. Take the script, introduce ten errors, and have learners correct it while listening. It sounds fiddly and takes two minutes to make.
6. Extensive listening at home. Generate a short serialised story, one episode per week at a steady level. Learners listen without a task, for pleasure and volume. This is the activity most likely to move a class and least likely to be assigned.

Getting the level right
Level in listening is not just vocabulary. It is speed, density, how much is said indirectly, and how much the listener has to hold in memory before the answer arrives.
Models are unreliable on all four. Ask for A2 and you will often get B1 vocabulary delivered in short sentences, which is a different thing entirely. Read the script against the CEFR descriptors published by the Council of Europe and edit rather than trusting the label.
Three quick checks that catch most problems. Count the unfamiliar words in the first thirty seconds: more than three and the class will disengage before the task starts. Look at the longest turn: anything over about forty words at A2 is a monologue pretending to be a conversation. And check where the answers sit, because generated scripts love to put every answer in the final ten seconds.
If you want the levelling done more carefully, the same CEFR logic applies to written material, and our guide to adapting teaching materials with AI covers the checking routine in detail.
Research
How are English teachers actually using AI?
We are running a survey of English teachers on what AI has changed in their work, what it has not, and where it gets things wrong. Twelve questions, about four minutes, no email address required.
The results will be published free on tefl.ai. There is very little independent data on this, so what teachers tell us here is what the report will say.
Accents, and why one voice is not enough
If every recording your class hears is the same polished accent, you are teaching them one accent. English is not one accent. Most text to speech systems offer several varieties, so rotate them across a term rather than settling on the one you find easiest to listen to.
Be careful about what you tell learners, though. A synthetic Irish or Indian or Nigerian voice is an approximation, and presenting it as a model of how people in a place speak is not accurate. Say plainly that it is a generated voice. For genuine variety, real recordings still win, and there is no substitute for hearing a person who is tired, in a hurry and standing next to a road.
Recording your own learners
Some of the best listening material in any classroom is the class. Learners interviewing each other, then transcribing what they said, produces material at exactly the right level about things they care about.
One rule though. A recording of a learner’s voice is personal data, and a recording of a child is sensitive. Get permission, keep it local where you can, and do not upload it to a tool your institution has not approved. The Irish Data Protection Commission guidance is short and worth reading once properly.
Where AI does not belong in a listening lesson
It does not belong in the part where learners fail. The purpose of a listening task is partly to let people not understand something and then work out how to cope, and a teacher who rescues that moment with a transcript on screen has removed the lesson.
It also does not belong in assessment without care. Generated audio is easier than authentic speech, so a learner who scores well on it has demonstrated something narrower than you might assume.
And it does not replace the second best listening resource in the building, which is you. Teacher talk, adjusted live to a class you are watching, does something no recording does.
Where to start this week
Take the topic your class is on now and build one dialogue. Six prompt elements, one target structure, two voices, roughly ninety seconds. Write three gist questions and five detail questions. Run it, then keep the transcript for a second lesson.
If it works, the rest follows quickly, and the whole lesson around it can be drafted the same way using our guide to ESL lesson plans with AI. If speaking practice is the next gap, our piece on using AI as a speaking partner covers the other half of the conversation, and the broader picture sits in our guide to AI in English language teaching.
Every free tool we run is on the AI tool overview page, and if you want the skills assessed rather than just practised, the AI-Skilled Teacher Certificate is the structured route.
