Spotting AI written work in an English class is a judgement, not a test result. Detection software cannot prove authorship, and it flags second language writers most often of all. What actually works is task design that generates evidence of the writing process, plus a short conversation with the learner about their own text.
Key takeaways
- No detector produces proof. Treat every score as a prompt to look closer, never as a verdict.
- False positives cluster on exactly the learners you teach: people writing in a second language.
- Process evidence beats forensic analysis. Drafts, notes and in class writing settle the question faster.
- The most reliable signal is a gap between the submission and everything else that learner has produced.
- A two minute oral check resolves most cases without anyone being accused of anything.
- Task design is the real fix. Some tasks invite generated text and some make it pointless.
Why detection tools cannot carry the weight
Detectors work by measuring how predictable a text is. Generated writing tends to sit in a narrow band of very probable word choices. Human writing wanders more.
The problem for an English classroom is immediate and structural. Learners writing in a second language also produce predictable text. They reach for the phrases they have been drilled on, they stay inside a limited range of structures, and they avoid the risky, surprising word. A B1 writer doing everything right looks, statistically, quite a lot like a language model.
This is not a minor calibration issue. It means the error falls hardest on the careful learner who has memorised useful phrasing, and lightest on the confident writer who takes risks. If you run a detector across a mixed class, the names that come back are not random.
Even the companies selling these tools qualify their own numbers. The sensible reading is that a high score tells you to read the work properly. It does not tell you who wrote it.
What the official guidance actually says
It is worth knowing the formal position, because it is more moderate than the staffroom conversation usually is.
The Joint Council for Qualifications guidance on AI use in assessments puts the emphasis on teachers assuring themselves that work is the student’s own, and on students referencing any AI generated content they have used. The burden sits with professional judgement and centre procedure, not with a software score.
The UK Department for Education takes a similar line in its policy paper on generative artificial intelligence in education, which notes that teachers know their pupils and are experienced at identifying their individual work. That is the actual mechanism being relied on, and it is the one you already have.

The signals that mean something, and the ones that do not
Most of the tells people share online are unreliable in an ESL context. A few hold up.
| What you noticed | How much it tells you | Why |
|---|---|---|
| Vocabulary well above the level this learner has ever shown | Strong | Range is hard to fake upwards and you have a baseline |
| The usual first language error patterns have vanished | Strong | Article or aspect errors do not disappear in a fortnight |
| Fluent English with no personal or local detail | Moderate | Models write about anywhere and therefore about nowhere |
| Very even paragraph lengths and tidy signposting | Weak | This is also exactly what exam preparation teaches |
| No spelling mistakes at all | Weak | Spellcheck has existed for forty years |
| A detector returned a high score | Weak on its own | Second language writing inflates these scores |
Notice that the two strong signals both depend on knowing the learner’s previous work. That is the whole game. Without a baseline you are guessing, and with one you usually do not need software at all.
“30hr Advanced grammar teaching course”
“Quizzes and assignments were straight forward and feedback was easy to follow […]”
Kate Bygrave · September 2026 · verified review of a course from The TEFL Institute, the accredited training provider behind tefl.ai · Read the full review · All our reviews
Build a baseline before you need one
Twenty minutes of in class writing in the first fortnight, kept in a folder, is worth more than any subscription. You want a sample produced under conditions you controlled, so you know what this person writes like when nothing is helping them.
Repeat it every few weeks. It doubles as progress evidence, which makes it easy to justify to a director of studies, and it means that when something arrives that does not match, you have the comparison ready rather than a feeling.
Keep the tasks similar in type. Comparing a timed opinion paragraph against a researched project report tells you very little, because those are different kinds of writing.
Design tasks that make generated text useless
This is the part that removes most of the problem, and it does not involve policing anyone.
Ask for the specific. A paragraph about the transport system in the learner’s own city, referencing something that happened on their commute, is difficult to generate convincingly and easy to check by asking a follow up question.
Ask for the personal. Opinions tied to experience, reactions to a text you read together in class, a response to something a classmate said. Models are weak at anything that requires having been in the room.
Ask for the staged. A plan, then a draft, then a revision, with the earlier stages handed in. The finished piece matters less once you can see how it got there.
Ask in the room. Some writing simply happens in class, on paper or on a monitored screen. Not everything, and not as punishment, but enough that the baseline keeps refreshing itself.

The conversation, and how to have it
When something does not match, the move is a short chat, not an email with a score attached.
Ask the learner to talk you through one choice in the text. Why this example rather than another one. What this word means here. How they would extend the third paragraph if they had another hundred words. A learner who wrote it answers easily, sometimes at length, occasionally with a better version of the argument than the one on the page.
A learner who did not write it usually says so quickly, and often seems relieved. The reasons are rarely dramatic. Deadline collision, panic, a genuine belief that this was allowed, or a previous teacher who said something different.
Keep the framing on the learning. The point you want to land is that you cannot help them improve English they did not write, so the shortcut costs them the thing they are paying for. That argument works far better than a malpractice lecture, particularly with adult learners.
Document what happened and what was agreed, briefly, in whatever system your centre uses. Not to build a case, but because the second occurrence needs the first one on record.
Do not accuse on a score
Worth stating plainly, because the consequences are serious and they land on people a long way from home.
An accusation that turns out to be wrong damages the relationship permanently, and in a language classroom it can end a learner’s willingness to write anything ambitious ever again. The safe formulation is about the work, not the person. Something like: parts of this do not look like your usual writing, so talk me through it. That opens a conversation that a denial can end cleanly, which an accusation cannot.
If your centre has a malpractice procedure, that procedure exists to protect you as well as the learner. Use it rather than improvising.
How are English teachers actually using AI?
We are running a survey of English teachers on what AI has changed in their work, what it has not, and where it gets things wrong. Twelve questions, about four minutes, no email address required.
The results will be published free on tefl.ai. There is very little independent data on this, so what teachers tell us here is what the report will say.
Where this sits in the wider picture
Detection is downstream of policy. If learners do not know what is permitted, some of them will guess wrong in good faith, and you will spend your term adjudicating rather than teaching. Our guide to writing an AI policy for your English classroom covers the upstream half of this, and the wider context is in our guide to AI in English language teaching.
On the marking side, there is a reasonable version of using AI yourself, covered in AI for ESL writing corrections and grading English writing against CEFR levels. Levels there are worth checking against the published CEFR descriptors rather than a tool’s own label.
Start with the folder, not the software
If you do one thing this term, make it twenty minutes of in class writing from every learner, filed and dated. It costs you a single lesson slot and it answers the question you will inevitably be asked later.
When you get to marking, our free AI CEFR Writing Grader and IELTS Writing Band Estimator are free to use for a first pass, and the full set is on the AI tool overview page.
