top of page

What ChatGPT gets wrong about IELTS (Proof for Teachers!)

  • Writer: Francis Carlisle
    Francis Carlisle
  • Jun 29
  • 4 min read

A student showed me a screenshot the other week. She'd pasted one of her IELTS Task 2 essays into ChatGPT, asked for a band score, and it had come back with a confident 7.5 and a paragraph of warm encouragement. She was thrilled. The trouble was, when I marked it, it was a 6.


I want to be fair to ChatGPT here, because this isn't one of those "AI is ruining everything" posts. It's a remarkable tool. I use it most days. For brainstorming topics, explaining a grammar point in three different ways, or turning a messy idea into a tidy paragraph, it's hard to beat. If your students are using it to push themselves to write more, that's a good thing.


The problem is narrower than "AI is bad." It's that ChatGPT is a general tool, and IELTS is a very specific exam with a very specific marking system. Here's where I see it go wrong most often, and why it matters for your students.



ChatGPT tells people what they want to hear

ChatGPT is built to be agreeable. It wants you to come away happy. So when a student asks "how's my essay?", the default instinct is to praise first, soften the criticism, and round the score up rather than down. 


For most uses, that's lovely. For exam preparation, it's dangerous, because the one thing a student needs before test day is an honest picture of where they actually are.

A student who thinks they're a 7 when they're a 6 won't put in the work to close that gap. They'll walk into the test relaxed and walk out confused, and the praise that felt so encouraging is part of why.



Argue with it and your IELTS score changes

Try this yourself. Paste an essay in, get a score, then reply "I think that's a bit harsh, I was aiming for 7." More often than not, the score creeps up. Push again and it'll creep up further.


A real examiner doesn't do that. The band you get is the band your writing earns against a fixed set of descriptors, and no amount of "but I really tried" moves it. 

That fixed standard is the entire point of IELTS. If you can talk the score upwards, it was never really a score, and the lesson the student takes away is that the criteria bend under pressure, when the whole exam is built on the fact that they don't.

ChatGPT will change the score if gave you if you challenge it. It is not a reliable way of telling your score.


The IELTS feedback is often shallow

When ChatGPT does give criticism, it tends to stay at the surface. "Try to use more linking words." "Vary your sentence structure." "Add more detail." 

It's not wrong, exactly. It's just the kind of advice you could give to almost any piece of writing in any context, and it doesn't tell the student what to fix in this sentence, right here.


The difference between Band 6 and Band 7 usually isn't a missing linking word. It's a pattern of small errors, a limited range of structures used safely rather than ambitiously, a paragraph that answers the question sideways instead of head-on. 

Catching that takes someone who can read the writing against the criteria, error by error. General advice doesn't get a student there.



ChatGPT doesn't really know the IELTS marking criteria

This is the root of the other three. IELTS Writing is marked on four things: task achievement, coherence and cohesion, lexical resource, and grammatical range and accuracy. Speaking has its own four. Each band within each criterion has a detailed descriptor, and an examiner is trained for a long time to apply them consistently.

ChatGPT has read about these criteria somewhere in its training, in the loose way it's read about everything. But reading about the band descriptors is not the same as being trained to apply them, the same way reading about driving is not the same as passing your test. 


So you get a number that sounds official and uses all the right vocabulary, but with very little of the disciplined, criteria-by-criteria judgement that should sit underneath it. It carries the authority of an examiner's mark without the training that an examiner actually has.



A general tool doing a specialist's job

Put all of that together and you get the real issue. ChatGPT is a magnificent generalist. IELTS rewards a narrow, trained, consistent kind of judgement that a generalist isn't built for. 


The model will always give you an answer, and it'll give it confidently, because that's what it does. That confidence is exactly the problem, because your students can't tell the difference between a well-calibrated score and a plausible-sounding guess, and neither tool will warn them which one they've got.


So where does that leave you as a teacher? Well, your judgement is the thing students actually need, and no general AI tool replaces it. The catch is that you can't sit with every student and mark every essay and every speaking attempt, several times a week, which is roughly what it would take.


That's the gap we built Keenu to fill. It's not a chatbot you talk into doing what you want. It scores writing and speaking against the actual IELTS criteria, trained on examiner scoring, with feedback that points at the specific errors rather than handing out generic tips. 


It won't flatter a student into a false sense of security, and arguing with it won't move the band. It's there to give your students honest, criteria-based feedback between your lessons, so the time you do spend with them goes on the interesting work rather than the marking.



The easiest way to see how this works is to try it yourself. You can take a mock test as a student would, see the scoring in action, and have a look at what the teacher dashboard shows. There's a 14-day free trial for schools, no card needed. If you'd rather have a quick call first, drop me a line at francis@keenu.io and we'll set something up.


 
 
bottom of page