I spent 15 years marking IELTS essays. Then I tried to explain it to a machine.
- Francis Carlisle

- Jul 23
- 6 min read
There's a moment every IELTS examiner knows. You read the last line of an essay, put your pen down, and you just know. Band 6.5.
You don't run through a checklist. You don't deliberate. I marked essays like that for years, and I was good at it. But when we started building Keenu, I discovered something uncomfortable: I couldn't explain how I did it.

The problem I kept ignoring
As an IELTS examiner in China, the same thing kept bothering me. Students who were working hard, paying for classes, putting in the hours, and still going into the IELTS writing exam without much sense of how they were performing.
Not because their teachers didn't care - teachers care enormously. But there are thirty or more students in a class and one teacher. Marking a single IELTS writing task properly, not ticking boxes but giving the kind of specific, calibrated feedback that actually moves a student's band score, takes lots of time. More time than a teaching schedule allows.
So most students got surface feedback. One big thing to work on, but not mentioning the rest. They'd get a band score the teacher had guessed and a general comment, useful but not enough.
To build a tool that marks like an examiner,
You have to be able to describe what an examiner does - in detail and in language precise enough that it can be turned into a system prompt, a set of criteria, a repeatable process.
I sat down to do this and realised I was trying to explain how to ride a bike.
You can ride a bike. You've ridden a bike for years. But if someone asks you to describe exactly which muscles you're using, exactly how you shift your weight, exactly what you do with your hands when you start to tip, you don't know how to explain it!
That's what examiner instinct is like. In 15 years I had marked thousands of essays - I knew what coherence looked like, I knew what distinguished a band 6 from a band 7 in task achievement. But knowing and explaining are completely different things.
So that became the work. What does "coherence and cohesion" actually mean in a band 6 essay, as opposed to a band 7?
Forget the official descriptor. What does it look like on the page? What is an examiner's eye actually catching? How do you divide coherence from task achievement when the two are so intertwined that one almost always affects the other?
We went through this for every criterion, every band level, every question type. And every time I thought we'd nailed something, someone on the team would ask me to be more specific, and I'd realise I had more to unpack.
It was slow and uncomfortable, and it was the most valuable thing we did.
Three versions, one feedback page
The feedback page that exists today is the third major version. Each time we rebuilt it, we started almost from scratch. Each iteration involved long, concentrated sessions where we argued, politely but properly argued, about things that might sound minor from the outside.
Should feedback expand on click or be visible upfront? How much detail is too much before a student switches off? When the system identifies a grammar issue, how do you display it in a way that's corrective without being crushing?
None of those questions have obvious answers. And the process of finding non-obvious answers, of sitting with something that isn't quite right yet and working out why, is what makes the difference between a 'thrown together' product that technically functions and a product that actually helps. We moved slowly on purpose, and then we moved again, and again, and each time the product got closer to what we'd been trying to build from the beginning.
The first time we looked back at our original mockups, we honestly couldn't believe what we'd thought was acceptable. The early work wasn't bad, it was the right start. But you don't know what the product needs to be until you've spent enough time sitting with the problem.
The feedback sandwich, at scale
One of the hardest design problems wasn't technical, it was human. Any teacher knows the feedback sandwich: you open with something the student did well, address what needs to improve, and close with encouragement. It isn't a trick, it's how effective feedback works. Students who feel seen and recognised are more likely to engage with criticism. Students who feel only criticised disengage.
The challenge was building that into a system that works for every student, every essay, at any band level, every time.
The Keenu feedback page doesn't just correct. It highlights what the student got right: specific vocabulary choices, structural decisions, content points they hit.
It tells them exactly which of the required Task 1 content points they included, and which they missed, and what those missing points actually are. Not "you need to include more detail." Here is point three. Here is what it requires. Here is how your essay handled it.
And the tone took serious work to calibrate. Too positive and it loses credibility. Too corrective and it demoralises. The system had to learn, through iteration and review, how to be honest and constructive at the same time. Not unlike a good teacher, which is the point.
The calibration no one else does
Before we launched, I did something that I think separates this product from anything else in this space. I marked 100s of essays myself.
We had collected real student essays. I ran them through the new feedback system and then sat down with the outputs as an examiner. I compared what the system said with what I would have said. Where they didn't match, we adjusted. Not the score, necessarily, but the framing, the emphasis, the balance between positive and corrective. The things that are hard to specify in advance but obvious when you see them done wrong.
This is ongoing, not a one-time calibration. As the product grows, so does the review process.
No general AI tool does this. There's no IELTS examiner sitting behind ChatGPT checking whether its feedback matches what a real marking session would produce.
With Keenu, there is.
What feedback students actually get on their IELTS Writing now
The feedback a student receives after a Keenu mock test isn't a grammar report or a generic AI summary. It's a structured, examiner-informed analysis of their writing, broken down by the criteria that IELTS examiners actually use.
It tells them which high-level vocabulary they used effectively, and names the words. It identifies overused words and shows them in context, then gives alternatives for those overused words, also in context.
For Task 1, it maps their essay against the data points the question required and shows them what they hit and what they missed. It gives an overall band score with a breakdown by criterion. And it does all of this in a tone that a good teacher would recognise: honest, specific, and (this was hard to get right) properly encouraging where encouragement is warranted.
A student can take a mock test, submit their essay, and receive feedback that previously only existed in a one-to-one session with an experienced teacher or examiner. At any hour, anywhere in the world, as many times as they need.
For language schools, this means your students are getting meaningful feedback on their writing practice whether or not you have time to mark every essay yourself. It doesn't replace your teaching. It extends what's possible.
What I learned from building the IELTS Writing simulator
I came out of this process understanding IELTS writing in a way I didn't before, and I've been in the IELTS game for nearly fifteen years.
When you have to make the implicit explicit, when you have to find words for the things you've always just known, you end up with a clearer, more articulable understanding of your own expertise. I can now explain, in detail, what makes a band 7 essay different from a band 6 at the level of coherence and cohesion, in a way I couldn't before we started building.
That knowledge is now inside the product. Every student who uses Keenu is getting access to it.
I think that's what we're most proud of. The technology is a means to an end. What we're proud of is that the feedback is good. It says what I would say. Now a student who can't afford a private tutor, or whose teacher doesn't have time to mark thirty essays, gets the same quality of feedback as one who does.
That's what we set out to build. We think we've built it.
Keenu is an AI-powered IELTS mock test platform for language schools. Teachers can run mock writing tests with their students and receive automated, examiner-quality feedback on every essay.

