How to write an exam rubric an AI can actually apply
Most rubrics are written for someone who already knows the subject. Making one explicit enough for a machine turns out to make it fairer for students too.
Ask ten examiners to mark the same script against the same rubric and you will get a spread. Not because anyone is careless — because the rubric is doing far less work than we assume. It is a set of reminders for someone who already knows what a good answer looks like.
That is fine when a human is marking. It stops being fine the moment you want consistency across a large batch, a second examiner, or any kind of automated support.
Here is what changes when you write a rubric to be applied literally.
1. Mark the reasoning, not the conclusion
A rubric that says “Correct answer: 4 marks” tells you nothing about the script where the method is flawless and the arithmetic slips in the last line.
Break the marks across the steps of the reasoning:
| Weak | Better |
|---|---|
| Correct answer — 4 marks | States the correct governing equation — 1 Substitutes given values correctly — 1 Solves for the unknown — 1 States the result with correct units — 1 |
This is not extra bureaucracy. It is what you are already doing in your head; writing it down means a second examiner, a moderation committee and a machine all do the same thing you would.
2. Say what partial credit looks like
The single biggest source of drift is partial credit. Every rubric needs an explicit answer to: what does a half-right answer earn?
Write the bands:
- Full marks — the criterion is met without qualification
- Partial — name the specific situations that earn it, e.g. “correct method, arithmetic error in a single step”
- Zero — name what fails entirely, e.g. “wrong governing principle, regardless of the arithmetic”
If you cannot write the partial band, the criterion is probably two criteria.
3. Name the acceptable alternatives
Students take routes you did not anticipate. A rubric that only encodes your preferred method will penalise a correct answer for being unfamiliar — and this is where automated marking most visibly goes wrong.
Add, explicitly: “Any method that establishes X is acceptable, including A, B or C.” You will not catch every route, but naming three is dramatically better than naming one.
4. Separate content from communication
Many rubrics quietly mix them: a student loses marks for a disorganised answer without the rubric ever saying presentation carries weight.
Decide, then state it. If communication is worth marks, give it its own criterion with its own band. If it is not, the messy-but-correct script must get full marks — and your rubric should say so, because otherwise different examiners will resolve it differently.
5. Write the criteria as observable checks
The test: could someone who is not you apply this criterion and get the same answer?
- ❌ “Shows good understanding of the concept”
- ✅ “Identifies that the system is in equilibrium and states the condition used”
“Good understanding” is not observable. The second version is a check you can run against a script. Most vague criteria become two or three concrete ones when you push on them, which is a sign the original was hiding a judgement you had not made explicit.
6. Decide what happens to unanswerable cases
Every batch throws up scripts the rubric does not cover — the answer to Q3 written inside the Q5 space, the student who answered a different question well, the illegible middle page.
You are already making these calls. Write down the common ones as policy. It takes twenty minutes once and removes a whole category of inconsistency.
The side effect nobody expects
Departments start this exercise because they want to use assessment software. They usually finish it having found something more useful: the places where their own examiners had been quietly disagreeing for years.
Making a rubric explicit enough for a machine makes it explicit enough for a new faculty member, a visiting examiner, and — importantly — for the student who wants to know why they lost two marks.
That last one matters more than it sounds. A student who can read the criteria and see exactly which check they failed has been taught something. A student who receives a number has not.
A practical starting point
Do not rewrite everything. Take one question from your last end-semester paper and rewrite its rubric using the six points above. It will take about half an hour. Then hand the original rubric and the rewritten one to two colleagues with the same five scripts.
The spread between their marks is your answer.
GunanQ applies the rubric you write — your criteria, your weightings, held steady from the first script to the three-hundredth. Book a session and we will grade one of your own papers against your own rubric.
AI-Shala Team
Research & Engineering
Written collectively by the people who build and teach here — engineers, researchers and mentors who spend their week with the problems these posts describe.