How Do You Grade Understanding?
The second any tool claims it can grade understanding, the question worth asking is: graded how, exactly? Because there's a lazy way of doing it that demos beautifully and comes apart the moment somebody's progress actually depends on the result.
The lazy way is handing an AI a student's explanation and asking it to score that out of a hundred. A number comes back. It looks authoritative sitting there. And on inspection it's barely distinguishable from a guess.
Why one big score doesn't work
A single overall number has three problems and they stack on each other.
It's unreliable, because ask twice and you can get 71 and then 84 for the identical answer, since the model is compressing a dozen fuzzy considerations into one opaque step and nothing is anchoring it.
It's opaque, because the number tells the student nothing about what was actually missing, and what was missing is the part that helps them improve (Hattie and Timperley, 2007). A grade with no reasons attached isn't feedback, it's a verdict.
And it's untunable, because when the standard turns out too harsh or too soft there's nothing to adjust except the wording of a prompt and your own optimism.
For a checkpoint that gates somebody's progress, trusting the vibe of a number isn't good enough.
Lots of small checkable judgments instead
Grading understanding properly means refusing the one big judgment and replacing it with a pile of tiny ones.
Break the thing you care about into concrete atomic questions. Did the explanation cover this key idea? Is this specific claim correct? Did it handle the follow-up or quietly dodge past it? Then judge each of those on its own, independently, with a short reason attached to each verdict.
There's a practical reason this works better, which is that small yes-or-no judgments are the thing language models are genuinely reliable at, in exactly the same proportion that holistic scores are the thing they're worst at. So you're asking the tool to do the job it can actually do.
And each little verdict, with its reason, turns out to be the feedback the student needed anyway. Not "73", but "you got the mechanism, your definition was slightly off here, and you skipped this consequence completely".
The final score then becomes transparent arithmetic over those verdicts rather than a number the model invented, which means you can see precisely which check produced which part of the result. It also means the standard can be moved by changing the arithmetic, in the open, rather than by going back and praying at a prompt.
Grade against the material, not against the model's memory
There's one more requirement for a grade worth trusting, which is that it judges the student's claims against the actual source material rather than against whatever the model happens to believe about the subject. A model's own priors are exactly where hallucinated corrections come from, and anchoring every judgment to the real text is the thing separating grading from opinion. That principle runs through the whole product and has its own post.
And when the material genuinely can't settle a point, the honest move is to skip that check rather than guess at it. A question the source can't support should lower how confident the grade is, not quietly drag the score up or down in whichever direction the model leaned.
Grading is deterministic arithmetic over many small, grounded judgments, never a single holistic score a model made up.
A number you can't interrogate isn't a grade. It's a guess wearing a lab coat.
The payoff here isn't only fairness, though it is that. It's that a grade built this way is legible in both directions: to the student, who learns from the reasons rather than the number, and to whoever's building the tool, who can see exactly where the standard currently sits and move it deliberately rather than by accident. A score you can argue with is worth considerably more than one you can only accept.
Sources
- Hattie, J. and Timperley, H. (2007) 'The Power of Feedback', Review of Educational Research, 77(1), pp. 81–112.
Read more
- The five question types and real mark schemes behind Norudit's mock exams: How Norudit Examines
- Why the real test of learning is applying it to a problem you've never seen: Learning That Travels
- Bloom's finding that one-to-one tutoring beats classroom teaching by two standard deviations: The 2 Sigma Problem