How Do You Grade Understanding?
The moment a tool claims to grade understanding, the honest question is: graded how? Because there is a lazy way that looks impressive in a demo and falls apart in practice.
The lazy way is to hand an AI a student's explanation and ask, "give this a score out of a hundred." You get a number back. It looks authoritative. It is, on inspection, close to a guess.
Why one holistic score fails
A single overall score has three problems that compound. It is unreliable: ask twice and you may get 71 and 84 for the same answer, because the model is integrating a dozen fuzzy considerations in one opaque step. It is opaque: the number tells the student nothing about what was missing, which, as Hattie and Timperley (2007) show, is the part that most helps them improve. And it is untunable: if the standard is too harsh or too soft, there is nothing to adjust except the wording of a prompt and your hopes.
For a checkpoint a student's progress depends on, "trust the vibe of a number" is not good enough.
The alternative: many small, checkable judgments
Grading understanding well means refusing the single big judgment and replacing it with many tiny ones. Break the thing you care about into concrete, atomic questions (did the explanation cover this key idea? is this specific claim correct? did it handle the follow-up rather than dodge it?) and judge each one on its own, with a short reason attached.
Small yes/no judgments are the thing models are actually reliable at, in exactly the way holistic scores are the thing they are worst at. And each little verdict, with its reason, is the feedback the student needs: not "73," but "you nailed the mechanism, your definition was slightly off here, and you skipped this consequence."
The final score is then a matter of transparent arithmetic over those verdicts, not a number the model conjured. You can see precisely which check produced which part of the result. And you can adjust the standard by changing the arithmetic, in the open, rather than by re-praying to a prompt.
Grade against the material, not the model's memory
There is one more thing a trustworthy grade requires: it has to judge the student's claims against the actual source material, not against whatever the model happens to believe. A model's own priors are exactly where hallucinated "corrections" come from. Anchoring every judgment to the real text is what separates grading from opinion.
And when the material genuinely cannot settle a point, the honest move is to skip that check, not to guess. A question the source can't support should lower how confident the grade is, not silently drag the score up or down.
Grading is deterministic arithmetic over many small, grounded judgments, never a single holistic score a model made up.
A number you can't interrogate isn't a grade. It's a guess wearing a lab coat.
The payoff is not just fairness. It is that a grade built this way is legible: to the student, who learns from the reasons; and to the people building the tool, who can see exactly where the standard sits and move it deliberately. A score you can argue with is worth more than one you can only accept.
Sources
- Hattie, J. and Timperley, H. (2007) 'The Power of Feedback', Review of Educational Research, 77(1), pp. 81–112.
Read more
- The five question types and real mark schemes behind Norudit's mock exams: How Norudit Examines: Five Question Types, Real Mark Schemes
- Why the real test of learning is applying it to a problem you've never seen, not memorising it: Learning That Travels
- Benjamin Bloom's 1984 finding that one-to-one tutoring beats classroom teaching by two standard deviations: The 2 Sigma Problem