7 min read

How Norudit Examines: Five Question Types, Real Mark Schemes

noruditproductassessmentlearning science

Most study apps test you by recognition. Tap the correct option, feel clever for a second, learn almost nothing. The exams here are built like the real thing instead: written the way a chief examiner would write them, marked against a scheme, and aimed deliberately at the places shallow knowledge likes to hide.

Every class closes its topics with a formal exam in Phase 6, and the Mock Generator will build timed, multi-topic papers whenever you ask for one. Both run off the same engine, and that engine has strong opinions about what a question is even for.

Five question types, each hunting a different failure

A good paper isn't a pile of questions, it's a set of instruments, and each of the five types is pointed at a different way understanding can be hollow.

Proof-flaw. A short argument or derivation with a subtle error planted somewhere in it, and your job is to find the flaw and explain what's wrong with it. This catches the student who can follow reasoning but can't audit it, which is a very common place to be without realising.

Comparison. Weigh two related ideas against each other and reach a judgement you can support. This catches knowledge that's been stored as a set of isolated facts with no relationships running between them.

Pattern. Several cases or examples, from which you have to induce the underlying rule, or spot which one of them breaks it. This catches rule-memorisation with no abstraction sitting underneath it.

Application. Take the concept into a scenario it has never been shown in. This catches understanding that only works inside the textbook's own examples and collapses the moment it's moved.

Theory. The conceptual why and how of the mechanism itself. This catches procedure-following with no model underneath, which gets people surprisingly far before it stops.

Every question carries a proper command word, so calculate, explain, compare, evaluate, assess, find the flaw, and gets tagged with the assessment objectives it's targeting: AO1 for knowledge, AO2 for application, AO3 for analysis and evaluation, in the style the real boards use.

The transfer-and-limits contract

Three rules bind the engine, and together they force every paper past pure recall.

  1. At least one application in a genuinely unfamiliar context. A different field, industry, era or scale from the source material, one that couldn't plausibly have appeared in it. Because if a question can be answered by remembering an example, it isn't testing transfer, it's testing memory wearing a costume.
  2. At least one boundary case. A scenario where the concept almost applies but doesn't, or applies counter-intuitively. Knowing a concept properly includes knowing where it stops being true.
  3. At least one stress case. A scenario where applying the concept naively hands you a wrong or misleading answer. That's the hardest and the most honest test available, because it's asking whether you know where the model breaks down.

That's desirable difficulty (Bjork and Bjork, 2011) turned into structure. The paper is contractually obliged to leave the textbook behind at least three times before it's finished.

The mark scheme does the examining

Every question ships with its own mark scheme, picked to suit that question, the way the real boards actually do it.

Points schemes for the closed, technical, recall and calculation questions: a list of discrete marking points whose marks add up to the question total, each descriptor stating exactly what earns it.

Levels schemes for extended and evaluative answers: three or four bands with holistic descriptors, marked best-fit, so you place the band first and then find the mark inside it.

Grading then applies whichever scheme mechanically. Points get walked one at a time with a reason attached to each award. Levels answers get placed in a single best-fit band and justified against that band's descriptor. Never an overall impression and never a vibe, which is the same discipline behind how the Feynman gate gets judged. You see the entire breakdown afterwards: what was awarded, what was missed, and why in both cases.

Mock papers: several topics, one clock

The Mock Generator puts together a timed paper across whichever topics you choose, sized honestly against the clock so questions get realistic minutes rather than optimistic ones. Two things make it more than a stack of topic quizzes glued together.

Interlinked questions. Roughly a third of any mock can only be answered by connecting two or more topics together. They're synthesis questions carrying higher tariffs, and what they're testing is the structure of your knowledge rather than the individual pieces of it.

Weak concepts dragged out of your own review deck. Every review card you answer updates a per-concept mastery estimate (Corbett and Anderson, 1994), and the generator reaches into that, takes the lowest-mastery concepts sitting inside the paper's topics, and deliberately weaves them back into the exam, reworded into fresh scenarios rather than copied across. So a concept you keep fumbling in daily review doesn't just sit in the deck waiting for you, it comes and finds you in the mock, in exam form. The mock ends up doubling as remediation aimed precisely where you're measurably weakest.

Retakes are parallel papers, not replays

Retake an exam and you don't get the same questions back. You get a parallel paper, where each original question is replaced by a new one testing the same concept and skill, the same type, the same marks, the same scheme approach, but reworded around a fresh scenario so it can't be answered from memory of the last attempt.

Same principle as concept-varied review: change the surface, and only the understanding survives the change.

Where the marks end up

No exam dead-ends at a score. Every marked paper produces an error analysis, which concepts failed and why each individual mistake happened, and that flows into the mastery model and from there into your review queue and the scheduler. The exam isn't the end of the flywheel, it's the thing that spins it.

The analysis also reads more than right and wrong. Every answer carries the confidence you declared when you gave it, and the exam sets that against whether you were actually correct. A confident wrong answer isn't treated like a hesitant one, because a belief you trust and act on does the most damage while it stands, and as Butterfield and Metcalfe (2001) showed, it's also the most durably corrected once you catch it. Surfacing that gap between what you think you know and what you do know trains the metacognitive calibration students are otherwise quite bad at (Dunlosky and Metcalfe, 2009).

And the why isn't free text either. Every mistake gets sorted into a fixed, closed set of error types, so a misread command word, a missing step, a boundary ignored, a stress case handled naively, a concept confused with its neighbour, and the engine can't invent new categories outside that set. Because the set is fixed it can be trended across papers, which means you stop seeing a scatter of unrelated mistakes and start seeing which single class of error keeps recurring. That's usually one fixable habit rather than ten separate problems.

A score tells you where you stand. A mark scheme tells you why, and what to do next.

The exam engine, the mock generator and everything sitting around them are part of the full academy.

Start your first class →

Sources

Read more