How Norudit Examines: Five Question Types, Real Mark Schemes
Most study apps "test" you with recognition: tap the right option, feel smart, learn little. Norudit's exams are built like the real thing: written by a chief examiner, marked to a scheme, and designed to probe the places where shallow knowledge hides.
Every class in Norudit ends topics with a formal exam (Phase 6), and the Mock Generator builds timed, multi-topic papers on demand. Both run on the same examination engine, and it is opinionated about what a question is for. Here is how it works.
Five question types, each hunting something different
A good paper is not a pile of questions; it is a set of instruments. Every Norudit exam draws on five types, each aimed at a different failure of understanding:
- Proof-flaw: a short argument or derivation with a subtle error planted in it; find the flaw and explain it. This catches students who can follow reasoning but not audit it.
- Comparison: weigh two related ideas and reach a supported judgement. This catches knowledge stored as isolated facts with no relations between them.
- Pattern: several cases or examples; induce the underlying rule (or spot which case breaks it). This catches rule-memorisation without abstraction.
- Application: apply the concept to a scenario it has never been shown in. This catches understanding that only works inside the textbook's own examples.
- Theory: the conceptual "why/how" of the mechanism itself. This catches procedure-following without a model underneath.
Each question carries a proper command word (Calculate, Explain, Compare, Evaluate, Assess, Find the flaw) and is tagged with the assessment objectives it targets (AO1 knowledge, AO2 application, AO3 analysis/evaluation), in the style of real exam boards (AQA / Edexcel / OCR).
The transfer-and-limits contract
The engine is bound by three rules that force every paper beyond recall:
- At least one application must sit in a genuinely unfamiliar context: a different field, industry, era or scale from the material, one that could not have appeared in the source. If the question could be answered by remembering an example, it isn't testing transfer.
- At least one boundary case: a scenario where the concept almost applies but doesn't, or applies counter-intuitively. Knowing a concept includes knowing its edges.
- At least one stress case: a scenario where applying the concept naively gives a wrong or misleading answer. This is the hardest and most honest test: does the student know where the model breaks down?
This is desirable difficulty (Bjork and Bjork, 2011) made structural. The paper is contractually obliged to leave the textbook behind.
The mark scheme is the examiner
Every question ships with its own official mark scheme, chosen to fit the question, exactly as real exam boards do:
- Points schemes for closed, technical, recall and calculation questions: a list of discrete marking points whose marks sum to the question's total, each descriptor saying exactly what earns it.
- Levels schemes for extended and evaluative answers: three to four bands with holistic descriptors, marked best-fit (the band first, then the mark within it).
Grading then applies the scheme mechanically: points are walked one by one with a reason for each award; levels answers are placed in a single best-fit band and justified against its descriptor. Never an overall impression, never a vibe, the same discipline behind how the Feynman gate is judged. You see the full breakdown: what was awarded, what was missed, and why.
Mock papers: many topics, one clock
The Mock Generator assembles a timed paper across the topics you choose, sized honestly to the clock (questions get realistic minutes, not wishful ones). Two things make a Norudit mock more than a stack of topic quizzes:
- Interlinked questions. Roughly a third of every mock can only be answered by connecting two or more topics: synthesis questions, marked at higher tariffs, that test the structure of your knowledge rather than its pieces.
- Weak-concept resurfacing, straight from your flashcards. The mock reaches into your review deck and finds the concepts you've been getting wrong. Every review card you answer updates a per-concept mastery estimate (Bayesian Knowledge Tracing; Corbett and Anderson, 1994); the generator pulls the lowest-mastery concepts within the paper's topics (weakest first) and deliberately reweaves them into the exam, reworded into fresh scenarios, never copied. So a concept you keep fumbling in daily review doesn't just wait in the deck; it comes for you in the mock, in exam form. The mock doubles as targeted remediation aimed precisely at your measured weak spots.
Retakes are parallel papers, not replays
Retake an exam and you don't get the same questions; you get a parallel paper: for each original question, a new one testing the same concept and skill, the same type, the same marks and scheme approach, but reworded with a fresh scenario so it cannot be answered from memory of the last attempt. The same principle as concept-varied review: the surface changes so only understanding survives.
Where the marks go
No exam dead-ends at a score. Every marked paper produces an error analysis of which concepts failed and why each mistake happened, and that analysis flows into the mastery model (Bayesian Knowledge Tracing) and from there into your review queue and the Smart Scheduler. The exam is not the end of the flywheel; it is what spins it.
The analysis reads more than right and wrong. Every answer also carries the confidence you declared when you gave it, and the exam sets that against whether you were actually correct. A confident wrong answer is not treated like a hesitant one: high-confidence errors are the priority target, because a belief you trust and act on does the most damage while it stands. And, as Butterfield and Metcalfe (2001) show, it is also the most durably corrected once caught. Surfacing the gap between what you think you know and what you actually know trains the metacognitive calibration students are otherwise poor at (Dunlosky and Metcalfe, 2009).
The why is not free text either. Every mistake is sorted into a fixed, closed set of error types (a misread command word, a missing step, a boundary ignored, a stress case handled naively, a concept confused with its neighbour) and the engine cannot invent new categories outside it. Because the set is fixed, it can be trended: across papers you see not just that you erred but which error class keeps recurring, so the pattern behind scattered mistakes becomes a single, fixable habit.
A score tells you where you stand. A mark scheme tells you why, and what to do next.
The exam engine, the mock generator, and everything around them are part of the full academy. Start your first class →
Sources
- Bjork, E. L. and Bjork, R. A. (2011) 'Making things hard on yourself, but in a good way: creating desirable difficulties to enhance learning', in Gernsbacher, M. A., Pew, R. W., Hough, L. M. and Pomerantz, J. R. (eds.) Psychology and the Real World: Essays Illustrating Fundamental Contributions to Society. New York: Worth Publishers, pp. 56–64.
- Butterfield, B. and Metcalfe, J. (2001) 'Errors committed with high confidence are hypercorrected', Journal of Experimental Psychology: Learning, Memory, and Cognition, 27(6), pp. 1491–1494.
- Corbett, A. T. and Anderson, J. R. (1994) 'Knowledge tracing: modeling the acquisition of procedural knowledge', User Modeling and User-Adapted Interaction, 4(4), pp. 253–278.
- Dunlosky, J. and Metcalfe, J. (2009) Metacognition. Thousand Oaks, CA: SAGE Publications.
Read more
- How a single explanation actually gets marked, fact by fact, not by vibe: How Do You Grade Understanding?
- Why you can answer the practice question but blank on the exam's reworded version: Learning That Travels
- Why a private tutor puts an average student in the top 2%, and what that means for a study tool: The 2 Sigma Problem