7 min read

How Norudit Examines: Five Question Types, Real Mark Schemes

noruditproductassessmentlearning science

Most study apps "test" you with recognition: tap the right option, feel smart, learn little. Norudit's exams are built like the real thing: written by a chief examiner, marked to a scheme, and designed to probe the places where shallow knowledge hides.

Every class in Norudit ends topics with a formal exam (Phase 6), and the Mock Generator builds timed, multi-topic papers on demand. Both run on the same examination engine, and it is opinionated about what a question is for. Here is how it works.

Five question types, each hunting something different

A good paper is not a pile of questions; it is a set of instruments. Every Norudit exam draws on five types, each aimed at a different failure of understanding:

Each question carries a proper command word (Calculate, Explain, Compare, Evaluate, Assess, Find the flaw) and is tagged with the assessment objectives it targets (AO1 knowledge, AO2 application, AO3 analysis/evaluation), in the style of real exam boards (AQA / Edexcel / OCR).

The transfer-and-limits contract

The engine is bound by three rules that force every paper beyond recall:

  1. At least one application must sit in a genuinely unfamiliar context: a different field, industry, era or scale from the material, one that could not have appeared in the source. If the question could be answered by remembering an example, it isn't testing transfer.
  2. At least one boundary case: a scenario where the concept almost applies but doesn't, or applies counter-intuitively. Knowing a concept includes knowing its edges.
  3. At least one stress case: a scenario where applying the concept naively gives a wrong or misleading answer. This is the hardest and most honest test: does the student know where the model breaks down?

This is desirable difficulty (Bjork and Bjork, 2011) made structural. The paper is contractually obliged to leave the textbook behind.

The mark scheme is the examiner

Every question ships with its own official mark scheme, chosen to fit the question, exactly as real exam boards do:

Grading then applies the scheme mechanically: points are walked one by one with a reason for each award; levels answers are placed in a single best-fit band and justified against its descriptor. Never an overall impression, never a vibe, the same discipline behind how the Feynman gate is judged. You see the full breakdown: what was awarded, what was missed, and why.

Mock papers: many topics, one clock

The Mock Generator assembles a timed paper across the topics you choose, sized honestly to the clock (questions get realistic minutes, not wishful ones). Two things make a Norudit mock more than a stack of topic quizzes:

Retakes are parallel papers, not replays

Retake an exam and you don't get the same questions; you get a parallel paper: for each original question, a new one testing the same concept and skill, the same type, the same marks and scheme approach, but reworded with a fresh scenario so it cannot be answered from memory of the last attempt. The same principle as concept-varied review: the surface changes so only understanding survives.

Where the marks go

No exam dead-ends at a score. Every marked paper produces an error analysis of which concepts failed and why each mistake happened, and that analysis flows into the mastery model (Bayesian Knowledge Tracing) and from there into your review queue and the Smart Scheduler. The exam is not the end of the flywheel; it is what spins it.

The analysis reads more than right and wrong. Every answer also carries the confidence you declared when you gave it, and the exam sets that against whether you were actually correct. A confident wrong answer is not treated like a hesitant one: high-confidence errors are the priority target, because a belief you trust and act on does the most damage while it stands. And, as Butterfield and Metcalfe (2001) show, it is also the most durably corrected once caught. Surfacing the gap between what you think you know and what you actually know trains the metacognitive calibration students are otherwise poor at (Dunlosky and Metcalfe, 2009).

The why is not free text either. Every mistake is sorted into a fixed, closed set of error types (a misread command word, a missing step, a boundary ignored, a stress case handled naively, a concept confused with its neighbour) and the engine cannot invent new categories outside it. Because the set is fixed, it can be trended: across papers you see not just that you erred but which error class keeps recurring, so the pattern behind scattered mistakes becomes a single, fixable habit.

A score tells you where you stand. A mark scheme tells you why, and what to do next.

The exam engine, the mock generator, and everything around them are part of the full academy. Start your first class →

Sources

Read more