The Deck That Rewrites Itself
Spaced repetition has proven to be the best method for memorising facts in the long term. In this blog I'll talk about one of the biggest limitations in how it gets applied, the problems with the well-deserved most popular spaced repetition tool, Anki, and Norudit's solutions to all of them.
Spaced repetition was first measured experimentally in 1885 by the German psychologist Hermann Ebbinghaus, who discovered the spacing effect and the forgetting curve. It was scaled into practical, large-scale studies in 1939 by H. F. Spitzer, validated as a premier cognitive strategy by comprehensive reviews like Dunlosky et al. (2013), and later tested at scale with Anki.
The major problem with current spaced repetition software
It's long been established that spaced repetition is the best approach when you need to memorise facts in the long term. But is that actually how you're tested? Whether someone has mastered a concept is judged on their understanding over the long term, not just their memory of it, and that's the major gap in current software.
That gap is what concept-varied spaced repetition exists to close. I invented it, and it's built on two principles, variable practice and desirable difficulties, both heavily researched by the cognitive psychologist Robert Bjork in the late 20th century. Before I get into it, though, it's worth being clear on why the current method falls short on understanding.
The problem is that you unintentionally spend your time memorising the visual pattern of the card. So it stops being a test of whether you understand the concept and becomes a test of how well you recognise that particular card, when the real exam expects both memorisation and understanding. That mismatch has always been there in conventional tools like Anki.
If pattern matching is the problem, the solution is obviously to prevent it, but how? What if, every time a card came up for review, a new question was generated for that same concept instead of showing you the old card again? That solves it, because you can't memorise a visual pattern that's never the same twice.
For the scheduling side I used FSRS-6, the Free Spaced Repetition Scheduler used by the latest version of Anki. FSRS-6 is the successor to SM-2, the algorithm Anki used previously, and it cuts the number of reviews you need by 20 to 30%. It can also be calibrated to your individual forgetting curve the more you use it, which matters because people forget at genuinely different rates, so the algorithm measures yours and personalises to it. So: FSRS-6 handles the card schedule, and an LLM handles the card variations.
That solved the pattern matching and the scheduling, but it created a new question. How different and how difficult should the new question be, and how much application should it demand when the question changes? After a lot of deliberation and research, the answer I landed on was that mastery should be tracked per concept, and the new question should be varied based on the student's mastery of that specific concept. For that I used BKT, Bayesian Knowledge Tracing, an edtech student-modelling algorithm for tracking how well a learner has actually mastered a skill.
With FSRS-6, BKT and an LLM, I accomplished my goal of fixing the major problem I and many others faced, with the invention of concept-varied spaced repetition.
One more thing: not all flashcards are concept-based, some are fact-based, and the two get tested in different ways. So I built a primary approach for each, which can vary if a card contains both. For concept-based cards the idea is to reword and re-angle the question. For fact-based cards, it's turned into an application-type question. Both are trying to mirror how that material actually shows up in an exam paper.
Concept-varied spaced repetition isn't really about flashcards, flashcards are just one application of it. The actual mechanism is: hold the underlying skill or concept fixed, regenerate and vary the retrieval surface every time to stop pattern-matching, gate how hard the variation gets on a per-skill mastery model, and schedule the whole thing with a forgetting-curve-fit algorithm.
For example, if you wanted to build something similar dedicated to language learning:
- Variation and skill/concept: for each new card, change the words or sentence so they're used in different contexts (or cases, for a language like Russian), and have verbs appear in different conjugations and tenses as part of a sentence.
- Mastery tracking: set the difficulty matching against language proficiency, A1 to C2.
- Scheduling: use the same FSRS-6 algorithm.
If you wanted to make it nicer you could add a TTS and STT model so the audio is generated automatically. That would solve a related problem, which is that new language learners tend to memorise words and phrases in isolation and then wonder why they can't actually speak. I'll be working on this after Norudit.
Two more examples, to show how much reach this has:
- Coding interview prep. The skill is the technique (two-pointer, DP on subsequences, whatever), and the surface is the specific problem. LeetCode either shows you the exact same problem again or throws you something unrelated. A version built this way could ask the same underlying question repeatedly, vary the difficulty against your tracked mastery, and prompt you on the homepage to review a skill based on your forgetting curve. Genuinely useful for a recent CS graduate prepping for interviews.
- Med-school case reasoning, which fits neatly given Anki's biggest audience is already medical students. The skill is the diagnostic reasoning pattern and the surface is the patient vignette.
So there's a lot of ground this covers. If there's a skill or concept, a way to test applying it, and a way to track mastery of it, there's a good chance concept-varied spaced repetition could work there, with review timing on top.
Anki's strengths, and its problems
Anki is the best known tool in this space. Ask any medical student and they won't shut up about how Anki is a life saver. It's free, and although it's a bit complicated to wrap your head around at the start, it's fairly easy to use once you have.
So what exactly are the problems with it? Setting aside concept-varied spaced repetition, there are two that I and plenty of other people have complained about: the backlog, and the lack of automation.
Problem 1: Anki burnout, the backlog trap
If you've ever been an avid Anki user you'll probably know the specific anxiety of opening it after a week off, or two weeks, or a month, and finding 238 cards due.
This is actually why I stopped using Anki. I'd see the number, do some of them, come back the next day to find even more waiting, and eventually get so overwhelmed that I never started again. While writing this I found out it has a name: Anki burnout, or the backlog trap.
I spent a while thinking about how to handle it. The first part of the answer was simple, cap the number of cards you have to review in a day, and you can already do that in Anki if you go looking for the setting. What Anki can't do is reorder the cards by how much they matter.
Most of your focus is at the very start of a session, which means the most important cards should be the ones at the top of the pile. Anki has no way of working out which those are. Norudit does, and it follows an order of priority:
- Cards you're actually weak on come first. Norudit tracks mastery per concept rather than per card, so cards testing your weakest concepts get pushed to the front. One thing I had to be careful about here: a concept you've never touched isn't the same as a concept you've touched and struggled with, so I never let an untested concept jump ahead of a genuinely weak one. Due cards take priority over new ones.
- Exam material jumps the queue. If you've got an exam coming up, anything tagged to it goes ahead of the rest of your backlog, so you're not grinding through unrelated cards with a deadline closing in. This is synced to the smart scheduler.
- New cards never get starved. Even with a big backlog, a few slots are reserved for new cards every session, so a backlog can't mean you simply never see anything new again.
- The backlog arrives in pieces. Instead of dumping everything on you at once, Norudit serves 30 cards in your first sitting, or fewer if you've turned your daily limit down. Finish those and you can ask for 20 more, and again after that, up to 50 a day per class, adjustable down to 15 if you want a gentler pace. So a frightening "238 cards due" never has to be faced all in one go.
That last part also fixes the problem with simply capping reviews in Anki, which is that a blind cap might cut exactly the cards you most needed to see.
A related issue worth mentioning is the streak. I never experienced this myself, but I've read a lot of people describing how fixated they got on maintaining one, how much anxiety that caused, and how losing a long streak killed their motivation for a week or more, which then led to Anki burnout as the overdue cards piled up. Norudit doesn't have streaks.
Problem 2: automation
Here's the other reason I left Anki: making cards is too time consuming and too stressful.
I needed that automated, but of course there are a million glorified flashcard makers (most study apps, basically) that already do it. So what makes this different?
- The obvious one is that it comes with everything above. Most of those apps also use the Leitner system, which is worse than SM-2, which has itself been surpassed by FSRS-6.
- The thing I found genuinely overwhelming: I'd upload a whole textbook and get flashcards for the entire thing at once, something like 354 cards, when what I wanted was a bite-sized version, or at least something organised. You could split the PDF yourself, but then you're doing the organisation manually, and with Norudit's backlog ordering and automatic splitting, the benefit of doing that by hand is almost nothing, leaving mostly just the headache.
But of course I wasn't going to make it that easy. So what kind of a pain in the ass thing did I do? Before any card gets generated, you have to prove you understand the topic using the Feynman technique in Phase 3, because being able to simplify something into your own words is the best measure of understanding there is. Then you have to retrieve the topic from your head, with the technical terms, in Phase 4. Once you've passed both, the cards get created and added automatically, or you can make them yourself if you'd rather.
Because here's the thing. Like I said earlier, real tests expect memorisation and understanding. So before you even get to spaced repetition, you need to have understood the concept and be able to remember and use the technical terms. Those two are the prerequisites.
Concept-varied spaced repetition preserves understanding and facts in the long term, so first you need the understanding and the facts locked in with Phase 3 (Feynman) and Phase 4 (blurting) before repetition is worth anything.
If you'd like to learn more about The Norudit Academy, here's the full tour. Or check out the landing page. If you feel ready to try it, go ahead and sign up.
Sources
- Bjork, R. A. (1994) 'Memory and metamemory considerations in the training of human beings', in Metcalfe, J. and Shimamura, A. J. (eds.) Metacognition: Knowing about Knowing. Cambridge, MA: MIT Press, pp. 185–205.
- Ebbinghaus, H. (1885) Über das Gedächtnis: Untersuchungen zur experimentellen Psychologie. Leipzig: Duncker & Humblot. (Translated as Memory: A Contribution to Experimental Psychology, 1913.)
- Ye, J., Su, J. and Cao, Y. (2022) 'A stochastic shortest path algorithm for optimizing spaced repetition scheduling', in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. New York: ACM.
Read more
- How Norudit fits its spaced-repetition schedule to your own forgetting curve: How Norudit Actually Personalises Your Learning
- How Norudit plans your study week like maths rather than a guess: The Scheduler