Where words go missing
Students capture 91% of a lecture's top-level ideas and 11% of the details. Vocabulary lives in that 11%. Here is where the words go, and what actually brings them back.
You know the notebook. The first three pages are neat. By page six it has thinned into a column of bare words with a hopeful arrow beside a few of them. Around week five the entries stop, and the notebook joins the drawer where a decade of other notebooks ended exactly the same way.
The usual verdict is discipline. It is almost never discipline. Cognitive psychology has spent forty years measuring what falls out of a set of notes, in what order, and what puts it back — and the answers are more specific, and considerably stranger, than “try harder.”
Your notes are already mostly empty
Start with what actually reaches the page. In 1987 Kiewra, Benton and Lewis did something almost nobody had bothered to do: they broke a lecture into a hierarchy of ideas and counted, level by level, how much of each level students wrote down.
The headline is the collapse. The interesting part is which level empties.
Everything a vocabulary learner needs lives at the bottom of that hierarchy. The pronunciation. The part of speech. Which preposition it takes. The register that makes it fine at a bar and wrong in an email. The one example sentence that shows what the word actually does. Those are precisely the things an instructor says once, in passing, on the way to the next slide — and precisely the things the page does not keep.
Baker and Lombardi came at it from another direction and landed in the same place: student notes contained under a quarter of the propositions independent judges had rated worth recording, and about half of the main ideas the lecture was built to deliver.
So the notebook was never full. It was thin on the first day, in a specific and predictable pattern, and the pattern has a physical cause.
The hand was never going to keep up
Speech runs at two to three words per second. Handwriting runs at 0.2 to 0.3. That is a tenfold gap, and no amount of resolve closes it.
What closes it, badly, is triage. You hold a sentence in working memory while deciding what to drop, and while you decide, the next sentence goes by. Piolat, Olive and Kellogg measured what that costs with a dual-task probe: people taking notes were slower to catch an interrupting signal than people reading the very same material. Taking notes is not transcription. It is live editing against a deadline you did not set.
Source and statistics
Interference reaction time: 661 ms while taking notes vs 590 ms while reading. Speech 2–3 words per second vs handwriting 0.2–0.3 words per second.
Piolat, A., Olive, T., & Kellogg, R. T. (2005). Cognitive effort during note taking. Applied Cognitive Psychology, 19, 291–312.
Now add a second language.
To take good notes in a language class, you have to already be good at the language.
That is not a figure of speech. Asaly-Zetawi and Lipka watched 63 bilingual undergraduates write in both of their languages. Same hand, same person, and switching to the second language cost them about a third of their writing speed.
Then they asked what predicted the quality of the notes, and the answer was not motor at all. Vocabulary size and reading fluency were the two strongest predictors by a wide margin — well ahead of how fast the hand moved, and ahead of working memory too.
Source and statistics
N = 63. Handwriting speed 68.79 in L1 vs 46.80 in L2, F(1,61) = 30.67, p < .001, ηp² = 0.34. Correlations with note quality: vocabulary r = .60, reading fluency r = .60, handwriting speed r = .37, working memory r = .34.
Asaly-Zetawi, M., & Lipka, O. (2019). Note-taking skill among bilingual students in academia. Frontiers in Psychology.
Read that back slowly. The skill you enrolled to acquire is the entry fee for recording the class in which you would acquire it.
Then the research turns on you
Here is where the story stops being about effort and starts being about structure.
If writing things down were the active ingredient, the effect of note-taking should be large. It is not. Kobayashi’s 2005 meta-analysis pooled 57 studies, 131 independent samples and 306 effect sizes, comparing people who took notes against people who did not. The pooled advantage was small — real, but nowhere near what the ritual of note-taking promises.
Source and statistics
k = 57 studies, 131 independent samples, 306 effect sizes. Note-taking vs no note-taking, weighted mean d = 0.22 (unweighted 0.29).
Kobayashi, K. (2005). What limits the encoding effect of note-taking? A meta-analytic examination. Contemporary Educational Psychology, 30, 242–262.
In 2025 the authors of an earlier second-language note-taking meta-analysis did something rare: they brought in a methodologist and re-ran their own paper, 27 studies and 55 effect sizes, under a multivariate multilevel model. Then they split the results by what kind of note was taken.
Left to your own devices, note-taking did not even reach significance. Sorted into a vocabulary notebook, it produced the biggest effect in the whole re-analysis. Same hour, same lecture, same student.
Look at the bottom row. The vocabulary notebook came out on top, and its lead over taking notes your own way was not a rounding artifact — that gap was tested on its own and held.
Source and statistics
27 studies, 55 effect sizes, immediate post-test (Gain1). Conventional note-taking 0.275, p = 0.299 (n.s.); strategy instruction 0.977, p = 0.007; framework notes 1.297, p = 0.001; vocabulary notebook 1.852, p = 0.003. Conventional minus vocabulary notebook = −1.576, p = .014.
The rest of the literature points the same way. Jin and Webb’s 2024 meta-analysis of 21 studies and 1,992 learners split its cases by whether the learners had ever been taught a note-taking strategy. With that instruction, the effect was more than five times the size it reached without it. Kobayashi’s earlier review of 33 studies concluded that providing a framework or the instructor’s notes outperformed pre-training and verbal instruction — and that the lower a learner’s academic level, the more they gained from it. The intervention works hardest on beginners.
Source and statistics
k = 21, N = 1,992. Overall g = 0.56 [0.24, 0.88]. With strategy instruction g = 0.84 vs without g = 0.16, Q_between p = .007.
Jin, Z., & Webb, S. (2024). The effectiveness of note taking through exposure to L2 input: A meta-analysis. Studies in Second Language Acquisition, 46(2), 404–426.
Even a lecturer’s phrasing counts as structure. Titsworth and Kiewra ran 60 students through lectures with and without spoken organizational cues; with the cues, students recorded 39% more organizational points and 35% more detail, and test achievement rose by about 13%.
Which narrows the whole problem to one question: who builds the frame? In most classrooms, nobody does. The frame is homework you were never assigned, due at 3× the speed your hand can write.
In the app: the frame gets filled for you
In class, your job is one word, written by hand. On-device recognition turns it into text in one to two seconds, and the rest of the card fills itself: the base form, IPA, part of speech, meaning, a grammar note, two example sentences with translations, and a CEFR level. Japanese and Chinese entries get pronunciation ruby alongside the characters.
What the literature calls framework notes is a skeleton an instructor prepares in advance and hands out. Vocady builds that skeleton per word, for the words you actually met. What accumulates is the thing at the top of the chart above: a vocabulary notebook.
Recognition covers 300+ languages and runs fully offline once the language model is downloaded.
Words survive only when you pull them out
A filled-in note is not the finish line. This is where most studying quietly fails.
In 2013 five cognitive psychologists graded the ten learning techniques students actually use. Two earned the top rating.
| Utility | Technique |
|---|---|
| High | Practice testing (retrieval) · distributed practice |
| Moderate | Elaborative interrogation, self-explanation, interleaving |
| Low | Summarizing, highlighting, mnemonics, imagery, rereading |
Highlighting is in the bottom row. So is rereading — the thing nearly everyone means by the word “review.” If you have ever finished a chapter with a page glowing yellow and no idea what it said, you have already run the experiment. Here is the controlled version.
Check five minutes later and rereading wins. That is why we trust it. The material is warm, it comes back easily, and the feeling in your chest says got it.
A week later, in the room where it counts, it is 42% against 56%. Rereading manufactures the sense of knowing now. Retrieval manufactures the state of knowing later. The exam is always later.
Source and statistics
N = 120. After 5 minutes: rereading 81% vs testing 75% (d = 0.52). After 2 days: 54% vs 68% (d = 0.95). After 1 week: 42% vs 56% (d = 0.83). Main effect F(1,117) = 36.39, ηp² = .24.
Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.
How unreliable is that feeling? In 2008 Science published the cleanest possible test. Forty students learned 40 foreign-language word pairs under different conditions, then predicted how many they would still recall in a week.
The predictions were effectively identical across all four conditions — about half, roughly 20 of 40, in every group.
A week later the actual scores were about 80% for the group that kept retrieving the words, against 36% and 33% for the groups that had been dropped from retrieval. The two distributions did not overlap anywhere — not one student in the retrieval group scored as low as the best student in the others. The paper’s own phrasing is “greater than 150% improvements in long-term retention.”
Source and statistics
N = 40. One week later, repeated retrieval ≈ 80% vs 36% and 33% for the conditions dropped from retrieval, d = 4.03. Score ranges did not overlap: retrieval 63–95%, non-retrieval 10–60%. Predicted recall was ≈ 50% in all four conditions.
Same confidence, same study time, more than double the outcome. Your sense of how well you know a word carries almost no information.
The evidence has only thickened since. A 2021 meta-analysis in Psychological Bulletin pooled 222 independent studies and 48,478 students in real courses — not lab cubicles — and the testing advantage came through at moderate strength.
Source and statistics
222 independent classroom studies, N = 48,478 students. g ≈ 0.50.
Yang, C., Luo, L., Vadillo, M. A., Yu, R., & Shanks, D. R. (2021). Testing (quizzing) boosts classroom learning: A systematic and meta-analytic review. Psychological Bulletin, 147(4), 399–435.
Retrieval does come with a condition, and it is worth stating plainly because most flashcard habits violate it. In Rowland’s synthesis of 159 effect sizes, showing learners the right answer after each attempt roughly doubled what retrieval practice bought them. And in the worst case — no feedback, on material so new that the first attempt succeeds half the time or less — the benefit vanished outright. New words are exactly that case. Testing yourself on a word you have not learned yet, with no answer shown, is not practice. It is a coin flip you feel bad about.
Delay works in your favor too: waiting a day or more before you retrieve beats retrieving the same day.
Source and statistics
159 effect sizes from 61 papers. Overall g = 0.50 [0.42, 0.58]. With feedback g = 0.73 vs without g = 0.39. No feedback plus initial retrieval success ≤ 50%: g = 0.03 [−0.21, 0.27], p = .79. Retrieval delayed a day or more g = 0.69 vs same-day 0.41.
Rowland, C. A. (2014). The effect of testing versus restudy on retention: A meta-analytic review of the testing effect. Psychological Bulletin, 140(6), 1432–1463.
Notes obey the same rule. A 2024 meta-analysis of 24 studies and 49 effect sizes split lecture notes by whether learners were given time to review them, and the notes that got reviewed were worth exactly twice what the unreviewed ones were.
Source and statistics
24 studies from 21 papers, 49 effect sizes. With review g = 0.421 [0.287, 0.554] vs without review g = 0.208 [0.124, 0.292], Q(2) = 7.94, p = .019. Overall achievement g = 0.248 [0.181, 0.315].
Flanigan, A. E., Wheeler, J., Colliot, T., Lu, J., & Kiewra, K. A. (2024). Typed versus handwritten lecture notes and college student achievement: A meta-analysis. Educational Psychology Review, 36:78.
A note pays when you take it out, not when you put it in.
In the app: something pulls the words back out
Saved words do not sit there. Daily Recap serves up to 10 questions a day, and all five question types are retrieval tasks: choose the meaning from the word, infer the missing word in a sentence, choose the word from the meaning, assemble a sentence, and write the word out by hand.
Nothing here is a card you flip. You produce the answer first, every time. Each question shows the correct answer and an example immediately after, which puts every question on the feedback side of Rowland’s split rather than the no-feedback side.
One word yields one question per session — never the same word twice in a night, because the second answer would come from working memory rather than from memory worth measuring.
When your vocabulary is too small to fill four plausible options, the question is demoted rather than dropped: four choices, then three, then two, then handwriting. Note the direction. Demotion bottoms out in producing the word rather than recognizing it, so the floor of the system is harder than the ceiling.
The same hours, spaced out
The other high-utility technique is distributed practice. Total study time held constant, how you slice it changes the result.
Cepeda and colleagues pooled 317 experiments from 184 papers and 14,811 participants: final recall of 47.3% when the study was spread out, against 36.7% when it was crammed. A 2008 follow-up with 1,354 participants measured the gaps directly and got recall improvements of up to 111% by changing nothing but the spacing. The best gap tracked the delay to the test — roughly 20% of it, so a test weeks away wants gaps of days, and a test a year out wants something closer to 5%.
Source and statistics
2006: 184 papers, 317 experiments, 271 comparisons, N = 14,811. Spaced 47.3% vs massed 36.7%, t(540) = 6.6, p < .001. 2008: N = 1,354. Recall improvement up to 111% (d = 1.7); optimal gap ≈ 20% of the retention interval, falling to ≈ 5% at one year.
Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Psychological Bulletin, 132(3), 354–380 · Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T., & Pashler, H. (2008). Psychological Science, 19(11), 1095–1102.
Two honest caveats, because the popular version of this finding oversells it. Expanding intervals — the “review at 1 day, 3 days, 7 days” schedules sold as optimal — did not beat plain uniform ones in a 2021 review. And when Bego and colleagues ran spacing in real undergraduate courses, only 2 of the 9 courses showed a significant effect. Spacing is well supported. Any specific schedule marketed as the correct one is not.
Source and statistics
Expanding vs uniform intervals g = 0.034, not significant. Spacing reached significance in 2 of 9 undergraduate courses.
Latimier, A., et al. (2021). Expanding versus uniform spacing schedules: a meta-analysis. Educational Psychology Review · Bego, C. R., et al. (2024). Spacing effects across undergraduate courses. International Journal of STEM Education.
In the app: what comes up tonight is not random
Recap ordering is a weight, not a shuffle. Every word starts at 10.0, or 15.0 if you have never been quizzed on it, and three terms are added:
+ (5 − mastery) × 3.0— weaker words first+ (wrong / attempts) × 5.0— words you keep missing first+ min(days since last review, 10)— words you have not seen in a while first
The ten questions you get tonight are the ten words you are currently losing.
Handwritten letters land differently
Why write at all, then, if the app is going to type it out anyway?
Because for an unfamiliar writing system, the hand does something the eye and the keyboard do not. Wiley and Rapp taught 42 adults a novel alphabet under three conditions — writing it by hand, typing it, or studying it visually.
Letter production came out at 91.0 for writing, 78.5 for typing and 64.5 for looking. Letter naming: 93.4, 86.2, 82.8. The result that matters most was on a task nobody had trained for — spelling whole words the participants had never practiced — where the handwriting group scored 76.3 against the typing group’s 62.3.
Source and statistics
42 adults, 12 per condition completing the post-test. Letter production 91.0 / 78.5 / 64.5 (writing / typing / visual study). Letter naming 93.4 / 86.2 / 82.8. Untrained word spelling 76.3 vs 62.3, p < .05.
Wiley, R. W., & Rapp, B. (2021). The effects of handwriting experience on literacy learning. Psychological Science, 32(7), 1086–1103.
The benefit reached tasks that were never taught. Only for the group that wrote.
Stroke order in Chinese characters, the curves of kana, the shape of Cyrillic and Arabic letters: one pass with a pen does what twenty passes with your eyes do not.
Now the boundary, stated clearly, because this is where handwriting claims usually go wrong. What that study measured is acquisition of letter forms, not comprehension of lecture content. And on the broader question of handwritten versus typed notes, the honest answer is that nobody has established a winner. Voyer and colleagues in 2022 and Urry and colleagues in 2021 both landed on a difference statistically indistinguishable from zero, and the 2024 meta-analysis found typing comfortably ahead on how much gets recorded. Its meaningful split was not the input method at all. It was whether anyone reviewed the notes.
Source and statistics
Handwritten vs typed notes: Voyer et al. g = −0.008; Urry et al. g = 0.04, passing an equivalence test against zero. Amount recorded favors typing, g = 0.919 (Flanigan et al.).
Voyer, D., et al. (2022) · Urry, H. L., et al. (2021) · Flanigan, A. E., et al. (2024). Educational Psychology Review, 36:78.
So the case for handwriting here is narrow and specific: it helps you learn what unfamiliar characters look like, and it is the fastest way to get one word out of your head and into an app during a class that is still moving.
In the app: the strokes are kept, not just the text
After recognition finishes, the original strokes stay. Coordinates are normalized to a device-independent grid, so a word scrawled large on a phone and one written small on a tablet open at the same size and proportion everywhere.
Timing is stored alongside them. Tap the handwriting on a word’s detail page and it redraws itself in the order and at the speed you wrote it, including the pauses where the pen was in the air between strokes. Watching how you built a character is information a photograph cannot give you.
One hour of class, two versions
| In a paper notebook | In Vocady | |
|---|---|---|
| During class | You try to catch the word, the meaning and the sound at once — 0.3 words per second of hand chasing 3 words per second of speech, a third slower again if it is a second language. | You handwrite the word. Nothing else, because the rest is coming. Your eyes and ears stay in the room. |
| Right after | 91% of the main themes, 11% of the detail. The pronunciation column is usually blank. | A card with base form, IPA, part of speech, meaning, grammar note, two examples with translations and a CEFR level — plus the handwriting you produced. |
| That night | You reread the notes, using the technique that sits in the bottom row of the table above. | Ten recap questions. All retrieval, answers shown, weakest words first. |
| Two weeks later | The main themes survive. The notebook is on the desk. | The words have been pulled out several times across spaced days, on the train and on planes. |
The difference is not better note-taking. It is writing less, and making sure that what you wrote comes back.
Common questions
How many words do I need before reviews start?
There is no threshold. One saved word produces a question that night. Under ten words, you get however many you have.
With a vocabulary of three, a four-option question cannot be built, so it drops to fewer options or to writing the word out by hand. Breaking the loop at the very start seemed like the worst available outcome.
Wouldn’t drilling the same word several times a day be better?
Asking twice in one session means the second answer comes out of working memory. Getting right what you saw ninety seconds ago is not evidence of anything durable, so the rule is one word, one question per session.
It comes back tomorrow, and the week after. Spacing the retrievals is the part with evidence behind it.
Does it work without a connection?
Handwriting recognition runs entirely on the device. Download a language model once and it works on a plane or underground; 300+ languages are supported.
Saving words and doing reviews are offline-first and sync across your devices when a connection returns. Only the AI analysis of a new word needs the network.
Can I import a whole vocabulary list from a photo?
Yes — photograph a textbook page or a word list and it imports up to 50 entries at once. Imported words are analyzed the same way handwritten ones are and enter that night’s recap on the same footing.
Studies cited
- Kiewra, K. A., Benton, S. L., & Lewis, L. B. (1987). Qualitative aspects of notetaking and their relationship to information-processing ability. Journal of Instructional Psychology.
- Baker, L., & Lombardi, B. R. (1985). Students’ lecture notes and their relation to test performance. Teaching of Psychology, 12(1).
- Titsworth, B. S., & Kiewra, K. A. (2004). Spoken organizational lecture cues and student notetaking as facilitators of student learning. Contemporary Educational Psychology, 29(4).
- Piolat, A., Olive, T., & Kellogg, R. T. (2005). Cognitive effort during note taking. Applied Cognitive Psychology, 19, 291–312.
- Asaly-Zetawi, M., & Lipka, O. (2019). Note-taking skill among bilingual students in academia. Frontiers in Psychology.
- Kobayashi, K. (2005). What limits the encoding effect of note-taking? A meta-analytic examination. Contemporary Educational Psychology, 30, 242–262.
- Norouzian, R., Jin, Z., & Webb, S. (2025). Increasing meta-analytic quality: A multivariate multilevel meta-analysis of note-taking through exposure to L2 input. The Modern Language Journal, 109(1), 171–193.
- Jin, Z., & Webb, S. (2024). The effectiveness of note taking through exposure to L2 input: A meta-analysis. Studies in Second Language Acquisition, 46(2), 404–426.
- Kobayashi, K. (2006). Combined effects of note-taking/-reviewing on learning and the enhancement through interventions: A meta-analytic review. Educational Psychology, 26(3), 459–477.
- Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving students’ learning with effective learning techniques. Psychological Science in the Public Interest, 14(1), 4–58.
- Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.
- Karpicke, J. D., & Roediger, H. L., III (2008). The critical importance of retrieval for learning. Science, 319(5865), 966–968.
- Yang, C., Luo, L., Vadillo, M. A., Yu, R., & Shanks, D. R. (2021). Testing (quizzing) boosts classroom learning: A systematic and meta-analytic review. Psychological Bulletin, 147(4), 399–435.
- Rowland, C. A. (2014). The effect of testing versus restudy on retention: A meta-analytic review of the testing effect. Psychological Bulletin, 140(6), 1432–1463.
- Flanigan, A. E., Wheeler, J., Colliot, T., Lu, J., & Kiewra, K. A. (2024). Typed versus handwritten lecture notes and college student achievement: A meta-analysis. Educational Psychology Review, 36:78.
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.
- Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T., & Pashler, H. (2008). Spacing effects in learning: A temporal ridgeline of optimal retention. Psychological Science, 19(11), 1095–1102.
- Latimier, A., et al. (2021). Expanding versus uniform spacing schedules: a meta-analysis. Educational Psychology Review.
- Bego, C. R., et al. (2024). Spacing effects across undergraduate courses. International Journal of STEM Education.
- Wiley, R. W., & Rapp, B. (2021). The effects of handwriting experience on literacy learning. Psychological Science, 32(7), 1086–1103.
Voyer et al. (2022) and Urry et al. (2021), cited above for the handwriting-versus-typing comparison, are referenced in text only.
Every effect size and test statistic on this page was checked against the source paper or its publisher abstract. Where a paper does not report a number, we do not report one either.