BlogTalking through reviews vs clicking through cardsComparison

Talking through reviews vs clicking through cards

In short

  • Matched for information, a spoken queue takes about 14 minutes against 11 for clicking.
  • Speaking every due item instead of clicking it would take roughly three times as long.
  • The minutes come from a different budget: spoken review works while walking, clicking does not.
  • The sensible split is voice for the conceptual items, clicking for the recognition bulk.

The pitch for reviewing out loud usually includes a claim about speed. It is worth checking, because our own numbers do not support it.

What we costed

Two earlier pieces supply the inputs.

From what it costs to hold two thousand concepts: a collection of 2,000 concepts settles at about 74 due items a day, which we costed at roughly eleven minutes — an average of 9 seconds per item, the normal figure for recognition-level review.

From oral exams vs multiple choice: a spoken answer marked on two independent dimensions carries about 2.46× the information of a single five-option multiple-choice item.

A spoken exchange takes longer than nine seconds. The question has to be spoken, you have to answer in sentences, and a follow-up costs more again. We have used 28 seconds per spoken item below. It is an estimate, and it is the number the conclusion is most sensitive to — the working is laid out so you can substitute your own.

Three ways to spend the day’s queue

Approach Items Seconds each Total
Click through everything due 74 9 11.1 min
Speak everything due 74 28 34.5 min
Speak the equivalent information 30 28 14.0 min

The third row is the fair comparison. If one spoken item measures 2.46 times as much as one clicked item, you need 74 ÷ 2.46 ≈ 30 spoken items to learn as much about yourself as the full clicked queue told you.

At which point speaking costs 14 minutes against 11 — about a quarter longer.

So the speed claim is wrong. Talking through your reviews is not faster than clicking through them. Anyone telling you otherwise is comparing a short spoken session against a full clicked queue, which is the second row against the first.

Why it is still the better trade

Because the fourteen minutes and the eleven minutes come from different budgets.

Clicking requires a screen, a free hand, and your eyes. It has to be carved out of the part of the day you could have spent studying something else. Every one of those eleven minutes is a real eleven minutes.

Speaking requires none of those. It works on the walk to campus, in a supermarket queue, while cooking. Those fourteen minutes are spent during time that was not going to produce any revision at all.

Measured as minutes elapsed, clicking wins. Measured as desk time consumed, the spoken session costs zero and the clicked session costs eleven minutes. For most students the second measure is the one that is actually scarce.

This is also why the honest version of the pitch is not “finish your reviews faster”. It is: your reviews stop needing a slot.

Where speaking is genuinely worse

Two places, and they matter.

Pure recognition items. A drug name, a value, a single association — these are what nine-second clicking is for. Wrapping them in a spoken exchange multiplies the cost without adding anything, because there was no structure to assess.

Anything you cannot check by ear. Diagrams, spatial relationships, anything where the answer is a picture. Labelling exists because pointing at a structure is not a thing you can do out loud.

The split that actually works

Sort the day’s queue by what kind of knowledge it is.

  • Conceptual, mechanistic, anything with a because — speak it. This is where the 2.46× comes from, and where a follow-up question finds something.
  • Recognition, recall of a single fact — click it. Nine seconds, no ceremony.

On a typical 74-item day that might be twenty spoken items on the walk in, and the remaining fifty-odd cleared at a desk in six or seven minutes. Total desk time: under ten minutes, down from eleven, with the harder half assessed far more thoroughly than a multiple-choice item could manage.

That is the realistic gain. It is smaller than the marketing version and it survives contact with a stopwatch.

Method and limits

The 74-items-per-day figure comes from our FSRS simulation of a 2,000-concept collection; the 2.46× comes from scoring formats with a graded-response model. Both are described, with their assumptions, in the two posts linked above.

The 28 seconds per spoken item is an estimate rather than a measurement, and it carries the result. At 20 seconds a spoken item, the matched-information session drops to 10 minutes and speaking wins on elapsed time too. At 40 seconds it rises to 20 minutes. We would rather publish the sensitivity than a single confident number.

In KoiSwarm, Forgetting decides what is due and Live Talk is what you speak it into. Nothing in the scheduling changes based on which one you pick — the same item, the same interval, a different way of answering it.