Frugal tutoring
A large model can work a hard task out largely by itself. A model of 9 to 25 million settings cannot, and needs to be taught. This report reviews what human education research has established about teaching, checks which of it has held up in machine learning, and proposes one teaching framework for the platform.
Show, then coach on the learner's own attempts, then fade
- How the material is presented matters most. The largest effects in both literatures come from the form of the material, not the order of lessons. Our own arithmetic model went from 0.01 to 0.99 when digits were split differently, and three independent groups report the same kind of result for small transformers.
- Novices learn from worked examples, not from trial and error. This is one of the most robust findings in education, and it holds in machine learning: small models learn far more from imitating a good solver than from reward alone. Our E1 pair shows it directly: 20 rounds of reward left it choosing messages at chance, while imitation raised it from chance (0.51) to 0.94, close to the rule it copies (0.985).
- Feedback must say what the right step was, on the learner's own attempts. Tutors that give feedback at every step nearly match a human tutor, and bare right-or-wrong verification is weak or even harmful. A game score is exactly that kind of verification. Machine learning's version of good tutoring is to let the model act, then label the states it actually reached.
- Support must track the learner and fade. Guidance that helps a novice hurts an expert. The machine-learning equivalents are training on problems the model solves sometimes, never always or never, and reducing the teacher's weight as mastery grows.
- Fixed easy-to-hard lesson plans do not help. This is one idea from education that does not carry over. Controlled studies in machine learning find ordered curricula no better than random order when the budget is adequate. Our own staged-fields curriculum lost by 0.029.
- Score on held-out and transfer tests only, never on training performance. Varied practice and spacing lower training scores while raising later ones, and transfer to new situations is narrow unless it is taught. Our reader scores 0.97 in familiar formats and almost zero on a public benchmark.
Twenty techniques, and where each stops working
Averages hide a lot: nearly every meta-analysis reports wide spread, and effects shrink under stricter designs. The chart shows typical sizes to give a sense of scale. It is not a league table.
Sources: Rowland 2014; Yang et al. 2021; Latimier et al. 2021; Brunmair & Richter 2019; Belland et al. 2017; VanLehn 2011; Nickow et al. 2020; Wisniewski et al. 2020; Bisra et al. 2018; Alfieri et al. 2011 and 2013; Sinha & Kapur 2021; Bertsch et al. 2007; Kobayashi 2019.
| Technique | What it is | Evidence | When it fails |
|---|---|---|---|
| Retrieval practice | Make the learner produce the answer instead of re-reading it | Strong g ≈ 0.50 in the lab and in 222 classroom studies | Short delays; retrieval that fails with no feedback |
| Spacing | Spread practice out; the best gap grows with how long it must be remembered | Strong Spaced retrieval g = 0.74; uniform gaps work as well as expanding ones | Gaps far too long; it looks worse during training |
| Interleaving | Mix problem types instead of blocking them | Conditional g = 0.42 overall, 0.83 in a large maths trial | Reverses for word lists (−0.39); needs categories that are easy to confuse |
| Worked examples and fading | Study full solutions, then complete partial ones, then solve alone | Strong for novices Replicated since 1985 | Expertise reversal: the same guidance hurts learners who already know |
| Scaffolding | Temporary support contingent on what the learner can do now, then withdrawn | Strong Computer scaffolding in STEM g = 0.46 (144 studies) | Support that is fixed or never fades |
| Step-level tutoring | Feedback and hints on every step, not just on the final answer | Strong Step-based tutors d = 0.76 against 0.79 for human tutors; answer-only tutors far lower | Finer than step level gains nothing more |
| Feedback | Information that closes the gap between current and target performance | Conditional d ≈ 0.48 on average, but over a third of feedback interventions made performance worse | Bare right/wrong verification on complex tasks; feedback about the person, not the task |
| Mastery learning | Move on only after reaching a criterion, with correction in between | Overstated 0.52 on researcher tests; near zero on standardised tests. Bloom's "2 sigma" never replicated; about 0.33 is realistic | Tests not aligned with teaching |
| Self-explanation and generation | The learner explains or produces material instead of receiving it | Strong Self-explanation g = 0.55; generation d = 0.40 | Shallow or wrong explanations from novices |
| Discovery before instruction | Attempt a problem first, then be taught | Conditional With instruction after: g = 0.36. Pure discovery with no instruction: d = −0.38 | No consolidation afterwards; learners without the prerequisites |
| Contrasting cases and variation | Compare cases that differ in one feature while the rest stays fixed | Strong Case comparison d = 0.50; comparing two cases transfers where studying them separately does not | Cases that differ in many ways at once |
| Desirable difficulties | Conditions that slow learning now and improve it later | Conditional Robust as a principle; applied effects small (0.19 against 0.57 in the lab) | A difficulty the learner cannot yet overcome is just a difficulty |
| Whole-task design (4C/ID) | Teach whole tasks from simple to complex versions; drill parts only when they must become automatic | Promising d = 0.79 in few, weakly controlled studies | Decomposing a skill into parts taught separately loses the coordination between them |
| Transfer | Using a skill in a new setting | Rare by default Far-transfer effects vanish under good controls | It must be taught (comparison, varied examples, stated principles) and tested at a named distance |
| Learning by teaching | Explaining to others | Moderate g = 0.56 after actually teaching | Groups recall less than the same people pooled (collaborative inhibition) |
| Confidence judgments | How well a learner knows what it knows | Strong Delayed judgments are far more accurate than immediate ones | Confidence measured straight after practice is inflated |
Some ideas transfer strongly, one does not transfer at all
Each row pairs a human finding with its machine-learning counterpart, the evidence for small models, and what this programme has already measured. Most large-model results (1.5–70 billion settings) are an inference at our scale; the small-model evidence is marked.
| Human principle | Machine-learning counterpart | Evidence in ML | What we have measured |
|---|---|---|---|
| Reduce needless load; present material well | Tokenisation, output order, position markers | Strongest, small models Reversed digits and scratchpads let tiny transformers learn addition from scratch; abacus embeddings reach 99% on 100-digit sums | Digit splitting: arithmetic 0.01 → 0.99. A bigger vocabulary made copying worse (0.08 → 0.03) |
| Pitch material at the learner's level | Data simplified and curated to the model's capacity | Strong, small models Models under 10 M write fluent stories when the vocabulary is a small child's; models under 3 B learn less from long reasoning chains than from short ones | Our narrow, verifiable tasks already follow this |
| Worked examples for novices | Imitation and distillation before reinforcement | Strong Distilling a strong model into a 32-billion-setting base beat large-scale reward training of the same base on maths (72.6 against 47.0) | E1: imitation lifted message choice from 0.51 to 0.94; reward alone reached 0.52–0.555 |
| Step-level feedback, on the learner's own work | On-policy labelling (DAgger), on-policy distillation, process supervision | Strong Training on the expert's states lets errors compound; labelling the learner's own states fixes it. On-policy distillation doubled gains over ordinary distillation | Copy supervision, which supervises how a value is produced: +0.087. E1 imitation already labels the model's own games |
| Contingent scaffolding at the edge of ability | Train on problems the model solves sometimes | Strong A problem the model always or never solves gives exactly zero learning signal in group-relative training; filtering to intermediate pass rates speeds learning | Not used yet. Not measured |
| Fading and expertise reversal | Anneal the teacher's weight; move from imitation toward the model's own values | Moderate "Kickstarting" reached from-scratch performance in about a tenth of the steps and then passed the teacher | Not used yet |
| Variation and contrasting cases | Augmentation; minimal pairs that differ only in the cause | Strong Networks learn the cheapest rule that fits; counterfactual pairs break it | Altered photos +0.077; varied field names 0.013 → 0.278; a picture set solvable by pixel count |
| Spacing and retrieval | Replay old material while learning new | Strong against forgetting About 5% replay matched full retraining in continued pre-training. Explicit spaced schedules: one paper, not replicated | Not used. Staged fine-tunes could forget earlier stages |
| Performance during training is not learning | Held-out selection; delayed generalisation | Strong Small models can memorise and generalise much later; weight decay matters | A curriculum arm read 0.8301 mid-run and finished at 0.7442; picture scores rose again from 100 to 300 epochs |
| Transfer is narrow unless taught | Out-of-distribution tests at named distances | Strong | 0.97 in familiar formats, near zero on the benchmark; renamed fields 0.013 |
| Fixed simple-to-complex sequence | Curriculum learning | Does not carry over Thousands of orderings: random order did as well; the BabyLM small-model challenges found curricula "largely unsuccessful" | Staged fields lost by 0.029 (2.9 times the noise) and the default was retired |
| Whole task over separate parts | One model over a chain of part-models | Consistent, not tested directly | A five-model chain lost 85% of its last model's skill to upstream errors; the team's win came from one specialist |
| Learning by teaching | Models teaching models | Preliminary Large models only, in-context | Not tested |
Read as a teacher would read them
The rule is a hand-written policy that tells the most useful fact. Imitation is the G2 line (22 rounds so far, still rising). The reward pilots are the six S1b trainings after 20 rounds.
E1, communication. Six finer rewards, trained for 20 rounds with a safe step size, all left the model choosing messages at about chance (0.52–0.555). In teaching terms, a game score is verification feedback on a complex task: it says how the game went, not which message was right, and the messages explain only 1–6% of the variation in the score. That is the kind of feedback the education literature finds weakest. Imitation gives the model a worked example of the right step, on the positions it actually reached, and it works.
Three gaps follow from the review:
- The feedback covers only the move made. Every reward so far scored the one message the model sent. Code can score every message it could have sent, so each position can be labelled with the value of every option. That is step-level tutoring at full resolution, and it is also how a model can learn to do better than the rule.
- Asking was never taught. 65% of the model's questions are unanswerable. That is what random asking would give (about 60%, since the partner sees each cell with probability 0.4), so asking is simply untrained. Imitation covered telling only.
- There is no fading. Imitation stays imitation. Nothing hands control back to the model's own judgement, so its ceiling is the rule. The rule scores about 0.965 on the stricter puzzles against about 0.993 for a player who sees everything, so there is some room above it.
E2, extraction and transfer. The reader has mastered its trained formats (0.97) and transfers poorly: renamed fields fall to 0.28 after the first redesign, and 51–68% of every model's errors are invented values. Three things stand out:
- Field names are learned as surface features. Knowledge tied to a surface is not retrieved when the surface changes, which is the textbook account of failed transfer. Variation and short descriptions (the second redesign, now training) are the evidence-backed remedy. Minimal pairs, where only the field name changes, would add the contrast the literature finds most effective.
- Invented values are never corrected on the model's own output. The model trains on the correct answers, never on its own mistakes with the correct value beside them. That is the behaviour-cloning gap: errors appear in situations training never showed.
- The worked example has no steps. Every training answer begins with the same fixed sentence ("Reasoning: extract each requested field from the document."). It shows no step. A short trace saying where each value was found would be a worked example. It must stay short: small models learn less from long reasoning.
Frugal tutoring: five stages, repeated per skill
A skill is one thing the model must do: extract a type of field, tell a useful fact, ask a useful question. Each skill goes through the loop. The loop repeats until the model's mastery on held-out items stops improving.
Put the material in a form the model can learn, check it is learnable, and vary everything that could be used to cheat.
Worked examples: a teacher's correct choices, with a short trace of the step where one exists.
The model attempts; the tutor labels the positions it actually reached, with the value of every option.
Reduce the teacher's weight as mastery grows; practise on items it solves sometimes; replay earlier skills.
Sealed held-out tests, transfer tests at named distances, and confidence measured on held-out items.
1Prepare the material
- Does
- Chooses the representation (tokenisation, output order, how numbers are written) so each output token can be computed from what comes before it. Checks the target can be learned at all from what the model sees. Varies every cheap statistic that could stand in for the answer. Adds minimal pairs that differ only in the thing that matters.
- Human basis
- Cognitive load (remove needless load); variation theory and contrasting cases (d = 0.50).
- ML basis
- Arithmetic formatting results for small transformers; TinyStories; shortcut learning and counterfactual pairs.
- On the platform
- Mostly exists: the digit-safe tokenizer, the shortcut-ceiling audit, the learnability check (mutual information between target and input), and augmentation. To build: one "lesson readiness" report that runs them together before any training is paid for, and a minimal-pair generator for field names.
2Show worked examples
- Does
- Trains on a teacher's correct choices. For extraction the teacher is the labelled data plus a short located-evidence trace. For E1 it is the rule, for telling and, new, for asking (ask about the unseen cell with the highest expected value to the solver).
- Human basis
- Worked-example effect for novices; modelling in cognitive apprenticeship.
- ML basis
- Imitation before reinforcement; short rationales as extra targets let a 770 M model beat a 540 B one; a teacher too far ahead hurts, so keep traces short.
- On the platform
- Exists: supervised training, the LLM teacher, E1's imitation credit. To build: a short trace in place of the fixed "Reasoning:" sentence (tested before it becomes a default), and an asking rule for E1.
3Coach on the model's own attempts
- Does
- Runs the current model, then labels each position it reached. Where code can score every option (E1's tells and asks), the label is the value of each option and the model is trained toward that whole distribution. Where only the right answer is known (extraction), the model's own wrong outputs on training documents become examples with the correct value.
- Human basis
- Step-based tutoring (0.76, near a human tutor); informative, task-level feedback, not verification.
- ML basis
- DAgger; on-policy distillation; cost-to-go imitation, which can exceed a weaker teacher; COMA-style counterfactual credit.
- On the platform
- Partly exists: E1 imitation already labels the model's own games; the
improveloop and human corrections feed extraction. To build: atutorcredit for E1 (values of all legal acts), and an automatic self-correction step for extraction.
4Fade the help, keep practice at the edge
- Does
- Lowers the teacher's weight as held-out mastery rises, moving from the rule's choices toward the model's own values. Picks practice items the model solves some of the time and drops those it always or never solves. Replays about 5% of earlier material when a new skill is added.
- Human basis
- Contingent scaffolding and fading; expertise reversal; spacing and retrieval.
- ML basis
- Kickstarting with an annealed teacher weight; KL anchoring to the imitation policy; filtering to intermediate difficulty; replay against forgetting.
- On the platform
- Partly exists: E1's trust region and rollback, the stricter-puzzle pool, hard-case mining. To build: a fading schedule tied to mastery, pass-rate item selection with a stall alarm when most groups carry no signal, and replay in sequential fine-tunes.
5Prove it
- Does
- Judges only on sealed held-out items, never on training performance. Tests transfer at named distances: new wording, new layouts, new documents, the public benchmark. Measures confidence on held-out items.
- Human basis
- Desirable difficulties (training scores mislead); transfer is narrow; delayed judgments of learning are the accurate ones.
- ML basis
- Held-out selection; delayed generalisation after memorisation; shortcut learning.
- On the platform
- Exists and strong: sealed test splits, trivial floors, shortcut ceilings, the calibration gate, the name-shift probes. To build: one transfer ladder per task, so every run reports how far its skill travels.
What the framework rules out
Random order did as well in controlled studies, and ours lost. Order practice by the model's current pass rate instead.
Use it after imitation and coaching, anchored to them, never instead of them.
Keep traces short and targets within reach; use an intermediate teacher when the gap is large.
Coordination is part of the skill. Split only where a part must be automatic or is truly separate.
Varied practice and spacing make training scores fall while learning improves.
They reward change, not usefulness. Use only shaping that leaves the best policy unchanged.
What to add, and where it lives
| Component | What it does | Where | Size |
|---|---|---|---|
| Lesson readiness report | Runs the learnability check, shortcut ceiling, representation check and variation audit before training, and blocks a run whose target cannot be learned | Pre-flight gate; a prepare job stage | Small: assembles existing checks |
| Worked-step traces | Replaces the fixed "Reasoning:" sentence with a short trace of where each value was found | Data builders and the teacher | Small, behind a switch |
| E1 tutor credit | Labels every position with the exact value of every legal tell and ask, computed by code; trains toward that distribution; anchored to the imitation model with an annealed weight | E1's credit rules and pilot | Medium |
| E1 asking rule | Asks about the cell with the highest expected value to the solver, weighted by the chance the partner can see it | E1's code controls | Small |
| Self-correction for extraction | Runs the model on training documents, turns its wrong or invented values into examples with the correct value, and fine-tunes on them with replay | A tutor stage beside improve | Medium |
| Mastery tracker and fading schedule | Tracks held-out mastery per skill and lowers the teacher's weight as it rises | Training engine; recorded in the results ledger | Medium |
| Edge-of-ability selection | Chooses practice items by the model's pass rate; raises a stall alarm when most groups carry no signal | E1 pilot first, then the improve loop | Small to medium |
| Transfer ladder | A fixed set of transfer tests per task, reported on every run beside the held-out score | Eval stage and the Results tab | Medium |
| Lesson plan view | Shows each skill's stage and mastery on the run's Overview, in plain words | Control room | Medium, after the pilots |
Four pilots, each registered before it runs
| Pilot | Question | Arms, at equal compute | Would count as a pass |
|---|---|---|---|
| T1 · E1 coaching | Does full-option coaching with fading beat imitation, and pass the rule? | Continue imitation · tutor credit with annealed anchor · tutor credit plus the asking rule, all from G2's final model on the stricter puzzles | Message choice and score above imitation on the same puzzles; then score above the rule by more than the measured noise |
| T2 · E2 self-correction | Does training on its own corrected mistakes cut invented values? | More epochs of the same data · self-correction with 5% replay | Fewer invented values and higher held-out accuracy, with no loss on trained names |
| T3 · Worked-step traces | Does a short located-evidence trace help a small reader? | Fixed sentence · short trace | Higher held-out and renamed-field accuracy; parse rate unchanged |
| T4 · Edge-of-ability selection | Does choosing items by pass rate speed learning? | Uniform items · pass-rate-band items | Same final score in fewer rounds, or a higher score in the same rounds |
Order. T1 first, because E1 is where the gap is largest and the tutor can score every option exactly. It starts after G2's read (about 2–3 October), since G2 decides whether the model can learn to tell at all. T2 and T3 can run on the E2 agents once the varied-names retraining finishes. T4 rides inside T1 at little extra cost.
Cost. T1 is CPU work of the same size as the reward pilots. T2 and T3 are GPU fine-tunes of the small agents, each about a day. None needs a larger model.
Decided on 1 October
- Adopted: frugal tutoring is the programme's default training framework. Fixed curricula and reward-only training for small models are ruled out as defaults, and E1 and E2 move onto the framework (the migration plan keeps, retires or replaces each part).
- Approved: T1, E1 coaching, to be registered after G2's read.
- Approved: T2 and T3, E2 self-correction and worked-step traces, after the varied-names retraining.
- Agreed build order: the lesson readiness report and E1 tutor credit first; the control-room lesson view last.
- Running jobs: all four running jobs are stages the framework keeps, so none was stopped.
Human learning. Dunlosky et al. 2013, Psych. Sci. Public Interest. Rowland 2014, Psych. Bull. Yang et al. 2021, Psych. Bull. Cepeda et al. 2006, 2008. Latimier et al. 2021, Ed. Psych. Rev. Brunmair & Richter 2019, Psych. Bull. Atkinson et al. 2000; Kalyuga et al. 2003 (expertise reversal). van de Pol et al. 2010; Belland et al. 2017 (scaffolding). VanLehn 2011, Ed. Psychologist. Nickow et al. 2020, NBER. Kluger & DeNisi 1996; Wisniewski et al. 2020; Shute 2008 (feedback). Kulik et al. 1990; Slavin 1987; von Hippel 2024 (mastery, 2 sigma). Bisra et al. 2018; Chi & Wylie 2014. Sinha & Kapur 2021; Alfieri et al. 2011; Kirschner, Sweller & Clark 2006. Alfieri et al. 2013; Gentner et al. 2003 (comparison). Barnett & Ceci 2002; Sala & Gobet 2017 (transfer). van Merriënboer's 4C/ID. Macnamara et al. 2014. Rhodes & Tauber 2011. Marion & Thorley 2016.
Machine learning. Wu, Dyer & Neyshabur 2021 (arXiv:2012.03107); BabyLM findings 2023 and 2024 (arXiv:2412.05149). Lee et al. 2024, teaching arithmetic to small transformers (arXiv:2307.03381); McLeish et al. 2024 (arXiv:2405.17399). Eldan & Li 2023, TinyStories (arXiv:2305.07759). Li et al. 2025, small models and strong reasoners (arXiv:2502.12143). Hsieh et al. 2023 (arXiv:2305.02301). Ross et al. 2011, DAgger (arXiv:1011.0686); Agarwal et al. 2024, on-policy distillation (arXiv:2306.13649). DeepSeek-R1 2025 (arXiv:2501.12948). Jiang et al. 2021, PLR (arXiv:2010.03934); Yu et al. 2025, DAPO (arXiv:2503.14476). Kaushik et al. 2020 (arXiv:1909.12434). Ibrahim et al. 2024 (arXiv:2403.08763). Power et al. 2022, grokking (arXiv:2201.02177).
Communication. Lowe et al. 2020, supervision and self-play (arXiv:2002.01093). Lowe et al. 2019, pitfalls of measuring communication (arXiv:1903.05168). Lewis et al. 2017, Deal or No Deal (arXiv:1706.05125). Lu et al. 2020, seeded iterated learning (arXiv:2003.12694). Ross & Bagnell 2014, AggreVaTe (arXiv:1406.5979); Sun et al. 2017 (arXiv:1703.01030). Anthony et al. 2017, expert iteration (arXiv:1705.08439). Schmitt et al. 2018, kickstarting (arXiv:1803.03835). Foerster et al. 2018, COMA (arXiv:1705.08926). Ng et al. 1999, reward shaping. Rao & Daumé 2018; Grand et al. 2024 (asking by information value). Lin et al. 2024, DialOp (arXiv:2305.20076).