Frugal Tutoring
Research report and proposal · Frugal programme

Frugal tutoring

A large model can work a hard task out largely by itself. A model of 9 to 25 million settings cannot, and needs to be taught. This report reviews what human education research has established about teaching, checks which of it has held up in machine learning, and proposes one teaching framework for the platform.

1 October 2026 · Built from three literature reviews (about 150 sources, key effect sizes checked against the papers) and the programme's own measurements. Human effect sizes are standardised mean differences (0.2 small, 0.5 medium, 0.8 large).

1 · The short answer

Show, then coach on the learner's own attempts, then fade

  1. How the material is presented matters most. The largest effects in both literatures come from the form of the material, not the order of lessons. Our own arithmetic model went from 0.01 to 0.99 when digits were split differently, and three independent groups report the same kind of result for small transformers.
  2. Novices learn from worked examples, not from trial and error. This is one of the most robust findings in education, and it holds in machine learning: small models learn far more from imitating a good solver than from reward alone. Our E1 pair shows it directly: 20 rounds of reward left it choosing messages at chance, while imitation raised it from chance (0.51) to 0.94, close to the rule it copies (0.985).
  3. Feedback must say what the right step was, on the learner's own attempts. Tutors that give feedback at every step nearly match a human tutor, and bare right-or-wrong verification is weak or even harmful. A game score is exactly that kind of verification. Machine learning's version of good tutoring is to let the model act, then label the states it actually reached.
  4. Support must track the learner and fade. Guidance that helps a novice hurts an expert. The machine-learning equivalents are training on problems the model solves sometimes, never always or never, and reducing the teacher's weight as mastery grows.
  5. Fixed easy-to-hard lesson plans do not help. This is one idea from education that does not carry over. Controlled studies in machine learning find ordered curricula no better than random order when the budget is adequate. Our own staged-fields curriculum lost by 0.029.
  6. Score on held-out and transfer tests only, never on training performance. Varied practice and spacing lower training scores while raising later ones, and transfer to new situations is narrow unless it is taught. Our reader scores 0.97 in familiar formats and almost zero on a public benchmark.
The proposal in one lineA teaching loop for every skill: Prepare the material, Show worked examples, Coach the model on its own attempts with step-level feedback, Fade the help while keeping practice at the edge of its ability, and Prove it on sealed and transfer tests. Most of the parts exist on the platform. The missing ones are an on-policy coaching step, a fading schedule, and difficulty-aware selection of practice items.
2 · What human learning science has established

Twenty techniques, and where each stops working

Averages hide a lot: nearly every meta-analysis reports wide spread, and effects shrink under stricter designs. The chart shows typical sizes to give a sense of scale. It is not a league table.

Typical effect of each techniqueMeta-analytic averages. Bars left of zero mean the technique did worse than its comparison.
well replicateddepends strongly on conditionsthe comparison that fails

Sources: Rowland 2014; Yang et al. 2021; Latimier et al. 2021; Brunmair & Richter 2019; Belland et al. 2017; VanLehn 2011; Nickow et al. 2020; Wisniewski et al. 2020; Bisra et al. 2018; Alfieri et al. 2011 and 2013; Sinha & Kapur 2021; Bertsch et al. 2007; Kobayashi 2019.

TechniqueWhat it isEvidenceWhen it fails
Retrieval practiceMake the learner produce the answer instead of re-reading itStrong
g ≈ 0.50 in the lab and in 222 classroom studies
Short delays; retrieval that fails with no feedback
SpacingSpread practice out; the best gap grows with how long it must be rememberedStrong
Spaced retrieval g = 0.74; uniform gaps work as well as expanding ones
Gaps far too long; it looks worse during training
InterleavingMix problem types instead of blocking themConditional
g = 0.42 overall, 0.83 in a large maths trial
Reverses for word lists (−0.39); needs categories that are easy to confuse
Worked examples and fadingStudy full solutions, then complete partial ones, then solve aloneStrong for novices
Replicated since 1985
Expertise reversal: the same guidance hurts learners who already know
ScaffoldingTemporary support contingent on what the learner can do now, then withdrawnStrong
Computer scaffolding in STEM g = 0.46 (144 studies)
Support that is fixed or never fades
Step-level tutoringFeedback and hints on every step, not just on the final answerStrong
Step-based tutors d = 0.76 against 0.79 for human tutors; answer-only tutors far lower
Finer than step level gains nothing more
FeedbackInformation that closes the gap between current and target performanceConditional
d ≈ 0.48 on average, but over a third of feedback interventions made performance worse
Bare right/wrong verification on complex tasks; feedback about the person, not the task
Mastery learningMove on only after reaching a criterion, with correction in betweenOverstated
0.52 on researcher tests; near zero on standardised tests. Bloom's "2 sigma" never replicated; about 0.33 is realistic
Tests not aligned with teaching
Self-explanation and generationThe learner explains or produces material instead of receiving itStrong
Self-explanation g = 0.55; generation d = 0.40
Shallow or wrong explanations from novices
Discovery before instructionAttempt a problem first, then be taughtConditional
With instruction after: g = 0.36. Pure discovery with no instruction: d = −0.38
No consolidation afterwards; learners without the prerequisites
Contrasting cases and variationCompare cases that differ in one feature while the rest stays fixedStrong
Case comparison d = 0.50; comparing two cases transfers where studying them separately does not
Cases that differ in many ways at once
Desirable difficultiesConditions that slow learning now and improve it laterConditional
Robust as a principle; applied effects small (0.19 against 0.57 in the lab)
A difficulty the learner cannot yet overcome is just a difficulty
Whole-task design (4C/ID)Teach whole tasks from simple to complex versions; drill parts only when they must become automaticPromising
d = 0.79 in few, weakly controlled studies
Decomposing a skill into parts taught separately loses the coordination between them
TransferUsing a skill in a new settingRare by default
Far-transfer effects vanish under good controls
It must be taught (comparison, varied examples, stated principles) and tested at a named distance
Learning by teachingExplaining to othersModerate
g = 0.56 after actually teaching
Groups recall less than the same people pooled (collaborative inhibition)
Confidence judgmentsHow well a learner knows what it knowsStrong
Delayed judgments are far more accurate than immediate ones
Confidence measured straight after practice is inflated
Weak or debunked: learning styles; the "learning pyramid"; Bloom's 2 sigma as a real effect size; 10,000 hours of deliberate practice as the cause of expertise (it explains 4% of the variation in education and under 1% in professions); brain training and far transfer from chess or music; ranking interventions by average effect size.
3 · What carries over to machine learning

Some ideas transfer strongly, one does not transfer at all

Each row pairs a human finding with its machine-learning counterpart, the evidence for small models, and what this programme has already measured. Most large-model results (1.5–70 billion settings) are an inference at our scale; the small-model evidence is marked.

Human principleMachine-learning counterpartEvidence in MLWhat we have measured
Reduce needless load; present material wellTokenisation, output order, position markersStrongest, small models
Reversed digits and scratchpads let tiny transformers learn addition from scratch; abacus embeddings reach 99% on 100-digit sums
Digit splitting: arithmetic 0.01 → 0.99. A bigger vocabulary made copying worse (0.08 → 0.03)
Pitch material at the learner's levelData simplified and curated to the model's capacityStrong, small models
Models under 10 M write fluent stories when the vocabulary is a small child's; models under 3 B learn less from long reasoning chains than from short ones
Our narrow, verifiable tasks already follow this
Worked examples for novicesImitation and distillation before reinforcementStrong
Distilling a strong model into a 32-billion-setting base beat large-scale reward training of the same base on maths (72.6 against 47.0)
E1: imitation lifted message choice from 0.51 to 0.94; reward alone reached 0.52–0.555
Step-level feedback, on the learner's own workOn-policy labelling (DAgger), on-policy distillation, process supervisionStrong
Training on the expert's states lets errors compound; labelling the learner's own states fixes it. On-policy distillation doubled gains over ordinary distillation
Copy supervision, which supervises how a value is produced: +0.087. E1 imitation already labels the model's own games
Contingent scaffolding at the edge of abilityTrain on problems the model solves sometimesStrong
A problem the model always or never solves gives exactly zero learning signal in group-relative training; filtering to intermediate pass rates speeds learning
Not used yet. Not measured
Fading and expertise reversalAnneal the teacher's weight; move from imitation toward the model's own valuesModerate
"Kickstarting" reached from-scratch performance in about a tenth of the steps and then passed the teacher
Not used yet
Variation and contrasting casesAugmentation; minimal pairs that differ only in the causeStrong
Networks learn the cheapest rule that fits; counterfactual pairs break it
Altered photos +0.077; varied field names 0.013 → 0.278; a picture set solvable by pixel count
Spacing and retrievalReplay old material while learning newStrong against forgetting
About 5% replay matched full retraining in continued pre-training. Explicit spaced schedules: one paper, not replicated
Not used. Staged fine-tunes could forget earlier stages
Performance during training is not learningHeld-out selection; delayed generalisationStrong
Small models can memorise and generalise much later; weight decay matters
A curriculum arm read 0.8301 mid-run and finished at 0.7442; picture scores rose again from 100 to 300 epochs
Transfer is narrow unless taughtOut-of-distribution tests at named distancesStrong0.97 in familiar formats, near zero on the benchmark; renamed fields 0.013
Fixed simple-to-complex sequenceCurriculum learningDoes not carry over
Thousands of orderings: random order did as well; the BabyLM small-model challenges found curricula "largely unsuccessful"
Staged fields lost by 0.029 (2.9 times the noise) and the default was retired
Whole task over separate partsOne model over a chain of part-modelsConsistent, not tested directlyA five-model chain lost 85% of its last model's skill to upstream errors; the team's win came from one specialist
Learning by teachingModels teaching modelsPreliminary
Large models only, in-context
Not tested
4 · Why our small models are stuck

Read as a teacher would read them

E1: reward alone against being shownHow well the model picks which fact to tell (0.50 is random, 1.0 always picks the most useful), on the same validation puzzles

The rule is a hand-written policy that tells the most useful fact. Imitation is the G2 line (22 rounds so far, still rising). The reward pilots are the six S1b trainings after 20 rounds.

E1, communication. Six finer rewards, trained for 20 rounds with a safe step size, all left the model choosing messages at about chance (0.52–0.555). In teaching terms, a game score is verification feedback on a complex task: it says how the game went, not which message was right, and the messages explain only 1–6% of the variation in the score. That is the kind of feedback the education literature finds weakest. Imitation gives the model a worked example of the right step, on the positions it actually reached, and it works.

Three gaps follow from the review:

  • The feedback covers only the move made. Every reward so far scored the one message the model sent. Code can score every message it could have sent, so each position can be labelled with the value of every option. That is step-level tutoring at full resolution, and it is also how a model can learn to do better than the rule.
  • Asking was never taught. 65% of the model's questions are unanswerable. That is what random asking would give (about 60%, since the partner sees each cell with probability 0.4), so asking is simply untrained. Imitation covered telling only.
  • There is no fading. Imitation stays imitation. Nothing hands control back to the model's own judgement, so its ceiling is the rule. The rule scores about 0.965 on the stricter puzzles against about 0.993 for a player who sees everything, so there is some room above it.

E2, extraction and transfer. The reader has mastered its trained formats (0.97) and transfers poorly: renamed fields fall to 0.28 after the first redesign, and 51–68% of every model's errors are invented values. Three things stand out:

  • Field names are learned as surface features. Knowledge tied to a surface is not retrieved when the surface changes, which is the textbook account of failed transfer. Variation and short descriptions (the second redesign, now training) are the evidence-backed remedy. Minimal pairs, where only the field name changes, would add the contrast the literature finds most effective.
  • Invented values are never corrected on the model's own output. The model trains on the correct answers, never on its own mistakes with the correct value beside them. That is the behaviour-cloning gap: errors appear in situations training never showed.
  • The worked example has no steps. Every training answer begins with the same fixed sentence ("Reasoning: extract each requested field from the document."). It shows no step. A short trace saying where each value was found would be a worked example. It must stay short: small models learn less from long reasoning.
5 · The framework

Frugal tutoring: five stages, repeated per skill

A skill is one thing the model must do: extract a type of field, tell a useful fact, ask a useful question. Each skill goes through the loop. The loop repeats until the model's mastery on held-out items stops improving.

stage 1Prepare

Put the material in a form the model can learn, check it is learnable, and vary everything that could be used to cheat.

stage 2Show

Worked examples: a teacher's correct choices, with a short trace of the step where one exists.

stage 3Coach

The model attempts; the tutor labels the positions it actually reached, with the value of every option.

stage 4Fade

Reduce the teacher's weight as mastery grows; practise on items it solves sometimes; replay earlier skills.

stage 5Prove

Sealed held-out tests, transfer tests at named distances, and confidence measured on held-out items.

repeat stages 3–5 while held-out mastery still rises · a stall sends the skill back to stage 1

1Prepare the material

Does
Chooses the representation (tokenisation, output order, how numbers are written) so each output token can be computed from what comes before it. Checks the target can be learned at all from what the model sees. Varies every cheap statistic that could stand in for the answer. Adds minimal pairs that differ only in the thing that matters.
Human basis
Cognitive load (remove needless load); variation theory and contrasting cases (d = 0.50).
ML basis
Arithmetic formatting results for small transformers; TinyStories; shortcut learning and counterfactual pairs.
On the platform
Mostly exists: the digit-safe tokenizer, the shortcut-ceiling audit, the learnability check (mutual information between target and input), and augmentation. To build: one "lesson readiness" report that runs them together before any training is paid for, and a minimal-pair generator for field names.

2Show worked examples

Does
Trains on a teacher's correct choices. For extraction the teacher is the labelled data plus a short located-evidence trace. For E1 it is the rule, for telling and, new, for asking (ask about the unseen cell with the highest expected value to the solver).
Human basis
Worked-example effect for novices; modelling in cognitive apprenticeship.
ML basis
Imitation before reinforcement; short rationales as extra targets let a 770 M model beat a 540 B one; a teacher too far ahead hurts, so keep traces short.
On the platform
Exists: supervised training, the LLM teacher, E1's imitation credit. To build: a short trace in place of the fixed "Reasoning:" sentence (tested before it becomes a default), and an asking rule for E1.

3Coach on the model's own attempts

Does
Runs the current model, then labels each position it reached. Where code can score every option (E1's tells and asks), the label is the value of each option and the model is trained toward that whole distribution. Where only the right answer is known (extraction), the model's own wrong outputs on training documents become examples with the correct value.
Human basis
Step-based tutoring (0.76, near a human tutor); informative, task-level feedback, not verification.
ML basis
DAgger; on-policy distillation; cost-to-go imitation, which can exceed a weaker teacher; COMA-style counterfactual credit.
On the platform
Partly exists: E1 imitation already labels the model's own games; the improve loop and human corrections feed extraction. To build: a tutor credit for E1 (values of all legal acts), and an automatic self-correction step for extraction.

4Fade the help, keep practice at the edge

Does
Lowers the teacher's weight as held-out mastery rises, moving from the rule's choices toward the model's own values. Picks practice items the model solves some of the time and drops those it always or never solves. Replays about 5% of earlier material when a new skill is added.
Human basis
Contingent scaffolding and fading; expertise reversal; spacing and retrieval.
ML basis
Kickstarting with an annealed teacher weight; KL anchoring to the imitation policy; filtering to intermediate difficulty; replay against forgetting.
On the platform
Partly exists: E1's trust region and rollback, the stricter-puzzle pool, hard-case mining. To build: a fading schedule tied to mastery, pass-rate item selection with a stall alarm when most groups carry no signal, and replay in sequential fine-tunes.

5Prove it

Does
Judges only on sealed held-out items, never on training performance. Tests transfer at named distances: new wording, new layouts, new documents, the public benchmark. Measures confidence on held-out items.
Human basis
Desirable difficulties (training scores mislead); transfer is narrow; delayed judgments of learning are the accurate ones.
ML basis
Held-out selection; delayed generalisation after memorisation; shortcut learning.
On the platform
Exists and strong: sealed test splits, trivial floors, shortcut ceilings, the calibration gate, the name-shift probes. To build: one transfer ladder per task, so every run reports how far its skill travels.

What the framework rules out

A fixed easy-to-hard lesson order

Random order did as well in controlled studies, and ours lost. Order practice by the model's current pass rate instead.

Reward as the only teacher for a small model

Use it after imitation and coaching, anchored to them, never instead of them.

A teacher far ahead of the student

Keep traces short and targets within reach; use an intermediate teacher when the gap is large.

Splitting one skill across a chain of models by default

Coordination is part of the skill. Split only where a part must be automatic or is truly separate.

Judging on training scores

Varied practice and spacing make training scores fall while learning improves.

Bonuses for variety or influence

They reward change, not usefulness. Use only shaping that leaves the best policy unchanged.

6 · Building it on the platform

What to add, and where it lives

ComponentWhat it doesWhereSize
Lesson readiness reportRuns the learnability check, shortcut ceiling, representation check and variation audit before training, and blocks a run whose target cannot be learnedPre-flight gate; a prepare job stageSmall: assembles existing checks
Worked-step tracesReplaces the fixed "Reasoning:" sentence with a short trace of where each value was foundData builders and the teacherSmall, behind a switch
E1 tutor creditLabels every position with the exact value of every legal tell and ask, computed by code; trains toward that distribution; anchored to the imitation model with an annealed weightE1's credit rules and pilotMedium
E1 asking ruleAsks about the cell with the highest expected value to the solver, weighted by the chance the partner can see itE1's code controlsSmall
Self-correction for extractionRuns the model on training documents, turns its wrong or invented values into examples with the correct value, and fine-tunes on them with replayA tutor stage beside improveMedium
Mastery tracker and fading scheduleTracks held-out mastery per skill and lowers the teacher's weight as it risesTraining engine; recorded in the results ledgerMedium
Edge-of-ability selectionChooses practice items by the model's pass rate; raises a stall alarm when most groups carry no signalE1 pilot first, then the improve loopSmall to medium
Transfer ladderA fixed set of transfer tests per task, reported on every run beside the held-out scoreEval stage and the Results tabMedium
Lesson plan viewShows each skill's stage and mastery on the run's Overview, in plain wordsControl roomMedium, after the pilots
The platform's rules stay in force. Every new component is off by default, arrives with a registered test against a control at equal compute, and becomes a default only after it passes. The teacher reaches training only as data, never the scored path.
7 · Proving the framework

Four pilots, each registered before it runs

PilotQuestionArms, at equal computeWould count as a pass
T1 · E1 coachingDoes full-option coaching with fading beat imitation, and pass the rule?Continue imitation · tutor credit with annealed anchor · tutor credit plus the asking rule, all from G2's final model on the stricter puzzlesMessage choice and score above imitation on the same puzzles; then score above the rule by more than the measured noise
T2 · E2 self-correctionDoes training on its own corrected mistakes cut invented values?More epochs of the same data · self-correction with 5% replayFewer invented values and higher held-out accuracy, with no loss on trained names
T3 · Worked-step tracesDoes a short located-evidence trace help a small reader?Fixed sentence · short traceHigher held-out and renamed-field accuracy; parse rate unchanged
T4 · Edge-of-ability selectionDoes choosing items by pass rate speed learning?Uniform items · pass-rate-band itemsSame final score in fewer rounds, or a higher score in the same rounds

Order. T1 first, because E1 is where the gap is largest and the tutor can score every option exactly. It starts after G2's read (about 2–3 October), since G2 decides whether the model can learn to tell at all. T2 and T3 can run on the E2 agents once the varied-names retraining finishes. T4 rides inside T1 at little extra cost.

Cost. T1 is CPU work of the same size as the reward pilots. T2 and T3 are GPU fine-tunes of the small agents, each about a day. None needs a larger model.

8 · Decisions

Decided on 1 October

  1. Adopted: frugal tutoring is the programme's default training framework. Fixed curricula and reward-only training for small models are ruled out as defaults, and E1 and E2 move onto the framework (the migration plan keeps, retires or replaces each part).
  2. Approved: T1, E1 coaching, to be registered after G2's read.
  3. Approved: T2 and T3, E2 self-correction and worked-step traces, after the varied-names retraining.
  4. Agreed build order: the lesson readiness report and E1 tutor credit first; the control-room lesson view last.
  5. Running jobs: all four running jobs are stages the framework keeps, so none was stopped.
One confirmation still neededTeaching E1's model what to say means its communication is installed, not discovered. E1's claims therefore become taught and beyond the teacher. A claim that good communication arises without teaching stays possible only from an untaught comparison arm.
What this cannot promiseMost machine-learning evidence for coaching and fading comes from models a hundred to a thousand times larger than ours. Human effect sizes are averages with wide spread. The framework is a strong prior, not a result: each stage enters the platform only after its own pilot passes.
Key sources

Human learning. Dunlosky et al. 2013, Psych. Sci. Public Interest. Rowland 2014, Psych. Bull. Yang et al. 2021, Psych. Bull. Cepeda et al. 2006, 2008. Latimier et al. 2021, Ed. Psych. Rev. Brunmair & Richter 2019, Psych. Bull. Atkinson et al. 2000; Kalyuga et al. 2003 (expertise reversal). van de Pol et al. 2010; Belland et al. 2017 (scaffolding). VanLehn 2011, Ed. Psychologist. Nickow et al. 2020, NBER. Kluger & DeNisi 1996; Wisniewski et al. 2020; Shute 2008 (feedback). Kulik et al. 1990; Slavin 1987; von Hippel 2024 (mastery, 2 sigma). Bisra et al. 2018; Chi & Wylie 2014. Sinha & Kapur 2021; Alfieri et al. 2011; Kirschner, Sweller & Clark 2006. Alfieri et al. 2013; Gentner et al. 2003 (comparison). Barnett & Ceci 2002; Sala & Gobet 2017 (transfer). van Merriënboer's 4C/ID. Macnamara et al. 2014. Rhodes & Tauber 2011. Marion & Thorley 2016.

Machine learning. Wu, Dyer & Neyshabur 2021 (arXiv:2012.03107); BabyLM findings 2023 and 2024 (arXiv:2412.05149). Lee et al. 2024, teaching arithmetic to small transformers (arXiv:2307.03381); McLeish et al. 2024 (arXiv:2405.17399). Eldan & Li 2023, TinyStories (arXiv:2305.07759). Li et al. 2025, small models and strong reasoners (arXiv:2502.12143). Hsieh et al. 2023 (arXiv:2305.02301). Ross et al. 2011, DAgger (arXiv:1011.0686); Agarwal et al. 2024, on-policy distillation (arXiv:2306.13649). DeepSeek-R1 2025 (arXiv:2501.12948). Jiang et al. 2021, PLR (arXiv:2010.03934); Yu et al. 2025, DAPO (arXiv:2503.14476). Kaushik et al. 2020 (arXiv:1909.12434). Ibrahim et al. 2024 (arXiv:2403.08763). Power et al. 2022, grokking (arXiv:2201.02177).

Communication. Lowe et al. 2020, supervision and self-play (arXiv:2002.01093). Lowe et al. 2019, pitfalls of measuring communication (arXiv:1903.05168). Lewis et al. 2017, Deal or No Deal (arXiv:1706.05125). Lu et al. 2020, seeded iterated learning (arXiv:2003.12694). Ross & Bagnell 2014, AggreVaTe (arXiv:1406.5979); Sun et al. 2017 (arXiv:1703.01030). Anthony et al. 2017, expert iteration (arXiv:1705.08439). Schmitt et al. 2018, kickstarting (arXiv:1803.03835). Foerster et al. 2018, COMA (arXiv:1705.08926). Ng et al. 1999, reward shaping. Rao & Daumé 2018; Grand et al. 2024 (asking by information value). Lin et al. 2024, DialOp (arXiv:2305.20076).

Prepared 1 October 2026 for the Frugal programme. Programme figures come from its own records: the findings scoreboard, the E1-R trainer-gate and pilot records, and the E2 v2 registrations. Figures marked as small-model evidence come from models of the programme's size or close to it; all others are inferences from larger models.