Measuring an i+1 language tutor from the outside: one robust effect, one language-dependent null, and one failed conformance bar

Krashen's comprehensible-input hypothesis holds that acquisition happens when input sits just beyond the learner's current competence - the so-called i+1 condition. Large language models make it possible, for the first time, to generate such input on demand rather than select it from a fixed corpus. We report three results from a production system that does so, each measured from outside the engine, using none of its internal beliefs about the learner. First, a robust positive: teaching a blank-slate forgetting learner from a comprehensibility-ordered stream rather than a shuffled one leaves it recalling 79.7 +/- 10.3 more features after a two-week gap and wasting 35.8 +/- 4.2 percentage points less exposure. Both our adaptive engine and a plain fixed CEFR ladder achieve this; sequencing for comprehensibility, not adaptivity, carries the effect. Second, a null that turns out to be language-dependent. Against a good fixed ladder, our adaptive engine gains 3.5 +/- 5.0 features in German - not significant - but 6.2 +/- 2.4 in Dutch, which is. We trace the difference to a single measured property, mean word length, and the chain of consequences it sets off, and state the falsifiable prediction that follows. Third, a negative result on the system's own terms. The engine's central guarantee is that each generated sentence introduces exactly one new element. Measured across six CEFR levels, conformance is 3.5% to 20.5% against an acceptance bar of 80%, and six independent expert evaluations of the served streams returned would_teach_this: false without exception. We give the diagnosis, including the part of the shortfall that is a defect and the part that is an artefact of how the metric was defined. We publish the negative result because the positive one is not interpretable without it: a system can demonstrate that comprehensible input works while failing to reliably produce it.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-12
DOI
https://doi.org/10.5281/zenodo.22729334
Primary Topic
Natural Language Processing Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Measuring an i+1 language tutor from the outside: one robust effect, one language-dependent null, and one failed conformance bar

Mateusz Wiącek
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
preprint

Measuring an i+1 language tutor from the outside: one robust effect, one language-dependent null, and one failed conformance bar

Mateusz Wiącek
preprint en

Abstract

Krashen's comprehensible-input hypothesis holds that acquisition happens when input sits just beyond the learner's current competence - the so-called i+1 condition. Large language models make it possible, for the first time, to generate such input on demand rather than select it from a fixed corpus. We report three results from a production system that does so, each measured from outside the engine, using none of its internal beliefs about the learner. First, a robust positive: teaching a blank-slate forgetting learner from a comprehensibility-ordered stream rather than a shuffled one leaves it recalling 79.7 +/- 10.3 more features after a two-week gap and wasting 35.8 +/- 4.2 percentage points less exposure. Both our adaptive engine and a plain fixed CEFR ladder achieve this; sequencing for comprehensibility, not adaptivity, carries the effect. Second, a null that turns out to be language-dependent. Against a good fixed ladder, our adaptive engine gains 3.5 +/- 5.0 features in German - not significant - but 6.2 +/- 2.4 in Dutch, which is. We trace the difference to a single measured property, mean word length, and the chain of consequences it sets off, and state the falsifiable prediction that follows. Third, a negative result on the system's own terms. The engine's central guarantee is that each generated sentence introduces exactly one new element. Measured across six CEFR levels, conformance is 3.5% to 20.5% against an acceptance bar of 80%, and six independent expert evaluations of the served streams returned would_teach_this: false without exception. We give the diagnosis, including the part of the shortfall that is a defect and the part that is an artefact of how the metric was defined. We publish the negative result because the positive one is not interpretable without it: a system can demonstrate that comprehensible input works while failing to reliably produce it.

Zenodo (CERN European Organization for Nuclear Research)
Baria Vungtau University (VN)
Quality Education
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Measuring an i+1 language tutor from the outside: one robust effect, one language-dependent null, and one failed conformance bar — Mateusz Wiącek · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS