Understanding the method

What Comprehensible Input Is, and Why Reading Works

Comprehensible input is language you understand. Not language you could work out with a dictionary and twenty minutes, and not language you have already mastered — language that reaches you as meaning rather than as a puzzle. The claim behind the term is that this, and not much else, is what builds a second language in an adult head.

It is a narrow claim with wide consequences, and it is worth being precise about what it does and does not say before deciding how to spend the next year of your reading.

Where the idea comes from

The phrase belongs to Stephen Krashen, who set it out in the early 1980s as the input hypothesis — one of five hypotheses that together made up a model of second language acquisition. The others matter less here, with one exception, but the shape of the argument does.

Krashen separated acquisition from learning. Acquisition is the process by which a first language arrives: unconscious, driven by exposure to meaning, producing the intuition that a sentence sounds right without any account of why. Learning is what happens in a grammar lesson: conscious, rule-shaped, available for inspection. His argument was that the two are not the same system, and that fluency comes from the first. Conscious rules can monitor and correct what you produce, but they do not become the thing that produces it.

If that is true, then the practical question is what drives acquisition. Krashen’s answer was input, and specifically input one step beyond the learner’s current competence — the formula he wrote as i+1, where i is where you are now. You move forward by understanding language that contains a little you do not yet have, using context, world knowledge and the rest of the sentence to cover the gap.

Krashen’s own summary of the argument leaves little room for anything else:

We acquire language in one way and only one way: when we understand messages. Stephen Krashen

The one other hypothesis worth keeping is the affective filter: anxiety, boredom and the fear of getting it wrong reduce how much of the input actually gets in. For a reader working alone this is less abstract than it sounds. A text that humiliates you every second line is not being processed for meaning, whatever your eyes are doing.

What i+1 does not mean

It does not mean one new word per sentence, or one new grammatical structure per page. Krashen never claimed i or +1 could be measured; the formula is a description of a relationship, not a specification. This is the most frequent criticism of the hypothesis, and it is a fair one: a theory that cannot be operationalised cannot easily be tested, and the input hypothesis has been attacked on exactly that ground for forty years.

The criticism damages the theory more than it damages the practice. You do not need a number to know the difference between a page that carried you along with three or four unfamiliar words in it and a page that stopped you six times in a paragraph. The first is plus one. The second is plus rather more than one, and what you get from it is not acquisition but translation practice.

The other standing objections are about what input leaves out. Merrill Swain argued from French immersion data that learners who received years of rich comprehensible input still produced grammar that had gone quietly wrong, and that being forced to produce language — to notice what you cannot yet say — does something input alone does not. Michael Long put the emphasis on interaction: on the negotiation that happens when a conversation breaks down and gets repaired. Both look right. Neither is an argument for reading less; they are arguments against believing that reading is the whole of a language.

Why reading, specifically

Input can be anything you understand: a conversation, a film, a podcast, a teacher speaking slowly. Reading has four practical advantages over all of them for an adult who is learning outside a classroom.

It is self-paced. Speech arrives at the speaker’s speed and is gone. A page waits. When a sentence does not resolve on the first pass you can read it again without asking anyone for anything, and the second pass is often where the structure clicks into place. For a learner below B2 this is not a convenience, it is the difference between input that is comprehensible and input that merely went past.

It is dense. An hour of conversation contains a surprisingly small number of words, most of them the same few hundred. An hour of reading at a modest 150 words a minute is nine thousand words of text, drawn from a wider register, with subordinate clauses that spoken language rarely bothers with. Written language is where the vocabulary above the two-thousand-word mark actually lives.

It is available. No partner, no schedule, no bandwidth, no embarrassment. Reading is the one form of input you can get for an hour a day for a year without arranging anything with anybody, and consistency over a year beats intensity over a fortnight in every account of how languages are learned.

Its difficulty can be controlled. This is the decisive one. You cannot easily order the world to speak to you at exactly your level, but a text is a fixed object that can be chosen, graded, footnoted and replaced. Everything in this series follows from that: if input has to be comprehensible to work, and written input is the kind whose comprehensibility you can actually engineer, then a graded text is the most reliable delivery mechanism there is.

What the evidence looks like

The strongest results in favour of reading do not come from testing the input hypothesis directly — that, again, is hard to do — but from programmes that put a lot of understandable text in front of learners and measured what happened. The best known are Warwick Elley’s “book flood” studies in Fiji in the early 1980s, where primary classes given large quantities of accessible English books outperformed classes on a conventional structural syllabus, in reading and in writing, with the gap widening over the following year. Comparable results have come out of extensive reading programmes in Japan, Taiwan and elsewhere since, with the usual caveats of educational research attached.

The vocabulary side has firmer numbers. Words are picked up from context in small increments, at a low rate per encounter — a few per cent per meeting is a common estimate — which sounds discouraging until you notice how many encounters a year of reading contains. It also explains why volume matters so much: incidental acquisition is a weak effect applied to an enormous number of repetitions.

The same literature gives the figure that governs everything about choosing a text. To read unassisted with good comprehension, you need to already know something like 98 per cent of the running words on the page — about one unknown word in fifty. At 95 per cent coverage comprehension is still workable with effort and support; below that it degrades quickly, and guessing from context starts to fail, because the context itself is made of words you do not know. Choosing texts at that density is a practical problem with a practical answer.

What this means in practice

If you take the input hypothesis seriously but not dogmatically, a few things follow for anyone learning to read a foreign language.

  • Understanding comes first. A text you understand at speed is worth more than a harder text you decode. The prestige of the source is irrelevant; the original Dostoevsky at B1 is not input, it is homework.
  • Volume beats intensity, but not alone. Reading a lot is what turns a weak per-encounter effect into a large one. Reading closely is what stops errors fossilising. The two are complementary, and the balance between them is the practical decision.
  • The unknown words are the point. A text with nothing new in it is comfortable and inert. What you want is the small residue of unfamiliar language that context can still carry — reading past it without stopping is a skill in itself.
  • Enjoyment is not a bonus. The affective filter argument and the plain arithmetic of habit point the same way: you will read a hundred pages of a story you like and eleven pages of a text you are enduring.

In short

Comprehensible input is language you understand while attending to what it means. Krashen’s claim is that this is what builds the system underneath fluency, and that the useful input sits just past your current level. The theory is hard to test and incomplete — output and interaction do work of their own — but the practical consequence is solid: read a great deal of text you can actually follow, at a level where the unknown is a thin margin rather than the whole page.

The rest of these articles are about doing that: how to read without knowing every word, how to judge whether a text is at your level, how much to read closely, and what a complete routine looks like. If you would rather see how the principle is applied to a book, the method behind these volumes sets it out.