Skip to content
LingunAI
Guide

How many words do you need to know to read a language?

To read unsimplified text comfortably, you need to know about 98% of its words. For English, research puts that at roughly 8,000–9,000 word families, but graded and simplified text lets you start reading with far fewer.

The short answer: about 98% of the words on the page

You need to know about 98% of the running words in a text to read it comfortably without a dictionary. For unsimplified English, Paul Nation (2006) estimated that this takes roughly 8,000–9,000 word families for written text such as novels and newspapers, and 6,000–7,000 for spoken text. You can read with effort at lower coverage, and you can read graded text far earlier.

Goal Coverage Approximate vocabulary Source
Unassisted reading of novels and newspapers 98% 8,000–9,000 word families Nation (2006)
Understanding spoken English 98% 6,000–7,000 word families Nation (2006)
Reasonable comprehension, with effort 95% 4,000–5,000 word families Laufer and Ravenhorst-Kalovski (2010)
A graded reader at your level About 98% A few hundred to a few thousand headwords The publisher’s word list

These figures come from research on English, which is where most of this work has been done. Treat them as a guide for other languages rather than a precise target; the section on Chinese, Turkish and Hungarian below explains why counts differ.

What vocabulary coverage means

Vocabulary coverage is the percentage of the running words in a text that you already know. Running words count every word on the page, repeats included: in “the dog saw the cat”, the word “the” counts twice. Coverage is a simple measure, and it’s closely tied to whether you’ll understand what you read.

Two numbers dominate the research. Laufer (1989) suggested that about 95% coverage was needed for reasonable comprehension. Hu and Nation (2000) had learners read a fiction text at different coverage levels and found that understanding rose as coverage rose, with about 98% needed for most readers to follow the text without help. Later work, including Schmitt, Jiang and Grabe (2011), found a fairly steady rise rather than a sudden threshold, so every extra point of coverage helps.

The gap between 95% and 98% looks small. On the page, it isn’t:

Coverage Unknown words On a 200-word page What reading feels like
100% None 0 Effortless, good for building speed
98% 1 in 50 About 4 Comfortable; many new words can be guessed from context
95% 1 in 20 About 10 Workable, with effort and some lookups
90% 1 in 10 About 20 Slow decoding; the thread keeps breaking
80% 1 in 5 About 40 Not really reading

Word families, lemmas and word forms

Vocabulary-size estimates count word families, not every distinct word you see, and the unit you count changes the answer by thousands. Three units are common:

  • A word form is any spelling that appears in text. Walk, walks, walked and walking are four word forms.
  • A lemma is a base form plus its inflections, within one part of speech. Those four forms are one lemma, the verb walk.
  • A word family is a base word plus its inflections and its common derived forms. The family of use includes uses, used, using, useful, useless and user.

Nation’s figures are in word families, which is why they sound smaller than you might expect: 8,000 families stand for a much larger number of individual forms. Counting by family also assumes that once you know a base word, you’ll recognize its relatives. That holds for transparent forms like useless, and much less for pairs like hard and hardly, where the derived word means something quite different.

Single-word counts also leave out multiword expressions. Phrases like “look after”, “by and large” and “let it slide” mean more than their parts, and you can know every word in them and still miss the point. Learning them as units, not word by word, is part of building a reading vocabulary.

Why the first few thousand words matter most

The most frequent few thousand words do most of the work in any text, so each one you learn early adds far more coverage than a rare word learned later. Word frequency is extremely uneven: a small set of words such as the, of, and, to and it appears constantly, and each step down the frequency list adds less coverage than the step before.

That’s why moving from 95% to 98% coverage takes thousands of extra word families, even though it’s only three percentage points. Beyond the high-frequency core, each new word turns up rarely: you might meet it twice in one book and not again for weeks.

Compare two sentences:

  • “She said she would call him after work, but the train was late.” Every word here belongs to the high-frequency core of English.
  • “The central bank raised rates for a third consecutive month, citing persistent inflation.” Most of these words are common too, but consecutive, citing, persistent and inflation are not, and they carry much of the meaning.

The lesson is to learn the high-frequency core deliberately and early, through a course, flashcards and easy reading. After that, which words are most useful depends more and more on what you read about, and the best way to meet them often enough is to read a lot, preferably several texts on the same topic.

How graded and simplified text lowers the bar

Graded and simplified texts lower the vocabulary you need by changing the text instead of waiting for the reader. A graded reader written within a 1,000-headword list gives someone who knows those words close to full coverage, so a story that would be unreadable in its original form becomes a comfortable read.

That’s what makes reading possible from the early levels. Instead of waiting until you know 8,000 word families, you read at about 98% coverage at every stage, using texts written for the vocabulary you have now. Because each book recycles its core words many times, it also provides the repeated meetings new words need before they stick. The guide to graded readers covers how to choose a level.

AI rewriting does something similar for texts you pick yourself: rewriting a news article at a lower CEFR level swaps rare words for common ones and shortens sentences. The control is looser than a fixed word list, so a few harder words slip through, but the effect on coverage is the same in kind.

Counting words in Chinese, Turkish and Hungarian

The idea of coverage applies to any language, but word counts don’t transfer directly, because languages package meaning into words differently.

Chinese has two counts: characters and words. Text is written without spaces, so word boundaries aren’t marked. A few thousand characters make up almost all everyday written Chinese, but characters aren’t words. Most words in a dictionary combine two characters, while many of the most frequent words in running text, such as 的, 是 and 我, are single characters. Knowing the characters helps without guaranteeing the word: 电 (electricity) and 脑 (brain) make 电脑, computer, which you might guess, but 东 (east) and 西 (west) make 东西, which usually means “thing”. For Chinese, the useful target is words you can read in context, characters included.

Turkish and Hungarian are agglutinative: they build words by adding suffixes to a root, often one meaning per suffix, so a single root can appear in dozens or even hundreds of forms. In Turkish, ev (house) becomes evler (houses), evlerimiz (our houses), evlerimizde (in our houses) and evlerimizden (from our houses). Hungarian does the same with ház: házak, házaink, házainkban and házainkból mean the same four things.

Count every distinct form as a word and a Turkish or Hungarian vocabulary looks enormous, while any word list seems to cover very little. Counting roots plus the suffix system is fairer. In practice, learning the suffixes is part of learning to read these languages; once you know them, a familiar root is usually easy to spot inside a long word.

A realistic plan to grow your reading vocabulary

The most reliable way to grow a reading vocabulary is to combine plenty of reading at about 98% coverage with spaced review of a few words from each session. Reading brings words back in context; review makes sure the ones you choose actually stay.

  1. Build the core first. Learn high-frequency words through a course, flashcards and very easy reading until short, simple texts feel readable.
  2. Read at 98%. Use graded readers or texts rewritten at your level, so only a few words per page are new. Twenty minutes a day at a modest 100 words per minute is 2,000 running words a day, about 60,000 a month.
  3. Keep a few words, not all of them. From each session, save five to ten words that recur, that you’d want to use, or that blocked your understanding. Five a day comes to more than 1,800 a year.
  4. Review on a schedule. Spaced repetition brings each word back just before you’re likely to forget it. It builds on the forgetting curve Hermann Ebbinghaus described in 1885 and on the spacing effect, confirmed in a large meta-analysis by Cepeda and colleagues (2006). The spaced repetition guide explains how to set it up.
  5. Read narrowly, then widely. Several texts on one topic recycle that topic’s vocabulary, which is how mid-frequency words go from recognized to known. Then change topics.
  6. Step up when it gets easy. When a level feels effortless, move up one. When you reach unsimplified text, start with subjects you already know well, since background knowledge makes unknown words easier to guess.

Building reading vocabulary with LingunAI

LingunAI is built around the same loop: read at your level, keep the words that stop you, review them until they stick. Paste an article into the article reader and it’s rewritten in the language you’re learning at any CEFR level from A1 to C2, keeping the facts. Or generate a short story in StorySplice built around words from your flashcards, so the words you’re reviewing turn up again in a new context.

Select any word that stops you and save it as a flashcard. Cards come back on a spaced schedule, by default after 1, 3, 7 and 15 days, and you can design your own intervals. Flashcards you make by hand, reviews and quizzes are free; stories and rewrites use AI credits, and the free plan includes a one-time allowance to try them (see pricing). As with any AI text, the level control is approximate and no editor checks it, so treat it as reading practice and confirm anything that matters.

Questions, answered

01 How many words do you need to know to read a book in another language?

To read an unsimplified novel comfortably without a dictionary, you need to know about 98 percent of its running words. For English, Paul Nation (2006) estimated that this takes roughly 8,000 to 9,000 word families. You can follow a text with effort at around 95 percent coverage, and graded readers let you read whole books with a few hundred to a few thousand words.

02 What does 98% vocabulary coverage mean?

Coverage is the percentage of running words in a text that you already know. At 98 percent coverage, one word in every 50 is unknown, which is about four on a page of 200 words. Hu and Nation (2000) found that most learners need roughly this level to understand fiction without help.

03 What is a word family?

A word family is a base word together with its inflected forms and its common derived forms. The family of use, for example, includes uses, used, using, useful, useless and user. Vocabulary-size research such as Nation (2006) counts word families, so 8,000 word families stand for many more distinct word forms.

04 How many words do you need to understand spoken language?

Paul Nation (2006) estimated about 6,000 to 7,000 word families for 98 percent coverage of spoken English, fewer than for written text because everyday speech draws on a narrower range of words. Listening can still feel harder than reading, because you usually can't slow the speaker down or reread a sentence.

05 How many characters do you need to read Chinese?

Chinese has two counts, characters and words, and they are not interchangeable. A few thousand characters make up almost all everyday written Chinese, but most dictionary words combine two characters, and knowing both does not always give you the meaning: 东西, built from east and west, usually means thing. The useful target is words you can read in context, characters included.

Put it into practice today.

Free plan, no card. Read a story written at your level and keep every word that stops you.

Start reading free