Lexical diversity

Tracks how varied your vocabulary is across the manuscript. Sudden drops often mean you're leaning on a few favourite words too hard in that stretch.

lexical-diversity report card

What it measures

The core metric is standardized type-token ratio (stTTR) — computed on a rolling window of roughly 1,000 words at a time. Inside each window, the analyzer counts distinct words (types) and total words (tokens); the ratio is types divided by tokens. Standardizing the window size matters because raw TTR is length-dependent — longer texts naturally repeat more words, so comparing a chapter-long window to a paragraph-long one is meaningless without normalization.

Values range 0–1. In practice most published prose lands between 0.45 and 0.75. Below 0.45 the prose is measurably repetitive (either intentional — a mantra, a chant, a character's fixed obsession — or unintentional — the writer leaning on pet words). Above 0.75 the vocabulary is unusually varied, which reads as either literary richness or, in bad drafts, thesaurus-abuse.

The card computes the stTTR mean across all windows for the whole-manuscript verdict, and a lexical evolution array with one score per window (or per chapter, depending on document length) so you can see where the drops happen. The analyzer marks passages as "diverse" (mean > 0.7), "moderate" (0.5–0.7), or "repetitive" (< 0.5).

The methodology has one important limit: it treats every distinct word form as a distinct type. "Run", "ran", "running", "runs" all count as separate types, which inflates the number relative to lemma-based metrics. Compared against another writer's stTTR, that's fine (both are counted the same way). Compared against academic vocabulary-richness papers using lemmatised inputs, our numbers will read higher.

Why it's useful

Every writer has a small set of pet words they lean on unconsciously — five to fifteen words that show up two or three times more often than they should. You cannot see this while writing. The words feel neutral to you because you use them constantly; your brain filters them out the same way it filters out the noise of your own breathing. Only when someone else reads the manuscript back to you do you notice that you've written "suddenly" ninety-two times in three chapters.

This report surfaces the specific chapters and passages where lexical density collapsed — where you slipped into a repetitive stretch. It doesn't tell you which words are being over-used (that's Stylometry's job — Yule's K, Zipf-law fit, hapax ratio) — it tells you where in the manuscript the vocabulary went flat.

The natural use pattern is: run this report to find the troughs, then jump to those passages in the editor and use ⌘F to search for likely candidates ("just", "very", "suddenly", your protagonist's name repeated in dialogue tags, the sensory verbs you default to). If a trough coincides with a scene you know is deliberately repetitive — a ritual, a fugue, a character's OCD moment — leave it alone. If it coincides with what should have been dynamic action prose, revise.

How to read it

The card headline reads stTTR … — the mean standardized type-token ratio for the whole manuscript. Expand the card and you see two blocks stacked vertically.

The top block is the stTTR mean row — a horizontal Row bar with the numeric mean on the right and a plain-English verdict as a hint ("too short" if the manuscript is under the minimum window length, "diverse" for above 0.7, "moderate" for 0.5–0.7, "repetitive" below 0.5). The bar itself is colored: green above 0.7, amber between 0.5 and 0.7, red below 0.5.

Below that is the Evolution across book section — a cyan bar-chart sparkline with one bar per window as you move left-to-right through the manuscript. Shorter bars are the troughs; taller bars are the diverse stretches. The visual is small (about 40px tall) but enough to spot the shape at a glance — a fairly even flat line with modest wiggle means the vocabulary is steady across the book, and one or two very short bars in a row means a specific stretch collapsed.

There are no per-passage click-throughs in this card — the sparkline is descriptive, not interactive. To find the specific passage that produced a trough, count roughly how far through the book the low bar sits and open the corresponding chapter in the editor.

The card header carries a small colored dot next to the label: green for a fresh run, blue for a cached one. Inside the expanded card there's a "fresh" / "cached" pill above the metric and a re-run button next to it to force a fresh compute.

When to ignore it

Some scenes are supposed to be low-diversity. A repetitive ritual scene, a mantra or prayer, a character's obsessive internal loop, a nursery rhyme, a mocking chant — anything where verbal repetition is the point — will produce a trough that is not a bug. If you know what scene the trough corresponds to and repetition was the intent, leave it.

Genre also matters. Children's books and middle-grade fiction score low on lexical diversity because their working vocabulary is deliberately restricted — that's a feature, not a flaw. Poetry and prose-poetry score high because the form encourages novel word choice. Genre-thriller prose sits in the middle-low range because short punchy repetition is part of the pacing. Compare against your target register before treating the number as prescriptive.

Third: passages with a lot of proper nouns (a battle scene with fifteen named characters, a diplomatic dinner with formal titles) inflate the diversity artificially because every proper noun counts as a unique type. Those passages will read as high-diversity even when the actual working prose is repetitive. A single high spike surrounded by average bars often reflects the arrival of new named entities more than any change in prose quality.

Finally, the metric penalises deliberately simple prose. Hemingway would score below 0.5 on huge stretches of The Old Man and the Sea, and rightly so — the vocabulary restriction is the aesthetic. If you're chasing that register, don't panic when the bar goes red.

How to run it

  1. Click Write in the left sidebar and open the document you want to analyze.
  2. Open the right rail: click the Reports button (bar-chart icon) in the editor's top toolbar.
  3. In the rail header, click the Reports tab.
  4. Scroll to the Lexical diversity card under the "Linguistic" heading.
  5. Click the card to expand. First-time runs take 5–20 seconds. Cached runs render instantly (blue "cached" pill instead of the green "fresh" one).
  6. The collapsed row headlines with stTTR … — the mean standardized type-token ratio. Expand to see the rolling-window curve and click into any amber trough to jump to the passage.
  • Stylometry — the whole-manuscript view of vocabulary
  • Reports — back to the panel overview