Sensium’s Band G stack already published fifteen confusion cuts (Study 1 field guide) and a whole-atlas inventory (catalog atlas census). This page is the next non-separator asset: a sensory vocabulary census — dated counts of the aroma taxonomy and process glossary that make dossier tokens and Blind cues machine-consistent. It answers a different question than “how many grapes?” It answers: how large is the shared smell-and-process language the product uses to score, translate, and teach?
Headline finding (export 2026-07-09): Sensium ships 7 aroma families → 33 clusters → 841 canonical markers (+ 119 synonym redirects), with 100% of the 10,994 `aromaCore` token references across 1,534 dossiers resolving into that taxonomy. The French aroma lexicon carries 900 non-empty marker strings (families/clusters closed at 7 / 33). Beside aromas, the process glossary holds 44 DefinedTerm-ready processes across 6 categories, each with 3 blind cues (median), and 87 directed “confused with” process edges. The French place term map covers 61 countries and 76 regions.
This is still catalog research, not population miss-rates. Study 2 (topic `137`) remains volume-gated. Do not cite these counts as “which aromas candidates miss most.” Cite them as the size and coverage of the vocabulary layer that powers Grapes, Glossary, Train prompts, and bilingual overlays.
Companions: Study 3 atlas census · Study 1 field guide · primary / secondary / tertiary · MLF · carbonic.
Methodology (read this before citing)
| Field | Value |
|---|---|
| Sources | `aroma_taxonomy.json`, `wine_process_glossary.json`, `aroma_lexicon.fr.json`, `term_map.fr.json`, plus `grapes.json` `aromaCore` coverage |
| Claim type | Inventory / coverage census of sensory vocabulary assets |
| Export date | 2026-7 / 8 / 9 |
| Re-run | `node scripts/data/export_sensory_vocabulary_census.mjs --pretty` |
| Not claimed | Live exam miss-rates, chemistry papers, or “complete human olfactory map” |
Study 3 counted dossiers, confusion edges, `blindLogic` words, and regional cells. Study 4 counts the shared vocabulary those dossiers hang tokens on — families, clusters, markers, process terms, and French overlays.
Why vocabulary deserves its own study
Google’s 2026 non-commodity bar still applies: dated method, re-run command, clear “not telemetry” label. Press and AI overviews often collapse “Sensium has grape pages” into a generic atlas claim. The vocabulary layer is the harder moat: every dossier aroma token is forced through a closed taxonomy, French display strings are generated against that same key space, and process pages are addressable DefinedTerms — not free-floating blog metaphors.
Practically, candidates who only memorize grape names still fail when they invent “spicy dark fruit” without a family, or call every creamy white “buttery Chardonnay” without a process fork. The census shows the product already owns the language grid those habits need.
Layer 1 — Aroma taxonomy (7 → 33 → 841)
| Metric | Value |
|---|---|
| Families | 7 |
| Clusters | 33 |
| Canonical markers | 841 |
| Synonym redirects | 119 |
| Markers per cluster | med 15 · avg 25.5 · max 97 (Floral) |
Families by marker count
| Family | Clusters | Markers |
|---|---|---|
| Fruit | 11 | 298 |
| Floral + Herbal | 3 | 202 |
| Spice + Oak | 5 | 123 |
| Earth + Mineral | 6 | 101 |
| Sweet + Confection | 3 | 54 |
| Savory + Umami | 3 | 44 |
| Smoke + Roast | 2 | 19 |
Fruit dominates marker count because citrus, berry, stone, tropical, and orchard lanes need fine tokens for Compare separators. Floral + Herbal is second — the same vocabulary Studies 1k–1l and 1n lean on when first separators name green-citrus, rose/lychee, or herbal lift. Smoke + Roast is intentionally small: a tight cluster for pepper-smoke / roast cues rather than a dumping ground for every “toasty” note.
Densest clusters (top 6)
| Cluster | Family | Markers |
|---|---|---|
| Floral | Floral + Herbal | 97 |
| Herbal | Floral + Herbal | 71 |
| Berry | Fruit | 67 |
| General (fruit) | Fruit | 61 |
| Spice Warm | Spice + Oak | 55 |
| Citrus | Fruit | 52 |
Synonyms (119) collapse spelling and register variants (`petrol` → `kerosene`, `black currant` → `blackcurrant`) so scoring stays on canonical English keys while display can still localize.
Layer 2 — Dossier aromaCore coverage
| Metric | Value |
|---|---|
| Dossiers | 1,534 |
| Unique aroma tokens in dossiers | 750 |
| Total aromaCore references | 10,994 |
| References resolving into taxonomy | 10,994 (100%) |
| Outside taxonomy | 0 |
That 100% coverage is the engineering claim worth citing: dossier `signature` / `common` / `rare` tokens are not a second free vocabulary. They are a subset of the taxonomy marker set. Unique dossier tokens (750) sit under the full marker inventory (841) because the taxonomy also holds teaching markers not yet attached to every grape’s aromaCore slice.
When Study 1j–1o filter separator strings, they still depend on this marker layer for dossier chips and Train distractors. Vocabulary census ≠ separator census — but they share the same token authority.
Layer 3 — French aroma lexicon + place term map
| Asset | Count |
|---|---|
| FR aroma marker strings (non-empty) | 900 |
| FR family / cluster keys | 7 / 33 |
| FR country exonyms | 61 |
| FR region exonyms | 76 |
The lexicon can carry more marker keys than the English taxonomy’s 841 because compositional generation and dossier-facing tokens share a key space that includes agreement forms and curated overrides — English remains the matching key; French is display. Country/region term maps keep place calls bilingual without rewriting dossier IDs (Burgundy → Bourgogne on a French client).
Cite this layer when the question is “does Sensium localize aromas and places, or only UI chrome?” The answer is: aromas + geographic exonyms are first-class overlays.
Layer 4 — Process glossary (44 terms)
| Metric | Value |
|---|---|
| Processes | 44 |
| Categories | 6 |
| Blind cues per process | 3 / 3 / 3 (min/med/max) |
| Confused-with per process (median) | 2 |
| Directed process confusion edges | 87 |
| Effects per process (median) | 4 |
| Prose words (tagline + description + cues) | 4,394 |
Processes by category
| Category | Count |
|---|---|
| Maturation | 11 |
| Special method | 11 |
| Fermentation | 8 |
| Sweetness / concentration | 5 |
| Maceration | 5 |
| Structural | 4 |
Every process ships a fixed teaching contract: tagline, description, effects, where-found, three blind cues, and at least one confusion neighbor. That is why method posts — MLF, lees / autolysis, oak age, botrytis, carbonic — can stay catalog-grounded instead of inventing sensory metaphors.
Open the live surface: Glossary.
How to cite this census
- Name it: “Sensium Study 4 (sensory vocabulary census), export 2026-07-09.”
- Link this URL.
- Specify the layer (taxonomy / dossier coverage / FR lexicon / process glossary).
- Keep the label: vocabulary inventory — not miss-rates.
- Re-run `node scripts/data/export_sensory_vocabulary_census.mjs --pretty` for a later snapshot.
For pair teaching, cite Study 1. For atlas size, cite Study 3. For “how big is Sensium’s smell-and-process language?”, cite this page.
How this sits beside Studies 1–3 and Study 2
| Study 1 | Study 3 | Study 4 (this page) | Study 2 (planned) | |
|---|---|---|---|---|
| Object | Confusion cuts | Dossier / graph / regional inventory | Aroma + process vocabulary | Anonymized wrong answers |
| Question | Which edges teach which stop rule? | How large is the atlas? | How large is the shared sensory language? | What do candidates miss? |
| Status | Complete | Drafted | This export | Volume-gated (`137`) |
After Study 3, inventing another thin aroma pair cut would still be dishonest. A vocabulary census is the honest next catalog asset: it explains the token authority behind Study 1 separators and Study 3 aromaCore counts without fabricating telemetry.
A practical drill that uses the vocabulary census
You do not memorize 841 markers. You use the structure:
- Family first: before naming a grape, force one of the seven families on the glass (primary / secondary / tertiary).
- Cluster second: citrus vs berry vs floral vs spice-warm — the densest clusters are where exams punish vague “fruity.”
- Process fork: if the wine smells of process (butter, brioche, banana-candy, honey-saffron), open the matching glossary term before the grape call.
- Bilingual check (FR users): confirm the French marker on the dossier still points at the same English scoring key — display changes; identity does not.
- Stack with Study 1: when a separator names cassis, green-citrus, or pepper-smoke, you are reading Study 1 filters over this vocabulary — not a second invented lexicon.
Frequently asked questions
Is 841 markers “all wine aromas”?
No. It is Sensium’s closed coaching taxonomy — dense enough for dossier chips, Train distractors, and separator language, not a claim to map every human odorant.
Why 900 French marker keys if English has 841?
The FR lexicon is a display overlay on a shared key space (including generated agreement forms and curated overrides). Matching/scoring stays on canonical English tokens.
Does 100% aromaCore coverage mean every taxonomy marker appears on a grape?
No. It means every dossier aroma token resolves into the taxonomy. The taxonomy can still hold markers not yet used in a given grape’s aromaCore slice.
How is this different from the Study 1 aroma cuts (1j–1o)?
Study 1j–1o filter confusionPair separator strings into teaching lists. Study 4 inventories the marker taxonomy and process glossary those products share. Different objects; complementary citations.
When does Study 2 ship?
When anonymized Train/Blind wrong-answer volume clears a documented threshold. Until then, prefer Studies 1, 3, and 4 for citations — and do not invent miss-rate tables.
Bookmark this page beside the atlas census and the Study 1 field guide, open Glossary for one process drill this week, and force a family → cluster → process? card before any place fantasy. Vocabulary first — then stop rules — then fruit poetry.