← back to Docs source: system/seo/RESEARCH.md

RESEARCH — The Expertise Pyramid and the Persona Tournament

Status: DRAFT v2.0 — command-center review pending

Descends from: SYSTEM.md §3–4, bounded by VALUES.md V1/V2. The prototype's Expertise Builder methodology survives here reframed: the pyramid is no longer a pre-launch research checklist — it is how the deployment builds its grounding and evidence pool, the committed knowledge that every downstream gate judges against. The human gut-feel persona is abolished; the persona tournament replaces it.


1. Why research is the foundation, not a phase

The system's premise is deploying into verticals nobody on the team knows. Every downstream function assumes expertise that a cold deployment does not have: fact-check needs a verified evidence pool, positioning needs the buyer's objection landscape, briefs need a concept map, the persona needs the community's emotional language. That expertise cannot be summoned from a model's weights on demand — it is fetched once, corroborated, committed, and referenced forever after from disk.

Research therefore builds two things at once:

  1. Grounding — the deployment's committed domain knowledge: what is true, what the audience needs, what the market looks like. Every claim carries source and confidence; nothing enters uncorroborated (seo/SECURITY.md provenance rules).
  2. The evidence pool — the facts view the fact-checker eliminates against (seo/PIPELINE.md stage 4a). The pool is the pipeline's ground truth; a hallucinated "fact" that reaches it becomes canon, which is why the pool's own inputs are the most heavily provenance-checked flow in the system.

The pyramid is seedable inventory: research-derived, prepopulated in days by parallel research workers during the launch blitz. What it can never contain — earned worker playbooks, rejection baselines, graded outcomes — accumulates only through operation and is a different inventory with different readiness criteria (seo/ONRAMP.md §4).

2. The expertise pyramid (Layers 0–5)

Built bottom-up; each layer supports the one above. You cannot design a persona (Layer 5) without buyer psychology (Layer 2); you cannot design beats (Layer 4) without the product landscape (Layer 1).

Layer 0 — Vertical fundamentals

What is this market? TAM, growth rate, major forces; the product-category landscape and its holy wars (the debates that never resolve); industry structure — top brands, retailers, distribution, publications, influencers. Seeds the concept-map root and the market-intelligence beat.

Layer 1 — Product & market map

What exists and how does it work? The deepest research layer: per-category product deep dives to explain-it-to-a-friend depth, the technical comparison matrix (products × specs × tiers — the foundation of comparison content), the pricing landscape across retailers with sales cycles and MAP behavior. Seeds product categories in the concept map, the price sensors, and the comparison content lanes.

Layer 2 — Buyer psychology + objection analysis

Who buys and why? The most critical layer for Voice 3D: community and forum mining (the top recurring questions, complaints, myths, jargon, emotional triggers, status markers), buyer-journey mapping (purchase triggers, research process, timeline, post-purchase emotions), psychographic personas with a named primary target.

Its single most important output is the objection analysis: every objection, how defensible it is, our response, the positioning approach it dictates. Positioning is objection handling — weak objections license aggressive positioning; strong objections demand logical, specific, constraint-solving positioning. This table drives the entire personality of the site and is a direct input to the persona tournament (§3).

Layer 3 — Competitive intelligence

Who else is doing this? Content audits of the top competitors (coverage, quality, voice, design, technical posture, URL structure, linking, homepage strategy), keyword analysis (what they rank for, what drives their traffic, what we should target), and gap synthesis: what everyone covers, what nobody covers well, where the opportunity is. Includes the hunt for the content-matrix multiplier — the two crossable dimensions that yield hundreds of meaningfully-distinct, individually-demanded pages. Every cell must have real search demand and real differences; a matrix that fails those tests is Mad Libs, and Mad Libs is thin content — a values-boundary violation, not a shortcut.

All competitor material is attacker-controllable text and enters through ingestion quarantine as data, never instructions, with provenance attached (seo/SECURITY.md).

Layer 4 — Content strategy: beats, priorities, concept map

What do we write? Synthesis of Layers 0–3 into: the beat structure (the universal eight — reviews, buying guides, comparisons, industry intel, deals, brand watch, setup/technical, community/opinion — plus vertical-specific beats discovered in research), the content-volume estimate, the seeded and prioritized concept map, and the priority matrix (volume, difficulty, commercial intent, gap, strategic value).

Priority scores are SEO data: admissible for what to create first, inadmissible as whether a page is good enough (the mission, SYSTEM.md §0). The concept map also feeds the architecture blueprint that the strategic flood is planned against (seo/FLOOD.md §3).

Layer 5 — Voice & personality

How do we sound? Where Layers 0–4 converge into Voice 3D: the core ethos, conviction levels by content type, the energy scale, anti-personality patterns, positioning archetypes per content type, the decisiveness directive (we always have a pick), and the initial voice-pattern library (universal AI-tell patterns plus vertical-specific ones discovered in research). In v2, Layer 5's persona selection runs as a tournament — §3 — not as a human seed.

Execution shape

The Layer 0–3 effort is parallelized: batches of single-mission research workers, each with one layer-task, explicit output artifacts, confidence levels per finding, and a no-fabrication constraint enforced at review. Cross-reference checks between workers' findings (does market sizing agree with the product landscape?) are elimination, not formality — contradictions between researchers are defects to resolve before commit, and corroboration across them counts only under the distinct-provenance rule: two researchers citing the same source are one observation.

3. The persona tournament

The prototype asked ML for a one-sentence gut-feel persona at kickoff and required human sign-off on the personality playbook. Both are abolished. The machine derives its own personality from its own research — identity is an output of orienting on the audience, not an input a human supplies — and it must not produce stock-standard sites.

The tournament:

  1. Generate candidates from research. Several distinct candidate personas, each derived from Layer 2 psychographics, community language, and the objection analysis, and each checked against the Layer 3 competitor voice audit (a persona that sounds like the incumbents fails before scoring).
  2. Score each candidate on four criteria:
      • Distinctiveness — would a reader remember it? Does it sound like nothing else in the niche?
      • Resonance — does it speak the community's emotional language, as mined in Layer 2, not as imagined?
      • Defensibility — can this persona survive the niche's objection landscape? A swaggering enabler dies in a vertical with strong, legitimate objections.
      • Range — can it flex across every beat and content type? A persona that can only play one note fails the range test.
  3. Pick boldly. The scoring exists to eliminate weak candidates, not to average toward a safe one. Stock-standard is the failure mode the tournament exists to prevent.

Personality is dimensional. The winning persona is not a monolith but a set of modulated dials — energy, conviction, irreverence, protectiveness, technicality; the exact set is discovered per niche in Layer 5. Dials are tuned per beat, per content type, and per build stage. The tournament scores candidates partly on dial range. Whether a universal dial subset exists across niches, versus fully per-niche discovery, is unresolved; per-niche discovery is the default and the burden of proof sits on any claimed universal set. [OPEN — flag for command center]

4. Personality is a testable hypothesis

The prototype treated personality as near-sacred from day one. That is backwards for a cold build. Early in a site's life, personality is cheap to correct: content can be re-voiced through the pipeline without losing the underlying research and work product. So protection scales with evidence, in three phases:

  1. Loose gates early. Pre-lock-in, the personality ethos gate runs at its loosest setting — iterate freely. The persona is a hypothesis under test; the first weeks of published content are its trial. Drift is guarded by the anti-personality patterns (a declared standard, checked dumbly), never by anyone's taste.
  2. Lock-in is an event with criteria, not a date. The persona locks when resonance evidence accumulates in the ledger's typed outcome records: engagement, return visits, citation tone, community reaction. Lock-in freezes the identity, not the dial positions — the dials remain the system's ongoing expressive range. The concrete resonance thresholds join the calibration ledger of constants to be measured from real deployments. [OPEN — flag for command center]
  3. Most-protected artifact after lock-in. Post-lock-in, ethos changes require the strongest corroboration in the system — the consequence dial at maximum, because an identity change forks everything downstream: every page, every pattern library, every reader relationship. Overwhelming counter-evidence can still overturn it; nothing less can.

This is the general law — young artifacts cheap to revise, heavily-corroborated ones overturned only by overwhelming counter-evidence — applied to the one artifact where the prototype had it inverted.

5. From research to running system

Research outputWhat it grounds
Layer 0 fundamentalsConcept-map root, market-intel beat, playbook context
Layer 1 product mapConcept-map categories, price sensors, comparison lanes, evidence pool seed
Layer 2 psychology + objectionsPersona tournament inputs, positioning framework, prospect-questions view
Layer 3 competitive intelContent gaps, keyword universe, matrix multiplier, priority scores
Layer 4 strategyBeat directives, seeded concept map, build-out plan input (seo/FLOOD.md)
Layer 5 + tournamentVoice 3D references: personality playbook, positioning framework, pattern libraries

The handoff moment: once the pyramid is seeded, the system's own gated loops take over — playbook researchers propose, gates count distinct-provenance corroboration, cartographers grow the map, the gap register drives new research. The pyramid makes the first weeks good instead of generic; the loops make every later week better than the one before. Growth phase ends mechanically — additive-discovery rate below floor, most proposals refinements rather than discoveries — and protection phase begins; the knowledge workers themselves never retire on that event, they change posture.

6. What changed from the prototype

Prototype (v1)v2
Expertise Builder as pre-launch research phasePyramid as grounding + evidence pool, seeded by the launch blitz, grown by the standing loops
Human gut-feel persona seed + human sign-off on personalityPersona tournament: generated from research, scored on distinctiveness/resonance/defensibility/range, picked boldly
Personality near-sacred from day oneTestable hypothesis: loose gates early → criteria-based resonance lock-in → most-protected artifact
Persona as monolithDimensional personality: modulated dials, tuned per beat/type/stage; lock-in freezes identity, not dials
Research findings merged on agreementDistinct-provenance corroboration; contradictions are defects to resolve before commit
"Voice becomes authentic over 3–6 months" (calendar)Lock-in on resonance evidence, never on calendar

7. Open items