MODELS — The Intelligence Spectrum, Operationalized
Status: DRAFT v2.0 — command-center review pending
Descends from:
SYSTEM.md§7 (model policy, ML-locked) · Cross-references:specs/CONSEQUENCE.md(the dial),corpus/COMMISSIONING.md(manifests, trials),corpus/FUNNEL.md(escalation),seo/ONRAMP.md(ramp settings).The policy itself is locked by ML and restated in
SYSTEM.md§7; this document operationalizes it. Where this document and §7 conflict, §7 wins.
1. The spectrum
Intelligence in this system is a spectrum with four rungs. Three are available to workers; the fourth is not a model assignment at all — it is the escalation ladder.
| Rung | What it is | What runs there |
|---|---|---|
| Tier 0 — Scripts | Deterministic code. Counting, set reconciliation, timestamp arithmetic, known-answer probes, schema checks, threshold comparison. No prompt, no context window, no persuadability. | Gates. Meters. Sentinels. Health monitors. The values-gate mechanical checks. Everything in the self-knowledge layer that can possibly be arithmetic. |
| Tier 1 — DeepSeek V4 Flash class | Cheap, fast LLM. Enough intelligence to read a rubric and apply it; not enough to be argued into a sophisticated mistake. | Orienters that need language comprehension a script can't provide. Condensers. Extractors. Formatters. Mature generators whose directives have absorbed the hard thinking. |
| Tier 2 — GLM 5.2 class (the ceiling) | The most capable model any worker may run. Hard ceiling below the strategic layer. | Generators on high-consequence or immature work: framing, diagnosis, proposal-writing, persona work, build-phase production. |
| Above the ceiling — the command center (Claude Fable) | Not a tier. Not rentable. Consequence that exceeds GLM's envelope does not get a bigger model — it climbs the escalation ladder to the command center, which decides within delegated bounds or carries it to ML. | Structural decisions, doctrine, deadlock arbitration, the four contact categories. |
Two consequences of the ladder shape everything below:
- There is no "rent a frontier model for this one job." A job too consequential for GLM is by definition a strategic decision, and strategic decisions belong to the command center — which has organizational context no rented model would have. The fix for under-powered workers is never a bigger model; it is a smaller, better-specified job, or an escalation.
- Dumbness is load-bearing. The orienter tiers (0 and 1) are not a cost compromise — they are the safety property (
corpus/LAWS.md: a clever judge can be argued into a sophisticated mistake; a counter cannot). An orienter is never promoted to Tier 2. If judging a thing seems to require GLM-class intelligence, the standard is under-specified; fix the standard, not the judge.
2. Assignment rubric by consequence
Assignment is priced by consequence-per-token, never context volume. A worker that reads 200K tokens of ledger and emits one count is a script's job no matter how big the input; a worker that writes 500 tokens which fork a quarter's strategy is a GLM job (or an escalation) no matter how small the output.
Consequence tiers per specs/CONSEQUENCE.md (the one dial: corroboration required scales with downstream forking). The rubric:
| Worker role | Consequence: low | Consequence: medium | Consequence: high | Touches the values boundary |
|---|---|---|---|---|
| Orienter / gate | Script | Script, DeepSeek-class if language is unavoidable | Script wherever possible; DeepSeek-class ceiling. The threshold moves, never the judge's brain. | Script only (the values gate is deliberately dumb and un-overridable) |
| Generator | DeepSeek-class | DeepSeek-class, GLM on evidence of under-power | GLM-class | GLM-class + highest-tier gating; verdict escalates if ambiguous (V6) |
| Condenser (funnel) | Script/DeepSeek | DeepSeek | DeepSeek — condensers compress, never re-judge; they do not need to be smart, they need to conserve valence | n/a — conservation is audited by script tracers |
| Monitor / sentinel / self-knowledge | Script | Script | Script — the four dumb primitives only | Script |
| Improvement layer: diagnose + propose | DeepSeek | GLM | GLM | Escalate |
| Improvement layer: verify | Script (counting a pre-registered trial is arithmetic) | Script | Script | Script |
Reading the rubric:
- Orienters and gates run scripts or DeepSeek-class regardless of consequence. Rising consequence raises the corroboration threshold the gate demands (
specs/CONSEQUENCE.md), never the gate's intelligence. This is the whole point of the one dial: it lets the judge stay dumb everywhere. - Generators scale with consequence, up to the GLM ceiling and never past it.
- Verification is always cheap. Only diagnosis and proposal are smart; counting the outcome of a pre-registered trial is arithmetic (
SYSTEM.md§4). - Exact tier boundaries (what counts as low/medium/high) are owned by
specs/CONSEQUENCE.md; this document does not duplicate them.
3. Assignment is a commit
Per-worker model assignment lives in the worker's manifest (manifests/<worker>.yaml, see ops/DEPLOYMENT.md §3) and is a commit in the full sense of corpus/LAWS.md — it crosses the commit boundary, so it enters on evidence and changes only on evidence:
- Entry. A new worker's initial assignment comes from the rubric in §2 plus the build-phase posture in §4, recorded in the manifest with its rationale. The commissioning gate checks the assignment against the rubric before the worker exists (
corpus/COMMISSIONING.md). - Trial. Any assignment change — up or down — is a champion/challenger trial on the Trial-Verifier path: pre-registered prediction, both assignments run on comparable work, outcomes read from the ledger, a dumb gate counts the result. No exceptions for "obviously" cheaper or "obviously" needed.
- Demotion. When the cheaper class matches recorded outcomes (rejection rate, gate-pass rate, outcome records — not vibes), the worker demotes. Matched is enough; the cheaper class does not need to win. Demotion is an auditable, reversible ramp event.
- Promotion. Requires metered evidence of under-power: a quality series showing this worker's class is the binding constraint (e.g., rejection rate elevated vs class baseline, with diagnosis reproducing the failure at Tier 1 and not at Tier 2). "The output would probably be better" is not evidence. Promotion never crosses the GLM ceiling — see §1.
- Sentinel. Every demotion leaves a drift alarm on the worker's quality series against its pre-demotion baseline. Regression re-opens the trial automatically.
Nobody hand-tunes model assignments — not the command center, not ML. The manifest records; the ledger decides.
[OPEN — flag for command center] Numeric trial parameters: minimum qualifying runs per champion/challenger arm, the "matches outcomes" equivalence band, and the regression threshold that re-opens a demotion. Placeholders pending calibration from the first v2 deployment's ledger data.
4. Build-phase posture: start capable, demote on evidence
Cheap-model-plus-immature-directive is a compounding error source at exactly the moment errors shape the foundations. Therefore new deployments invert the steady-state cost picture — within the worker ceiling:
- Generators launch at GLM-class while their directives are immature, their workspace playbooks empty, and their rejection baselines nonexistent.
- Orienters launch dumb and stay dumb. The ramp never touches gate intelligence — build-phase loosening is a threshold setting (
seo/ONRAMP.md), not a model setting. - Per-worker demotion, no global event. Each generator demotes individually when its maturity criteria hit: directive stable N runs, workspace playbook past a depth floor, rejection rate floored under threshold, reviewer approvals clean. Criteria are evidence-based, never calendar-based, and the demotion runs through the §3 trial machinery like any other assignment change.
- The principle: intelligence migrates from the model into the artifacts. Model strength compensates for directive immaturity; as directives stabilize, playbooks deepen, and institutional memory accumulates on disk, the same job needs less model. A mature deployment running mostly DeepSeek-class is not degraded — it is the evidence that the knowledge plane works.
Note the difference from the v1 doctrine ("strong models everywhere at launch"): in v2 the ceiling is absolute. The build phase floods GLM-class compute, never command-center-class compute. The command center's build-phase role is structural (commissioning, doctrine, the ramp itself) — it does not write content, ever.
[OPEN — flag for command center] Numeric maturity criteria for demotion (N stable runs, playbook depth floor, rejection-rate threshold) — placeholders pending the same calibration as §3's trial parameters.
5. Cost posture: spartan design, not token thrift
Restated from corpus/LAWS.md because model policy is where it is most often misread:
- The system is built to crush tokens. Flood compute during the build; run loops at whatever cadence the evidence-epoch rule permits; never make a worker smaller-context or lower-cadence to save money.
- The constraint is clean design: zero duplicated jobs, zero dead wiring, zero redundant workers, zero orphan functions. A duplicated worker is a design defect at any price; a busy fleet of cheap workers is the intended shape.
- Cost control falls out of the architecture, not out of restraint: orienters and the entire meta-plane are scripts or DeepSeek-class arithmetic over the ledger (four of the six closure functions —
corpus/LOOP-CLOSURE.md), so headcount can run 2:1 to 4:1 meta:operating while cost stays operating-dominant. - The only model-cost lever anyone is permitted to pull is the §3 demotion trial. If spend looks wrong, the diagnosis is a design audit (duplication, dead wiring, cadence above evidence rate), never a unilateral downgrade.
6. Escalation economics
The ladder above the GLM ceiling is an attention budget, spent on exactly one scarce resource: the command center's (and ultimately ML's) judgment.
What climbs to the command center:
- Structural decisions — commissioning/retiring worker classes, wiring changes, pod composition, org-chart restructuring.
- Doctrine — changes to anything in
corpus/,seo/,specs/; amendment proposals forVALUES.md(decided only by ML). - Deadlocks that survived the ladder — a gate stuck past its deadlock clock after the remediation ladder (
ops/DEPLOYMENT.md§6) and arbiter mechanics have run out. The command center resolves within delegated bounds; only deadlocks it cannot resolve reach ML as Tiebreaker. - The four contact categories (
corpus/FUNNEL.md) — items that must reach ML, carried by the command center via the Executive Brief's NEEDS A DECISION section. - Values ambiguity — V6: genuinely ambiguous boundary questions stop the action and climb.
What NEVER climbs:
- Content. No draft, brief, rewrite, or publish decision touches the command center. The pipeline's gates are the authority; the command center is not a better gatekeeper, it is the wrong one.
- Domain judgments. Keyword priority, persona dials, playbook content, link targets — the operating layer and its knowledge plane own these. The command center is the organization's expert on exactly one thing — the organization — and never becomes the domain expert (
SYSTEM.md§4). - Routine operations. Retries, reroutes, staleness, credit alerts, schedule drift — the health monitor and remediation ladder own these; the command center hears about them only through the Brief, as valence-conserved summary.
The dumb enforcement: escalation is a typed record on the ledger, and a script audits the type against this list. An escalation of a never-climbs type is itself a defect, filed against whoever raised it — the ladder is protected from spam the same way every other channel is: by a counter, not a judge.
Ownership: this document is doctrine tier 5 (ops). It operationalizes but never overrides SYSTEM.md §7. Model class names (GLM 5.2, DeepSeek V4 Flash, Claude Fable) are current bindings of the tiers, replaceable by ML without changing the tier structure.