CONSEQUENCE — The Stakes Rubric and the One Dial
Status: DRAFT v2.0 — command-center reviewed 2026-07-20 · ML review pending What this is: the mechanical contract for evidence scales with consequence (LAWS.md). v1 hand-set fifteen per-gate thresholds and got them wrong in both directions — dead checks, overkill retire thresholds, threshold-1 build cliffs. v2 replaces the per-gate threshold table with one dial: grade the decision's stakes with five yes/no questions, read every control off the tier. The gate stays dumb everywhere; only its threshold moves.
1. The five questions
For any decision a gate must judge, answer five binary questions. No weights, no partial credit, no clever scoring — a yes-count. A counter cannot be argued with.
| # | Question | Yes when the decision… |
|---|---|---|
| Q1 | Does it touch a commit? | …creates, changes, or removes permanent state: a page, a worker, a directive, a standard, a schedule, a config. |
| Q2 | Does it touch the world? | …crosses the sandbox boundary: publishes, sends, posts, requests, purchases — anything an outsider can observe. |
| Q3 | Does it touch ML's attention? | …pushes to ML: escalations, Discord pings, brief content, anything on the four-category contact surface. |
| Q4 | Does it touch privileged access? | …uses or alters credentials, keys, payment instruments, publishing rights, or the grant ledger. |
| Q5 | Does it touch the values floor? | …is within reach of a VALUES.md floor clause: proof-grade legality (VF1), scam-shaped presentation or transaction (VF2/VF3), operator-integrity invariants (reporting valence, granted perimeter). Quality/thin-content judgments are NOT Q5 — they are mission machinery graded by Q1/Q2. |
Tier = the yes-count (0–5), with one override: a yes on Q5 sets tier 5 regardless of the other answers. Values-boundary touches are never averaged down by an otherwise-low-stakes decision.
One asymmetry modifier: a decommit (retire a worker, prune a page, revoke a capability, unpublish) grades one tier above the equivalent commit, capped at tier 4 (Q5 still owns tier 5). Removing capacity demands stronger evidence than adding it (SYSTEM.md §6); this is where that law becomes arithmetic.
Grading is mechanical: each gate's ramp-config entry carries its five pre-graded answers (stakes: [Q1..Q5]) reviewed at commissioning. Workers do not grade their own proposals; the gate reads the tier from config. A decision class whose answers change (e.g., a knowledge artifact starts feeding published pages) is re-graded through the worker-update path.
2. The tier table
The dial. Every control that used to be a hand-set per-gate constant reads off this row. Numeric values are launch defaults — placeholders pending calibration from deployment data [OPEN — flag for command center: the numbers, not the shape].
| Tier | Corroboration threshold (distinct run × provenance pairs) | Generation budget | Review requirement | Spot-check sampling | Model class allowed |
|---|---|---|---|---|---|
| 0 | 1 | unbounded (sandbox economics) | none | 2% | script / DeepSeek-tier |
| 1 | 2 | high — flood freely | none; reviewer samples | 5% | DeepSeek-tier |
| 2 | 3 | standard per-class budget | reviewer on sample | 10% | GLM-tier |
| 3 | 4, min_distinct_provenance ≥ 2 | reduced; per-item accounting | reviewer required, every item | 25% | GLM-tier (ceiling) |
| 4 | 5, min_distinct_provenance ≥ 2, consecutive_runs ≥ 2 | minimal; each generation justified | reviewer + command-center verdict | 100% | command center |
| 5 | n/a — default answer is BLOCK | zero until cleared | command center mandatory; ML notified (category-gated); unappealable in the operating layer | 100% | command center only |
Reading the columns:
- Corroboration threshold feeds
count_corroboration(GATES.md §3). Ramp ladders may loosen within a build phase but never below the tier-0 floor of the gate's grade minus two [OPEN — flag for command center: exact ramp-vs-tier floor rule; the lean is "ramp phases may drop at most 2 below steady tier, never below 1"]. - Generation budget bounds the proposer, not the judge — exploration stays unbounded in the sandbox; the budget applies at the commit boundary where generations become registered candidates.
- Review requirement names who must render an additional
verdictbefore commit. - Spot-check sampling is the monitoring layer's audit rate *of the judges themselves* — passed items re-examined by fresh context.
- Model class prices consequence-per-token (SYSTEM.md §7). GLM is the hard ceiling below the strategic layer; tier 4+ does not rent a bigger model, it climbs the escalation ladder to the command center.
Tier 5 semantics: a values-boundary touch is judged by the values gate first (GATES.md §6) — block + log, unappealable in the operating layer. What escalates is the question, not the action; the action waits. Bias toward inaction at the boundary (V6).
3. Migration table — the fifteen v1 hand-set thresholds → the dial
Grades shown as Q1 Q2 Q3 Q4 Q5 (Y/n). Where v2 differs from v1, that is the dial working: v1 numbers were cadence-denominated guesses; these are stakes-denominated.
| v1 gate | v1 setting | Stakes grade | Tier | v2 threshold | Notes |
|---|---|---|---|---|---|
| strategy-consensus | 3 / 7d consecutive-days | Y Y n n n | 2 | 3, consecutive_runs per ramp | Consecutive runs replace calendar-day adjacency (audit R6). |
| playbook | 2 / 14d + AUTO_PASS_SOURCES | Y n n n n | 1 | 2 | The source: auto-pass bypass is deleted (GATES.md §5). |
| worker-create | 2 / 14d + budget 60/65 | Y n n n n | 1 | 2 | Budget cap survives as a deterministic pre-filter (§4). |
| worker-update | 3 / 14d + dead MIN_DAYS | Y n n n n | 1 | 2 | v1's 3 was a guess; dead-code MIN_DAYS deleted. Updates to workers holding elevated scope re-grade Q4=Y → tier 2. |
| worker-retire | 5 / 30d + 2 lenses | Y Y n n n | 2 +1 decommit → 3 | 4 | v1's 5 approximated the decommit asymmetry by hand; the modifier makes it systematic. |
| voice-pattern | 2 / 14d | Y n n n n | 1 | 2 | |
| concept map | (template, ungated in v1 scripts) | Y n n n n | 1 | 2 | Additive knowledge; ramp builds at 1, graduates on coverage-stabilized (Decision 19). |
| pulse | (template, ungated in v1 scripts) | Y n n n n | 1 | 2 | Internal knowledge only. |
| content-pruning | 2 / 30d | Y Y n n n | 2 +1 decommit → 3 | 4 | v1 undervalued removing live pages; pruning is a decommit of world-facing state. |
| rich-media-opportunity | 2 / 7d + 50/day cap | Y n n n n | 1 | 2 | Daily cap survives as deterministic pre-filter, re-based to newly admitted keys (fixes audit R5). |
| rich-media-budget | pure budget filter | — | outside dial | — | Deterministic filter, §4. |
| link-quality | tiered DA/spam auto-pass + 2-date borderline | borderline: Y Y n n n | 2 | 3 | DA/spam metric tier survives as deterministic pre-filter; only the borderline judgment is consensus. |
| link-building-strategy | 3 / 7d | Y Y n n n | 2 | 3 | Outreach touches the world. |
| personality-ethos | 2 / 14d | Y Y n n n | 2 | 3 (locked phase: ramp raises to tier-4 controls) | Pre-lock-in the ramp phase drops to threshold 1 (personality is a correctable hypothesis); post-lock-in the ladder applies tier-4-equivalent settings — strongest consensus in the system. RESOLVED (ML, F3): locked-ethos changes are NOT values events — personality is fully system-governed through consensus gating; human input optional. |
| positioning-framework | 2 / 14d | Y n n n n | 1 | 2 | |
| domain-knowledge | (template, ungated in v1 scripts) | Y n n n n | 1 | 2 |
Grading notes:
- Q2 for strategy-consensus, pruning, link gates, and ethos: these decisions determine what the world sees even when the gate itself doesn't publish. The question is "does the decision touch the world," not "does this script call the publish API."
- Publishing itself (the gatekeeper) is not in this table because it was never a consensus gate; it is the mechanical three-nevers enforcement (V5) and grades Y Y n n Y → tier 5 controls on any values-clause hit, tier 2 controls otherwise.
- Escalation/brief machinery grades Q3=Y by construction; anything writing to ML's surface is tier ≥ 1 even when purely informational.
4. What stays outside the dial
The dial governs consensus — how much corroboration a judgment needs. Deterministic filters are not consensus and are not on the dial. They remain plain scripts with config-file parameters, unchanged in kind from v1:
| Filter | Parameters (in gate-ramp.yaml under params:) | Character |
|---|---|---|
| rich-media-budget | total_budget, conservation_at: 0.80, block_at: 1.00 | Arithmetic over spend records. Exit 4 on block. |
| link-quality DA tier | auto_pass_da: 30, auto_pass_spam: 30, borderline bounds | Metric comparison; only borderline cases enter consensus at tier 2. |
| worker-create budget | budget_soft: 60, budget_hard: 65 | Census count vs. cap; counts commissioned workers from the census, never files in a directory (fixes audit R7). |
| rich-media daily cap | daily_cap: 50 | Counts newly-admitted direction keys per UTC day. |
| velocity valves | per valve-map | Trust-canary-driven caps (V1 enforcement); close on degradation regardless of any tier. |
A deterministic filter can block and can never pass anything into commit by itself — passing is the dial's job. Filters compose with the dial as pre-filters (cheaper check first) and their blocks are ledgered as verdict/block with the filter named in reason.
5. Interaction with the ramp
The tier is the steady-state setting. gate-ramp.yaml ladders (GATES.md §2) modulate within a tier during build — lower thresholds, mandatory reviewers as compensation, evidence-based graduation back to the tier row. The dial answers "how much evidence at maturity"; the ramp answers "how do we get there from a cold start." Neither overrides the other's column: a build-phase gate at threshold 1 still carries its tier's review requirement or stricter, and tier-5 semantics are never ramp-loosened.
Worked example: strategy-consensus in build phase sits at tier 2 (steady threshold 3). Its ladder's build phase runs threshold 2 with consecutive_runs: 2 — within the "at most 2 below steady, never below 1" floor — and graduates back to the tier row when market data is live and the first graded outcomes exist. The dial never moved; the ramp moved within it.
6. Amendment
The five questions and the Q5 override are doctrine — command-center proposal, ML-reviewed (they define contact with the values boundary). Tier-table numerics are calibration — command-center amendable with ledgered rationale and a trial where one is possible. Per-gate stakes grades live in gate-ramp.yaml and change only through the worker-update path with the new grade in the proposal.
Provenance: rubric from LAWS.md "evidence scales with consequence" and the three real surfaces (SYSTEM.md §8); v1 settings from the gate-mechanic audit §1.1; decommit asymmetry from SYSTEM.md §6; model columns from SYSTEM.md §7 (ML, locked).