Who covers a seat when it caps, and how a seat's status changes.
A substitution is a CAP/DEATH response, not a preference: when a seat's quota signature fires (L5/L146), sweep every lane on that executor to the ladder below in one move. One retry per rung; skipping down is free; stepping UP is the fix for drift. Kit-measured confidence outranks benchmarks when they disagree.
CEO / Conductor (claude-fable-5)→claude-opus-4-8→claude-sonnet-5
Fable's own bucket caps independently of Opus's. Opus conducts with modest loss (SWE-Pro 69.2 vs 80.3) at half the price; Sonnet 5 is an emergencies-only conductor — fine judgment, thinner ceiling. Never conduct from the builder family (judgment seat stays in-family, KB-family-parity applies to WORKERS).
Lane Pilot (gpt-5.6-sol)→gpt-5.6-terra→claude-sonnet-5
Terra pilots at −1.2 SWE-Pro / −1.4 Terminal-bench for half cost — acceptable for most lanes. If the WHOLE codex family is capped (5h window), the pilot crosses families to Sonnet 5 (commit-capable, so the adopt-commit step disappears too).
Feature seat (gpt-5.6-terra)→gpt-5.6-luna→gpt-5.3-codex→gpt-oss-20b
Luna substitutes UP-conditional: only in lanes already piloted; drift = step back up. gpt-oss-20b takes only the fully-specified residue.
Crank (gpt-5.6-luna)→gpt-5.6-terra→gpt-oss-20b
The crank substitutes UPWARD by default (Terra) — the seat's failure mode is drift, and the fix is capability, not a cheaper twin. The free floor covers only bounded checklist passes.
Chief Auditor (gemini-3.1-pro-preview)→gemini-3.6-flash→claude-opus-4-8
The auditor's quota caps FAST (light subscription, multi-day reset) — this ladder fires often. 3.6 Flash first (in-family, fresher cutoff, unmeasured precision — pair its first rounds with spot-checks). Opus second: cross-family from the builder holds (Claude audits OpenAI builds). grok-4.5 is NEVER sole auditor — measured precision too low; it stays the second lens.
Red Team (grok-4.5)→claude-opus-4-8→gemini-3.1-pro-preview
SuperGrok quota is the durable pool (outlasts agy) — this ladder fires RARELY. When it does: Opus for exploit-class depth, the Chief Auditor for a security-flavored correctness pass. Both are precision lenses, not recall lenses — expect fewer, truer leads while the seat is empty. Pending clearances (mythos-5, flash-cyber) slot here the day their badges land.
Free floor (gpt-oss-20b)→qwen3:1.7b→gpt-5.4-nano→gemini-3.5-flash-lite
Ollama down = worker_death, restart serve first. Then: nano-local for trivial shapes, the paid nano floors for volume. All floor substitutions keep the same build-failing gates — a cheaper seat never buys a cheaper gate.
Edge Officer (qwen3:8b, Pi)→no substitute — by design
NO substitute by design: the seat certifies Holt-as-deployed, and any bigger stand-in falsifies the certification (the honest-exam rule). If the Pi is down, the exam waits.
Archivist (nomic-embed-text)→no substitute — by design
No wired substitute; embeddings are cheap and local. If it ever fails, re-embedding with a different backbone is a MIGRATION (vectors don't mix across models), not a substitution — plan it as one.
Image Desk (Nano Banana 2)→openai · image generation (gpt-image-2)→xai · image generation (grok-imagine, budget)
Per-JOB choice more than failover: Google for layout/text-in-image, OpenAI for flagship quality tiers, xAI for $0.02 bulk. All three live behind keys the staff already hold.
Video Desk (Veo 3.1)→openai · video generation (sora-2/-pro)→xai · video generation (grok-imagine-video)
Cost ladders inside each family (lite/fast/pro tiers); batch everything — the desk is slow by nature.
Voice Desk (whisper/gpt-4o-transcribe)→xai · speech (STT $0.10/hr + cloning TTS)→google · text-to-speech & music (90+ langs, Lyria music)
STT failover is real failover; TTS is per-job (steerable voice = OpenAI, languages/dual-speaker = Google, cloning = xAI, music = Google only).
Embeddings Desk (nomic local)→openai · embeddings (3-small $0.02/1M)→google · embeddings (multimodal)
Same migration caveat as the Archivist: switching embedders re-embeds the corpus. Choose ONCE per corpus; the ladder is for NEW corpora, not mid-life swaps. xAI has no embedder (confirmed).
Safety Gate (omni-moderation)→no substitute — by design
FREE and unique — no substitute needed; if it's down, the gate fails CLOSED (hold the publish), never open.
Benchmarks NOMINATE; pilots + gates PROMOTE; the human ratifies (L0). Every promotion below is a ticket with a measurable trigger — when it fires, run the pilot, show the gate evidence, record the outcome in models.yaml confidence.kit, and re-render this chart. Demotions use the same mechanism in reverse; a flagged benchmark regression opens a review, never an automatic bench.