Company Org Chart

the standing company structure — every registry model in a seat · canonical in kernel/company.yaml
← Kit Report ◈ Dashboard Model Registry

Company Org Chart

the standing company structure — every registry model in a seat · canonical in kernel/company.yaml. Tap a status to filter seats. 46 seats across 10 departments, 14 substitution ladders, 11 promotion-board tickets.

FAMILYOPENAI ANTHROPIC XAI GOOGLE LOCAL HUMAN

Chart 1 — The Org Chart

Executive at the top, six operating departments hanging below it, reserves and gated/retired seats in the band underneath.

EXECUTIVE

Judgment, sequencing, final accountability. The most expensive tokens in the building — spent on decisions, never on bulk.

CEO / Conductor

claude-fable-5
SEATED
ANTHROPIC

The org exists to concentrate judgment in one seat, and Fable 5 is the strongest reasoner in the registry by a wide margin — SWE-bench Pro 80.3% (11+ points above every other seat), Terminal-bench 2.1 88.0%, HLE-with-tools 64.5% — with always-on adaptive thinking and a 1M window. It is also the priciest seat ($10/$50) with its OWN smaller rate bucket: exactly the profile you conduct with and never bulk-build with (P2). The packet workspace already pins it; the chart makes that law.

Chief of Staff / Principal Reviewer & Closer

claude-opus-4-8
SEATED
ANTHROPIC

The ceiling executor for wrong-is-expensive rounds and the company's closer: measured ~4x less likely than Opus 4.7 to let code flaws pass, SWE-bench Verified 88.6%, GPQA 93.6% — and commit-capable, which the entire builder department is not (L127: codex cannot take .git/index.lock; Opus adopt-commits their trees). At $5/$25 it absorbs the senior-partner work the CEO's scarcer bucket must not.

The Approver (human)

none — Wolf
SEATED
HUMAN

L0: no agent self-approves. Charter ratification, promotion sign-off, spend posture, and anything irreversible terminate here.

Engineering

7 seats

The build value stream. OpenAI family holds it: the roster's measured BUILDER — structure-solid, detail-audited, every build paired with a cross-family review before ship.

FAMILY NOTE Known family defect (kit-measured): correct-STRUCTURE wrong-DETAIL — unit-scaling bugs pass its own selftests (Hz→GHz ×1000, milli-°C, Gi/Mi regex; Observify 2026-07-18). The Verification department owns that gate. Commit-incapable by default (L127).

VP Engineering / Lane Pilot

gpt-5.6-sol
SEATED
OPENAI

Every new lane is PILOTED by the strongest builder — Sol establishes the procedure and the judgment calls, then hands the crank down (kit-measured pattern, terminus 07-18). Top of the codex family on every axis: Terminal-bench 2.1 88.8% (91.9% Ultra), GPQA 94.6%, ARC-AGI-2 92.5%, SWE-bench Pro 64.6%. Ultra mode (4 parallel agents) is its org-chart superpower for the hardest tickets — used deliberately, piloted first (L52).

Senior Engineer (features)

gpt-5.6-terra
SEATED
OPENAI

The everyday workhorse for well-specified feature tickets: −1.2 SWE-bench-Pro and −1.4 Terminal-bench vs Sol at HALF the cost ($2.50/$15). Bulk feature work lands here so Sol stays free to pilot.

Line Engineer (the crank)

gpt-5.6-luna
SEATED
OPENAI

Kit-MEASURED, not inferred: inside a lane Sol established, Luna held +30% throughput at a flat error rate (terminus 2026-07-18, 14.2 vs 10.9 responses/min). Cheapest 5.6 tier ($1/$6). Two hard limits written into the seat: it does NOT establish lanes, and its long-context weakness (MRCR ~41.3%, flagged) keeps its tickets short-context. Drift from the piloted shape IS its ceiling — step the work back up.

Long-Haul Specialist (autonomous coding sessions)

gpt-5.3-codex
SEATED
OPENAI

The family's agentic-coding specialist for long unattended sessions — OSWorld-Verified 64.7%, Terminal-bench 77.3% at a 400K window and $1.75/$14. Holds marathon refactors and migration lanes where the 5.6 generalists would be overqualified per token.

Staff Engineer (in-family implementer, commit-capable builds)

claude-sonnet-5
SEATED
ANTHROPIC

The Claude-family implementer for builds that need commits, Claude-strengths, or a second opinion of a different lineage: best wired computer-use score in the registry (OSWorld-Verified 81.2%), SWE-bench Verified >80%, intro-priced $2/$10 through 2026-08-31. One explicit exclusion on the seat card: NOT for exploit-class security work (0% on the Firefox exploit set — route that to Opus 4.8).

Apprentice (free coding floor)

gpt-oss-20b
SEATED
OPENAI

Zero-quota local coding via `codex --oss` — every bounded, fully-specified ticket tries the apprentice BEFORE a paid token (the ladder's whole point). ~o3-mini class (GPQA 71.5% no-tools, AIME-with-tools 98.7%), fits the Mac's 48GB. Seat conditions: oriented (seven levers), num_ctx pinned ≥32768 (the silent-truncation scar), audited at the SAME build-failing gates as every paid seat. Doubles as the Mac's glue tier.

Junior Engineer (mechanical / chores / docs)

claude-haiku-4-5-20251001
SEATED
ANTHROPIC

~90% of Sonnet-4.5's agentic coding at $1/$5 (SWE-bench Verified 73.3%) — the cheap fast seat for mechanical edits, doc chores, and light subagent work where the free floor's reliability isn't enough but a 5.6 tier would be waste.

Verification & Audit

2 seats

Standing organ (non-negotiable). Cross-family from Engineering by law — the auditor never shares a family with the builder it audits.

Chief Auditor (correctness & architecture, READ-ONLY)

gemini-3.1-pro-preview
SEATED
GOOGLE

The kit's MEASURED most-reliable reviewer: clean, line-cited findings — caught 4 real unit/logic bugs plus a request-path issue that codex shipped and its own selftests missed (Observify 2026-07-18). The 1M window takes a whole monorepo plus external docs in one read. Seat is READ-ONLY at the launch posture (`--mode plan`), not by promise — given write access it edited the single-writer file it was reviewing (L46 breach). Owns the numeric-detail gate that Engineering's family defect requires. GPQA 94.3%, SWE-bench Verified 80.6%. Known seat risk: a LIGHT subscription that caps fast — the bench below exists because of it.

Volume Reviewer (high-throughput passes)

gemini-3.6-flash
BENCH
GOOGLE

The registry's FRESHEST knowledge cutoff (2026-03), OSWorld 83.0% (highest anywhere in the registry), SWE-bench Pro 58.7%, at half Pro's price with a free tier. The natural reviewer for high-volume passes and the first substitution when the Chief Auditor's quota caps — which is a WHEN, not an IF (multi-day resets, L146).

HIRE TICKET

Pilot on one repo A/B'd against the Chief Auditor on the SAME diff set; hire if line-cited precision holds.

Security

4 seats

Standing organ. Adversarial lens on every sensitive surface; findings are LEADS adjudicated against code, never ship-blockers on their own say-so.

Red Team Lead (adversarial security lens)

grok-4.5
SEATED
XAI

High-recall / low-precision BY MEASUREMENT, and seated for exactly that shape: across four Observify review rounds it produced two all-false-positive rounds, one empty — and one round with 2 REAL findings (a TOCTOU and a tail-path arbitrary-read) that the builder AND the Chief Auditor both missed. That is what a red team is for. Live web/X search gives it fresh-CVE context no other seat has. SuperGrok subscription quota OUTLASTS the auditor's — the durable review pool. Every finding is adjudicated (family law); AA-Omniscience hallucination 54% (flagged, up from 25% on 4.3) is why.

Exploit-Class Escalation

claude-opus-4-8
SEATED
ANTHROPIC

Dual-hatted from Executive: the only WIRED model with measured high-end cyber capability for Wolf's authorized engagements. Sonnet 5 is explicitly barred from this desk (0% exploit set); the two true specialists are still in clearance (below).

Security Specialist (safeguards-lifted)

claude-mythos-5
PENDING CLEARANCE
ANTHROPIC

Same weights as the CEO with safety classifiers removed — the ONLY Claude with cyber safeguards lifted, directly relevant to authorized bug-bounty/pentest work. Gate: Project Glasswing partnership (~150 orgs, invite-only, no waitlist). Un-wireable today; the badge IS the promotion.

Vulnerability-Discovery Specialist

gemini-3.5-flash-cyber
PENDING CLEARANCE
GOOGLE

Gemini fine-tuned for vulnerability discovery & fixing (the CodeMender pilot). Governments + trusted partners only; no public spec sheet. Recruiting target with a named gate — partner access — not a seat.

Research & Recon

3 seats

Wide reads, deep research, extraction. Cheap breadth so the executive floor never burns its bucket on reconnaissance.

Deep Research Squad (single-call fan-out)

grok-4.20-multi-agent-0309
BENCH
XAI

Native 4-or-16-agent parallel research in a single call, 1M window, $1.25/$2.50. Benched rather than seated because it overlaps what the kit's own fan-out already does well — hire on evidence, not novelty. Low rate ceiling (9 req/s) noted on the card.

HIRE TICKET

Wire for wide recon where ONE call genuinely beats kit-orchestrated fan-out; measure against a Workflow sweep on the same question.

Corpus Reader (big-context cheap recon)

grok-4.3
BENCH
XAI

Twice grok-4.5's context (1M vs 500K) at ~60% of its price. Also the silent redirect target of every retired grok slug — a seat card that exists partly so billing surprises have a name.

HIRE TICKET

Wire as the cheap 1M-window reader for whole-corpus sweeps when the Chief Auditor's window is needed elsewhere.

Strict Extractor (schema-forced, low-hallucination)

grok-4.20-reasoning-and-non-reasoning
BENCH
XAI

The non-reasoning variant's lowest-hallucination / strict-prompt-adherence profile is purpose-shaped for extraction pipelines; the reasoning variant is a cheaper 1M-ctx alternate. 'Interesting for schema-forced extraction' is the registry's own note — the ticket makes it testable.

HIRE TICKET

Pilot the non-reasoning variant on schema-forced extraction; hire if adherence beats the current glue tier on the same feeds.

Operations & Edge

6 seats

The free floor and the estate's edge. Glue, triage, retrieval, and the on-device brain — zero-quota rungs that every job tries before a paid token.

Edge Officer (Holt's on-device brain, the Pi)

qwen3:8b
SEATED
LOCAL

The Pi's model ceiling (14b times out) and therefore the HONEST exam model: it certifies Holt-as-deployed — grading on a bigger Mac model would inflate the grade. Hybrid thinking toggle, tool-calling, Apache-2.0, free.

Archivist (retrieval / embeddings backbone)

nomic-embed-text
SEATED
LOCAL

The only non-chat seat: text→vector for RAG/clustering/classification. MTEB 62.28 beats ada-002 and text-embedding-3-small; 274MB runs on anything including the Pi. Seat discipline on the card: task prefixes REQUIRED (search_document:/search_query:), and Ollama's 2K default context raised to the native 8192 for long-doc work.

Nano Triage (tight-RAM hosts)

qwen3:1.7b
BENCH
LOCAL

1.4GB Pi-class hybrid-thinking tier below gpt-oss-20b for trivial classification/labeling where even the 20B floor is oversized. Not confirmed pulled — the ticket says prove first.

HIRE TICKET

Pull + prove on one classification feed on the Pi-class host before any routing points at it.

Paid Nano Floor (API glue when local is busy)

gpt-5.4-nano
BENCH
OPENAI

The cheapest current GPT text model ($0.20/$1.25, 400K ctx) — the free-adjacent floor for trivial volume when the local rungs are saturated. No public benchmarks: volume-only mandate, never judgment.

HIRE TICKET

Wire as overflow glue; route only classification/label/normalize shapes.

Cheap Fresh Triage (multimodal, free-tier-eligible)

gemini-3.5-flash-lite
BENCH
GOOGLE

$0.30/$2.50, ~350 tok/s, 2026-03 cutoff, 1M window, computer-use built in — the strongest spec sheet of any floor-priced tier. Benched only because the free local rungs come first by law.

HIRE TICKET

Wire for high-volume triage needing a FRESH cutoff or multimodal input at floor price.

Subagent Fleet Default (candidate)

gpt-5.4-mini
BENCH
OPENAI

OpenAI's own positioning: 'strongest mini model for coding, computer use, subagent deployments' — $0.75/$4.50 at 400K ctx sits exactly in the subagent sweet spot between nano glue and the 5.6 line tiers.

HIRE TICKET

RESOLVES A FLAGGED AMBIGUITY: the kit's 'mini-class' label names no real 5.6 model (mini lineage stopped at 5.4). Wire gpt-5.4-mini explicitly OR redefine mini-class as luna — decide, then measure vs luna on cost/quality for subagent fan-outs.

Studio & Shared Services

9 seats

The non-agentic wing: tool models, not employees. Nothing here runs a loop or holds a ticket — any staffed agent with the family's key CALLS these desks mid-task (Wolf's rule: a subscription/API key makes them agentically reachable). Desks are picked per job by cost/quality; the substitution ladders below make the cross-family coverage explicit.

FAMILY GAPS Two structural gaps shape this wing: Anthropic ships NO non-agentic models (Claude staff reach across to OpenAI/Google/xAI or local for every media/embedding tool), and xAI ships NO standalone embedder (grok staff borrow OpenAI/Google/local for RAG). The org chart routes around both by design.

Image Desk

google · image generation
SEATED
GOOGLE

Nano Banana 2 (gemini-3.1-flash-image) leads the desk on layout/text-in-image strength at $0.045-0.151/img, with gemini-3-pro-image for 4K finals. gpt-image-2 is the OpenAI alternate (flagship quality, $0.006-0.211 by tier); grok-imagine-image is the budget bulk lane ($0.02 flat). ⚠ Imagen 4 shuts down 2026-08-17 — no new work.

covers
gemini-3.1-flash-imagegemini-3.1-flash-lite-imagegemini-3-pro-imagegemini-2.5-flash-image (legacy)

Video Desk

google · video generation (Veo)
SEATED
GOOGLE

Veo 3.1 holds the desk: NATIVE AUDIO (sora-2 has audio too, but Veo's lite/fast tiers give a $0.05-0.30/sec cost ladder in one family). sora-2/sora-2-pro is the OpenAI alternate ($0.10-0.70/sec); grok-imagine-video the budget lane ($0.05-0.08/sec). Heavy/slow either way — batch it (the registry's own note). ⚠ Veo 2/3 are past announced shutdown yet still doc-listed — smoke-test before depending.

covers
veo-3.1-generate-previewveo-3.1-fastveo-3.1-litegemini-omni-flash-preview (conversational edit)

Voice Desk (TTS + STT)

openai · speech-to-text
SEATED
OPENAI

Transcription anchors on whisper-1/gpt-4o-transcribe ($0.006/min; -diarize adds speakers; mini at $0.003/min for volume). Voice OUT is per-job: gpt-4o-mini-tts for steerable instruction-guided voice, gemini-3.1-flash-tts for 90+ languages/dual-speaker, xAI /v1/tts for voice cloning (eve/ara/leo/rex/sal). Music is Google-only (Lyria 3, $0.04-0.08/song).

covers
whisper-1gpt-4o-transcribegpt-4o-mini-transcribegpt-4o-transcribe-diarizegpt-realtime-whisper (streaming)

Embeddings Desk

nomic-embed-text
SEATED
LOCAL

The LOCAL floor stays the default (free, MTEB-competitive, already the Archivist). Hosted alternates when scale or hosting demands: OpenAI text-embedding-3-small ($0.02/1M, the price/perf pick) or 3-large ($0.13/1M, matryoshka dims); Google gemini-embedding-2 is the only MULTIMODAL embedder (text/image/audio/video/PDF → one vector space) — the pick the moment cross-modal RAG appears. xAI: none (confirmed negative).

Safety Gate (pre/post-flight moderation)

openai · moderation
SEATED
OPENAI

omni-moderation-latest is FREE — a zero-cost guard any agent can call before acting on or emitting content (harassment/hate/self-harm/sexual/violence, text+image). A free build-failing gate is the kit's favorite kind; wire it wherever generated content ships.

covers
omni-moderation-latest

Realtime Voice Desk

google · Live API / realtime voice
BENCH
GOOGLE

No current product needs a live-voice loop — benched, not seated. The desk exists so the capability is a lookup, not a re-research.

HIRE TICKET

Wire when a product needs live dialogue: bidirectional realtime audio + live speech-to-speech translation. OpenAI's realtime-whisper/translate lane is the alternate.

covers
gemini-3.1-flash-live-previewgemini-3.5-live-translate-previewgemini-2.5-flash-native-audio-preview

Action Models Desk (computer-use / robotics)

google · specialist (computer-use, robotics)
BENCH
GOOGLE

Specialists only earn seats by beating the incumbents' measured scores — the same rule as every other bench ticket.

HIRE TICKET

Purpose-built screen-control/embodied models — evaluate against the wired generalists' computer-use (sonnet-5 OSWorld 81.2, gemini-3.6-flash 83.0) before any hire; the generalists may already cover it.

covers
gemini-2.5-computer-use-preview-10-2025gemini-robotics-er-1.6-preview

Media Desks (OpenAI/xAI capability cards)

openai · image generationopenai · video generation (Sora)openai · text-to-speechopenai · embeddingsxai · image generationxai · video generationxai · speech (TTS + STT)google · text-to-speech & musicgoogle · embeddings (multimodal)
SEATED
OPENAIXAIGOOGLE

The alternate-family desks named in the ladders above — carried as first-class cards so every family's media reach is a lookup. Notable retirements on watch: gpt-image-1.5/-mini retire 2026-12-01; dall-e-2/3 already shut down.

covers
gpt-image-2gpt-image-1.5 (retires 2026-12-01)gpt-image-1-mini (retires 2026-12-01)sora-2sora-2-protts-1tts-1-hdgpt-4o-mini-ttstext-embedding-3-smalltext-embedding-3-largetext-embedding-ada-002grok-imagine-imagegrok-imagine-image-qualitygrok-imagine-videogrok-imagine-video-1.5gemini-3.1-flash-tts-previewgemini-2.5-flash-preview-ttsgemini-2.5-pro-preview-ttslyria-3-pro-previewlyria-3-clip-previewlyria-realtime-expgemini-embedding-2 (multimodal)gemini-embedding-001 (text)

(structural gap cards)

anthropic · (text-only family)xai · embeddings — NONE
SEATED
ANTHROPICXAI

Confirmed-negative cards, kept deliberately: they encode WHY the cross-family routing above is mandatory, so no future session re-derives the gap.

Reserves, Hardware-Gated & Offboarding

Reserves & Reproducibility

7 seats

Version-pinned fallbacks. Not staffed for new work; exist so a regression or an API-shape need has a named, tested landing spot.

Opus pin (if 4.8 regresses)

claude-opus-4-7
BENCH
ANTHROPIC

Same price as 4.8, strictly weaker — its ONLY value is version-pinned reproducibility. Never a capacity hire.

budget_tokens API shape (Opus)

claude-opus-4-5
BENCH
ANTHROPIC

Last Opus supporting BOTH classic budget_tokens AND effort — the landing spot if a component needs a fixed thinking-token ceiling. 200K ctx.

budget_tokens API shape (Sonnet)

claude-sonnet-4-5
BENCH
ANTHROPIC

Pre-4.6 Sonnet with classic thinking. EARLIEST retirement floor of any Active model (not before 2026-09-29) — plan exits, don't build on it.

no seat

claude-sonnet-4-6
BENCH
ANTHROPIC

More expensive than Sonnet 5's intro price and weaker — the registry's own verdict: no reason to add. Listed so nobody re-litigates it.

Gemini Pro fallback

gemini-2.5-pro
BENCH
GOOGLE

Older-gen Pro if the 3.x quota exhausts entirely: GPQA 86.4%, classic thinking_budget shape. A brownout seat, not a hire.

no seat (superseded)

gemini-3.5-flash
BENCH
GOOGLE

Strictly dominated by 3.6 Flash (cheaper output, fresher cutoff). Named here so the dominated choice is never made by accident.

Cheap Coding Pilot (challenger)

grok-build-0.1
BENCH
XAI

xAI's dedicated coding tier at $1/$2 — predecessor claimed 70.8% SWE-bench Verified at 190 TPS. The named challenger to the Apprenticeship: pilot 3 bounded tickets vs gpt-oss-20b; hire only on gate evidence.

Hardware-Gated

1 seat

Apprentice, Senior (free coding — bigger box)

gpt-oss-120b
HARDWARE GATED
OPENAI

~o4-mini class free coding (GPQA 80.1% no-tools) — a strict upgrade to the Apprenticeship the estate cannot host (needs ~80GB; the Mac has 48GB). Promotion is blocked on HARDWARE, not benchmarks: an 80GB+ host joining the estate unlocks it automatically.

Offboarding

4 seats

Do-not-staff. Every id here is a live-bug tripwire: grep configs for them at intake.

claude-opus-4-1
OFFBOARDING
ANTHROPIC

RETIRES 2026-08-05 (~2 weeks) at 3x the price of Opus 4.8. Any config hardcoding it is a LIVE BUG about to 404.

claude-haiku-3-5
OFFBOARDING
ANTHROPIC

Retired on the first-party API 2026-02-19. Listed only so nothing references the dead id.

grok-4-and-grok-3-legacy
OFFBOARDING
XAI

Retired 2026-05-15; slugs SILENTLY redirect to grok-4.3/grok-build-0.1 and bill at successor rates — a billing surprise wearing a familiar name.

gemini-3-pro-preview-and-2.0-legacy
OFFBOARDING
GOOGLE

Shut down (2026-03-09 / 2026-06-01). Dead ids fail loudly at least — but migrate any references to 3.1 Pro Preview.

Chart 2 — Substitutions & Promotions

Who covers a seat when it caps, and how a seat's status changes.

Substitution ladders

A substitution is a CAP/DEATH response, not a preference: when a seat's quota signature fires (L5/L146), sweep every lane on that executor to the ladder below in one move. One retry per rung; skipping down is free; stepping UP is the fix for drift. Kit-measured confidence outranks benchmarks when they disagree.

CEO / Conductor (claude-fable-5)claude-opus-4-8claude-sonnet-5

Fable's own bucket caps independently of Opus's. Opus conducts with modest loss (SWE-Pro 69.2 vs 80.3) at half the price; Sonnet 5 is an emergencies-only conductor — fine judgment, thinner ceiling. Never conduct from the builder family (judgment seat stays in-family, KB-family-parity applies to WORKERS).

Lane Pilot (gpt-5.6-sol)gpt-5.6-terraclaude-sonnet-5

Terra pilots at −1.2 SWE-Pro / −1.4 Terminal-bench for half cost — acceptable for most lanes. If the WHOLE codex family is capped (5h window), the pilot crosses families to Sonnet 5 (commit-capable, so the adopt-commit step disappears too).

Feature seat (gpt-5.6-terra)gpt-5.6-lunagpt-5.3-codexgpt-oss-20b

Luna substitutes UP-conditional: only in lanes already piloted; drift = step back up. gpt-oss-20b takes only the fully-specified residue.

Crank (gpt-5.6-luna)gpt-5.6-terragpt-oss-20b

The crank substitutes UPWARD by default (Terra) — the seat's failure mode is drift, and the fix is capability, not a cheaper twin. The free floor covers only bounded checklist passes.

Chief Auditor (gemini-3.1-pro-preview)gemini-3.6-flashclaude-opus-4-8

The auditor's quota caps FAST (light subscription, multi-day reset) — this ladder fires often. 3.6 Flash first (in-family, fresher cutoff, unmeasured precision — pair its first rounds with spot-checks). Opus second: cross-family from the builder holds (Claude audits OpenAI builds). grok-4.5 is NEVER sole auditor — measured precision too low; it stays the second lens.

Red Team (grok-4.5)claude-opus-4-8gemini-3.1-pro-preview

SuperGrok quota is the durable pool (outlasts agy) — this ladder fires RARELY. When it does: Opus for exploit-class depth, the Chief Auditor for a security-flavored correctness pass. Both are precision lenses, not recall lenses — expect fewer, truer leads while the seat is empty. Pending clearances (mythos-5, flash-cyber) slot here the day their badges land.

Free floor (gpt-oss-20b)qwen3:1.7bgpt-5.4-nanogemini-3.5-flash-lite

Ollama down = worker_death, restart serve first. Then: nano-local for trivial shapes, the paid nano floors for volume. All floor substitutions keep the same build-failing gates — a cheaper seat never buys a cheaper gate.

Edge Officer (qwen3:8b, Pi)no substitute — by design

NO substitute by design: the seat certifies Holt-as-deployed, and any bigger stand-in falsifies the certification (the honest-exam rule). If the Pi is down, the exam waits.

Archivist (nomic-embed-text)no substitute — by design

No wired substitute; embeddings are cheap and local. If it ever fails, re-embedding with a different backbone is a MIGRATION (vectors don't mix across models), not a substitution — plan it as one.

Image Desk (Nano Banana 2)openai · image generation (gpt-image-2)xai · image generation (grok-imagine, budget)

Per-JOB choice more than failover: Google for layout/text-in-image, OpenAI for flagship quality tiers, xAI for $0.02 bulk. All three live behind keys the staff already hold.

Video Desk (Veo 3.1)openai · video generation (sora-2/-pro)xai · video generation (grok-imagine-video)

Cost ladders inside each family (lite/fast/pro tiers); batch everything — the desk is slow by nature.

Voice Desk (whisper/gpt-4o-transcribe)xai · speech (STT $0.10/hr + cloning TTS)google · text-to-speech & music (90+ langs, Lyria music)

STT failover is real failover; TTS is per-job (steerable voice = OpenAI, languages/dual-speaker = Google, cloning = xAI, music = Google only).

Embeddings Desk (nomic local)openai · embeddings (3-small $0.02/1M)google · embeddings (multimodal)

Same migration caveat as the Archivist: switching embedders re-embeds the corpus. Choose ONCE per corpus; the ladder is for NEW corpora, not mid-life swaps. xAI has no embedder (confirmed).

Safety Gate (omni-moderation)no substitute — by design

FREE and unique — no substitute needed; if it's down, the gate fails CLOSED (hold the publish), never open.

Promotion board

Benchmarks NOMINATE; pilots + gates PROMOTE; the human ratifies (L0). Every promotion below is a ticket with a measurable trigger — when it fires, run the pilot, show the gate evidence, record the outcome in models.yaml confidence.kit, and re-render this chart. Demotions use the same mechanism in reverse; a flagged benchmark regression opens a review, never an automatic bench.

Promotions 6

PROMO-1
gpt-5.6-luna

FROM Line Engineer (crank) TO Senior Engineer (features) — piloted lanes only

TRIGGER Terminus-style evidence already exists (+30% throughput, flat errors). Promote per-LANE: after any lane where Luna holds N=3 consecutive feature tickets with zero step-ups, it keeps that lane at Terra's tier.

BENCHMARK BASIS SWE-Pro 62.7 within 0.7 of Terra; the gap that matters (MRCR long-context) is avoidable by lane design.

PROMO-2
gemini-3.6-flash

FROM bench (Volume Reviewer) TO co-Auditor

TRIGGER A/B on the SAME diff set as the Chief Auditor: hire if line-cited precision is within tolerance; promote to co-equal if it also catches something Pro missed (its 14-month-fresher cutoff makes that plausible on new-API surfaces).

BENCHMARK BASIS OSWorld 83.0 (registry-best), SWE-Pro 58.7, cutoff 2026-03 vs Pro's 2025-01.

PROMO-3
grok-build-0.1

FROM bench (challenger) TO Apprentice co-seat (paid-cheap coding floor)

TRIGGER 3 bounded tickets vs gpt-oss-20b, same gates: promote if gate-pass rate strictly beats the free floor AND cost/ticket stays under the Junior Engineer's.

BENCHMARK BASIS Predecessor's 70.8% SWE-bench Verified at 190 TPS (vendor claim, old slug — hence pilot, not trust).

PROMO-4
gpt-5.4-mini

FROM bench TO Subagent Fleet Default

TRIGGER Resolve the flagged mini-class ambiguity (models.yaml): wire it explicitly, run one fan-out A/B vs luna; promote on cost-per-accepted-output.

BENCHMARK BASIS Vendor positioning only (no public numbers) — which is precisely why the ticket demands a measured A/B.

PROMO-5
gpt-oss-120b

FROM hardware_gated TO Apprentice, Senior

TRIGGER HARDWARE, not benchmarks: an 80GB+ host joins the estate. Then prove on one low-stakes unit like any floor hire.

BENCHMARK BASIS GPQA 80.1 vs 20b's 71.5 (no-tools) — a whole tier of free capability waiting on RAM.

PROMO-6
claude-mythos-5gemini-3.5-flash-cyber

FROM pending_clearance TO Security Specialists

TRIGGER Access grants (Glasswing partnership / CodeMender pilot). The paperwork IS the promotion; capability is already established by lineage.

BENCHMARK BASIS Mythos-5 = Fable-5 weights (SWE-Pro 80.3 lineage); flash-cyber has no public sheet — clearance-gated evaluation on arrival.

Demotion watch 3

DEMO-1
grok-4.5

WATCH AA-Omniscience hallucination 54% (flagged, was 25% on 4.3). The seat TOLERATES low precision by design, but if adjudication cost (false-positive burden per real find) doubles from the 07-18 baseline, demote the seat to grok-4.3 and re-measure.

DEMO-2
claude-sonnet-5

WATCH Price cliff 2026-09-01 ($2/$10 → $3/$15): re-run the Staff Engineer seat-cost math then. Also standing: the exploit-desk bar (0% Firefox set) — any routing that drifts security work here is a config bug.

DEMO-3
gpt-5.3-codex

WATCH ⚠ verify current ChatGPT gating (was Pro-gated at launch per secondary sources). If the lane's access changes, the Long-Haul seat falls back to terra until re-verified.

Watchlist 2

WATCH-1
gemini-3.1-pro-preview

WATCH Still 'Preview' (3.5 Pro delayed) — a GA replacement or a silent repoint is a reproducibility risk; pin ids, re-probe at intake, and treat a Pro-GA release as an automatic review of the Chief Auditor seat.

WATCH-2
studio retirement calendar

WATCH Imagen 4 shuts down 2026-08-17; gpt-image-1.5 and gpt-image-1-mini retire 2026-12-01; dall-e-2/3 and grok-2-image already gone; Veo 2/3 past announced shutdown but still doc-listed (smoke-test). Grep any media pipeline for these ids at intake — the same live-bug tripwire rule as the offboarding department.

updated just nownext 3m 00s