52 models · 5 families · canonical in kernel/models.yaml, rendered for orchestration.
Tap a status to filter. Specs verified from primary sources 2026-07-23 — IDs/prices drift, re-probe at intake.
WIREDAVAILABLE — NOT WIREDNON-AGENTIC · same keyHARDWARE-GATEDRESTRICTEDDEPRECATEDRETIRED(⚠ = the family offers it, the kit isn't using it — address in the kit)
OpenAI — GPT / Codex
5 wired · 14 total
ROLE IN KIT the roster's BUILDER — single-writer owned-file edits, bounded well-specified coding
KIT CONFIDENCE BUILDER. Reliably ships correct-STRUCTURE with wrong-DETAIL — numeric/unit-scaling bugs slip past its own selftest gate (Observify 2026-07-18: Hz->GHz off 1000, milli-C unconverted, `free -h` Gi/Mi regex). PAIR EVERY codex build with a read-only agy correctness review (+ grok security on a sensitive surface) before ship. Commit-incapable by default (L127 — cannot create .git/index.lock; the conductor adopt-commits).
Fast/floor tier — crank-turning in an ESTABLISHED lane (review cycles, checklist passes, extraction, classification, routing).
COST$1.0in /1M$6.0out /1M$0.1 cachedbatch: 50% off
CONFIDENCE
◆ KIT — measured
A REAL executor tier for pattern-following inside a piloted lane: +30% throughput vs Sol at flat error rate (terminus 2026-07-18). NOT for establishing a lane. If output drifts from the piloted shape, that IS its ceiling — step back up.
▲ BENCHMARK — external
GPQA Diamond92.3%SWE bench Pro62.7%Terminal bench 2.184.7%long context MRCR~41.3% (3rd-party, weaker at long lengths — flagged)
notes⚠ verify current ChatGPT gating (was Pro-gated at Feb 2026 launch per secondary sources).
gpt-oss-20b
WIRED
FREE local CODING floor (via `codex --oss`) — bounded, fully-specified coding: boilerplate, mechanical refactor, test scaffold, codegen, tight first-draft. Also the glue tier.
COSTFREEFREE — local compute only. Runs in ~16GB RAM (fits this Mac's 48GB).
CONFIDENCE
◆ KIT — measured
The free floor: every bounded well-specified coding ticket TRIES it before a paid token. NOT for ambiguous/architectural/high-stakes work — fills ambiguity with plausible-wrong; L92/L95 bite HARDEST here. MUST be oriented (seven levers) + audited at the SAME gates. Prove on one low-stakes unit before trust. ⚠ Ollama defaults to a small context window — set num_ctx>=32768 or it loses the spec.
▲ BENCHMARK — external
class~o3-miniMMLU85.3%GPQA Diamond71.5% (no-tools)AIME 2025 with tools98.7%
Larger free local model (~o4-mini class) — near-parity with o4-mini on core reasoning.
COSTFREEFREE local — but needs a single 80GB GPU (H100/MI300X class).
CONFIDENCE
◆ KIT — measured
⚠ NOT hostable on this Mac (needs ~80GB; the Mac has 48GB). Would be a stronger free coding tier on an 80GB+ host. Available to wire ONLY if a bigger box joins the estate.
▲ BENCHMARK — external
class~o4-miniMMLU90%GPQA Diamond80.1% (no-tools)AIME 2025 with tools97.9%SWE bench Pro16.2% (general model, not an agentic-coding specialist)
exhaustive specs ▾
ollama taggpt-oss:120b
params
total117B
active_per_token5.1B (MoE, 36 layers)
quantizationMXFP4 (MoE)
context window128000
modalities
inputtext
outputtext
reasoning
is_reasoning_modelyes
effort_tierslow, medium, high
featuresfunction calling, structured outputs, fine-tunable on a single H100, license: Apache-2.0
Cheap high-volume subagent/coding model (OpenAI's current 'mini' — the strongest mini for coding/computer-use/subagents).
COST$0.75in /1M$4.5out /1M$0.075 cached
CONFIDENCE
◆ KIT — measured
⚠ AVAILABLE — NOT WIRED. The kit's `mini-class` label is AMBIGUOUS: there is NO gpt-5.6-mini (the mini lineage stopped at 5.4). Decide whether the kit means gpt-5.4-mini (a real mini id, 400K ctx, cheaper) OR gpt-5.6-luna (the 5.6 floor). Wire one explicitly.
▲ BENCHMARK — external
noteOpenAI: 'strongest mini model for coding, computer use, subagent deployments'
Text → speech. Callable as a tool for voice output / narration.
COSTtts-1: $15/1M chars · tts-1-hd: $30/1M chars · gpt-4o-mini-tts: $0.60/1M in + $12/1M audio out
CONFIDENCE
◆ KIT — measured
TOOL-CALLABLE with the same key. gpt-4o-mini-tts is steerable (instruction-guided voice). Plus gpt-realtime-translate ($0.034/min) for live speech→speech translation.
TOOL-CALLABLE with the same key — a paid alternative to the local nomic-embed-text floor when a hosted embedder is wanted. 3-large supports the `dimensions` param to shorten output.
ROLE IN KIT in-family implementer / conductor / CEILING — exhausted LAST (most expensive), spent on what weaker vendors cannot do. Commit-capable (unlike codex).
KIT CONFIDENCE The ceiling tier and the conductor's own family (its scarcest pool BY CONSTRUCTION — the orchestrator spends it all day). Commit-capable where codex is not. Buffers -p output, so monitor claude workers via GROUND TRUTH (commits/ledger), not log tail. Route delegable side-tasks to OTHER families first (family_parity + quota thrift).
COST$5.0in /1M$25.0out /1Mcache write 5m: 6.25 · cache read: 0.5 · batch: 50% off · fast mode: $10/$50 flat (beta)
CONFIDENCE
◆ KIT — measured
Ceiling. Spent LAST, on what the weaker/cheaper vendors cannot do. ~4x less likely than Opus 4.7 to let code flaws pass — a strong auditor as well as builder.
notesRequires 30-day retention (a 'Covered Model') — NOT available under zero-data-retention. Sister model Mythos 5 = same weights, classifiers removed (restricted — see below).
claude-sonnet-5
WIRED
The kit's standard / implementer tier — strong agentic coding at a fraction of Opus cost.
COSTintro $2 in / $10 outintro through 2026-08-31: $2 in / $10 out · standard from 2026-09-01: $3 in / $15 out · cache read intro: 0.2 · batch: 50% off
CONFIDENCE
◆ KIT — measured
The default implementer for standard work; adaptive thinking ON by default; first Sonnet with xhigh effort + hi-res vision. NOTE: substantially WEAKER than Opus 4.8 at cyber-exploit dev (0% on a Firefox exploit set) — route authorized-security work to Opus/Mythos, not Sonnet.
The cheap fast tier for mechanical/light work. Pre-4.6 API shape: supports CLASSIC budget_tokens thinking, NO adaptive/effort. Anthropic's 'safest model yet' by internal metrics.
▲ BENCHMARK — external
SWE bench Verified73.3%Terminal bench40.2% (no thinking) / 41.75% (32K budget)
exhaustive specs ▾
aliasclaude-haiku-4-5
context window200000
max output64000
knowledge cutoffreliable 2025-02 / training 2025-07
notesThe only currently-wired model below 1M context / 128K output.
claude-mythos-5
⚠ RESTRICTED
Same weights as Fable 5 with SAFETY CLASSIFIERS REMOVED — for vetted cyberdefense/biosecurity work.
COST$10.0in /1M$50.0out /1M
CONFIDENCE
◆ KIT — measured
⚠ RESTRICTED — invite-only via Project Glasswing (~150 orgs). NO self-serve key, no Console toggle, no waitlist. The ONLY Claude with cyber safeguards lifted — directly relevant to Wolf's authorized bug-bounty/pentest work IF a Glasswing partnership is approved. Un-wireable today: no code path makes it selectable.
⚠ DEPRECATED — RETIRES 2026-08-05 (~2 weeks). 3x the price of 4.8 for less capability. If ANY kit config hardcodes `claude-opus-4-1`, that is a LIVE BUG to fix before it 404s.
⚠ RETIRED on the first-party API (2026-02-19). Listed ONLY so nothing references `claude-3-5-haiku-20241022`. (Still on Bedrock/Vertex with their own schedules.)
Anthropic ships NO non-agentic models — no image, video, audio, or embedding models. Claude is a text/vision-in → text family only.
COSTn/a
CONFIDENCE
◆ KIT — measured
For image/audio/video/embedding TOOLS, an agent on the Claude key must reach ANOTHER family (OpenAI/Google/xAI) or a local model. Claude itself accepts image + PDF INPUT (vision) but only emits text.
ROLE IN KIT adversarial SECURITY-review lens — the SECOND lens next to agy; also wide cheap parallel recon / N-variant prototyping.
KIT CONFIDENCE HIGH-RECALL / LOW-PRECISION and inconsistent — treat EVERY output as a LEAD, never a final answer. Adjudicate every grok finding against the actual code; NEVER block a ship on an unverified grok flag. It earns its seat by occasionally catching the real one nobody else did (Observify 2026-07-18: 2 rounds ALL false-positives, 1 round 2 REAL findings — a TOCTOU + a tail-path arbitrary-read the builder AND agy both missed, 1 round empty). SuperGrok is the DURABLE review pool — its window outlasts agy's; when agy caps for days, grok survives.
LANEgrok -p "<prompt>" -m <id> (OAuth ties to SuperGrok; --effort low|medium|high)
⚠ AVAILABLE — NOT WIRED. A candidate BIG-CONTEXT + cheaper grok lane (1M ctx vs 4.5's 500K). Also the silent redirect target for all retired grok-4/grok-3 slugs. Benchmarks mostly 3rd-party/low-confidence.
exhaustive specs ▾
context window1000000
modalities
inputtext, image
outputtext
reasoning
is_reasoning_modelyes
release~2026-04-30 (AA article; no x.ai/news post)
⚠ AVAILABLE — NOT WIRED. A cheap dedicated CODING lane worth piloting against codex/gpt-oss on bounded tickets. Predecessor grok-code-fast-1 scored 70.8% SWE-Bench Verified, 190 TPS (xAI's own claim, old slug).
⚠ AVAILABLE — NOT WIRED. A native parallel-agent deep-research lane; overlaps what the kit does with its own fan-out, so lower priority — but a candidate for single-call wide recon. Lower rate limits (9 req/s).
⚠ AVAILABLE — NOT WIRED. The non-reasoning variant's 'lowest hallucination / strict adherence' profile is interesting for schema-forced extraction; the reasoning variant is a cheaper 1M-ctx alternative to 4.5.
⚠ RETIRED 2026-05-15 — these slugs silently redirect to grok-4.3 / grok-build-0.1 and bill at the successor's rate. Listed so nothing hardcodes a dead grok-4/grok-3 id. (grok-3-mini fully retires 2026-08-15.)
xAI exposes NO standalone embeddings model (confirmed negative).
COSTn/a
CONFIDENCE
◆ KIT — measured
⚠ For RAG/embeddings, an agent on the xAI key must use ANOTHER family's embedder (OpenAI text-embedding-3, Google gemini-embedding-2) or the local nomic-embed-text. xAI's Collections API embeds internally with no selectable model/price.
ROLE IN KIT the free FLOOR — zero quota, spend before any API token. Glue + retrieval; Holt's on-device brain runs here.
KIT CONFIDENCE Free and unlimited, but WEAKER — fills ambiguity with plausible-wrong and emits confident, fluent, WRONG output with NO cap signature to warn you (L92/L95 bite hardest). Use for glue and bounded work; orient it (seven levers) and audit at the SAME build-failing gates. ⚠ Ollama defaults to a SMALL context window — a silently-truncated weak model is worse than a weak model; set num_ctx explicitly. (gpt-oss:20b / :120b live in the OpenAI family above — they are OpenAI open-weights also run here via `codex --oss`.)
LANEollama run <model> "<prompt>" (needs `ollama serve`; model pulled first)
qwen3:8b
WIRED
Holt's on-device BRAIN — closed-book graduate-exam answering on the Raspberry Pi (the Pi's model ceiling; 14b times out on the Pi). Hybrid thinking.
COSTFREEFREE local. q4_K_M 5.2GB / q8_0 8.9GB / fp16 16GB. Runs on a Raspberry Pi 5 (kit's own claim, not a vendor spec).
CONFIDENCE
◆ KIT — measured
The Pi's ACTIVE brain for Holt closed-book answering. Certifies Holt-as-deployed, so it is the HONEST exam model (using a bigger Mac model would inflate the grade). Instruct/thinking benchmarks NOT public at 8B; base-model figures only.
▲ BENCHMARK — external
note8B thinking-checkpoint benchmarks NOT public. Qwen3-8B-Base: MMLU 76.9, GPQA 44.4, GSM8K 89.8, MATH 60.8
Embedding backbone — text→vector for retrieval/RAG, clustering, classification (NOT a chat model).
COSTFREEFREE local. 274MB — runs on anything incl. a Pi.
CONFIDENCE
◆ KIT — measured
Pulled on the Mac; the retrieval/embedding tier. ⚠ REQUIRES task-instruction prefixes (search_document: / search_query: / clustering: / classification:) or quality degrades. ⚠ Ollama defaults its context to 2K vs the model's native 8192 — raise num_ctx for long-doc RAG.
▲ BENCHMARK — external
MTEB avg62.28 (768-dim); beats OpenAI ada-002 & text-embedding-3-small on MTEB + long-context LoCo
exhaustive specs ▾
params
total137M (nomic-bert-2048 encoder)
context window8192
embedding
dimensions768 (Matryoshka truncatable to 512/256/128/64)
⚠ AVAILABLE — NOT WIRED (not confirmed pulled). A candidate nano glue model below gpt-oss:20b for trivial classification/triage on tight-RAM hosts. Base figures only: MMLU 62.6, GPQA 28.3.
notesQwen3 also ships 0.6b / 4b / 14b / 32b dense + 30B-A3B / 235B-A22B MoE — other unused sizes if a bigger local tier is wanted.
Google — Gemini
1 wired · 13 total
ROLE IN KIT the big-window READ-ONLY CORRECTNESS / architecture auditor — whole-monorepo + full external docs in one context.
KIT CONFIDENCE The MOST RELIABLE correctness/architecture reviewer in the roster — clean, line-cited findings (Observify 2026-07-18: caught 4 real unit/logic bugs + a request-path issue codex shipped). READ-ONLY BY CONVENTION — enforce at the LAUNCH POSTURE (`--mode plan`), not the prompt: given write access it EDITED the single-writer file it was reviewing (correct fix, but an L46 breach). The big-window seat. LIGHT subscription — caps FAST (multi-day / ~6.5-day reset); a green 1-token probe does NOT prove run-room. Pin explicit model ids, not `-latest` aliases — a silent repoint mid-audit is a reproducibility risk. ⚠ the kit does not record which Gemini tier agy uses — confirm with `agy models`.
LANEagy -p "<prompt>" --mode plan --model <id> (READ-ONLY posture: --mode plan, NOT --dangerously-skip-permissions)
gemini-3.1-pro-preview
WIRED
Flagship Pro — the correctness/architecture auditor's tier: 1M context, strongest reasoning, best for whole-repo + docs audits.
COST$2.0in /1M$12.0out /1M$0.2 cachedover 200k: $4 in / $18 out · batch: $1/$6 · cache storage: $4.50/1M/hr · search grounding: 5k/mo free then $14/1k
CONFIDENCE
◆ KIT — measured
The likely agy auditor tier (big window + top correctness). ⚠ Still labeled 'Preview' (Gemini 3.5 Pro delayed). Deep Think mode is GATED (AI Ultra / early-access) — do not assume agy reaches it. Confirm the exact tier with `agy models`.
⚠ AVAILABLE — NOT WIRED. A cheaper, FASTER auditor/reviewer tier with the NEWEST cutoff (Mar 2026 vs Pro's Jan 2025) — attractive for high-volume review passes or when agy's Pro quota is scarce. Free tier available (rate-limited).
▲ BENCHMARK — external
SWE bench Pro58.7%OSWorld Verified83.0%Terminal bench 2.178.0%MLE Bench63.9%MRCR 128k91.8%
⚠ AVAILABLE — NOT WIRED. A very cheap Gemini glue/triage/classification tier ($0.30/$2.50) with a fresh cutoff — a candidate free-tier-eligible reviewer for high volume when the local floor is busy.
▲ BENCHMARK — external
SWE bench Pro54.2%Terminal bench 2.154.0%OSWorld Verified74.0%
Gemini fine-tuned for VULNERABILITY DISCOVERY & fixing — the CodeMender pilot model.
COST
CONFIDENCE
◆ KIT — measured
⚠ RESTRICTED — governments + trusted partners only (CodeMender pilot); NOT generally available, no public spec sheet. Directly relevant to Wolf's authorized bug-bounty/pentest work IF partner access is granted — un-wireable to `agy` today.
⚠ AVAILABLE — NOT WIRED. Older-gen fallback if the 3.x Pro quota is exhausted; classic thinking_budget (128-32768, no full disable). Weaker than 3.1 Pro.
exhaustive specs ▾
context window1048576
max output65536
knowledge cutoff2025-01
reasoning
thinking_budget128-32768 or -1 dynamic (cannot fully disable)
⚠ RETIRED — gemini-3-pro-preview shut down 2026-03-09 (migrate to 3.1 Pro Preview); Gemini 2.0 Flash/Flash-Lite shut down 2026-06-01. Listed so nothing targets a dead Gemini id.
TOOL-CALLABLE with the same key. Native audio (Veo 3.x). ⚠ Veo 2/3 past their announced shutdown but docs still list them — smoke-test before depending.
exhaustive specs ▾
modelsveo-3.1-generate-preview · veo-3.1-fast · veo-3.1-lite · gemini-omni-flash-preview (conversational video edit)
Text/image/audio/video/PDF → unified embedding vector. The first MULTIMODAL embedder — a strong RAG tool.
COSTgemini-embedding-2: text $0.20/1M · image $0.45/1M · audio $6.50/1M · video $12/1M · gemini-embedding-001: $0.15/1M text-only
CONFIDENCE
◆ KIT — measured
TOOL-CALLABLE with the same key. gemini-embedding-2 embeds ALL modalities into one space (128–3072 flexible dims, 8192-tok input) — the hosted upgrade over local nomic-embed-text for multimodal retrieval. (text-embedding-004 shut down 2026-01-14; -001 and -2 spaces are incompatible.)
Bidirectional real-time audio dialogue + live speech-to-speech translation.
COSTgemini-3.1-flash-live: audio $3/1M in ($0.005/min), $12/1M out ($0.018/min) · gemini-3.5-live-translate: ~$0.0053/$0.0315 per min
CONFIDENCE
◆ KIT — measured
TOOL-CALLABLE with the same key (WebSocket). A2A voice + 70-lang live translation. Also fills Gemini's STT gap — there is NO standalone Gemini transcription model; for pure STT use Google Cloud Speech-to-Text (a SEPARATE product/key) or send audio to a chat model.
Purpose-built action models: screen control and embodied/robot reasoning.
COSTgemini-2.5-computer-use-preview: $1.25/$10 per 1M · gemini-robotics-er-1.6-preview: $1/$5 per 1M
CONFIDENCE
◆ KIT — measured
TOOL-CALLABLE with the same key, but lean agentic (they emit ACTIONS): computer-use returns click/type/navigate for browser automation; robotics-ER does spatial/physical task planning. Borderline agentic — flagged.