Model Registry

every model the kit can use — confidence · specialty · cost
← Kit Report ◈ Dashboard

Model Registry

52 models · 5 families · canonical in kernel/models.yaml, rendered for orchestration. Tap a status to filter. Specs verified from primary sources 2026-07-23 — IDs/prices drift, re-probe at intake.

WIRED AVAILABLE — NOT WIRED NON-AGENTIC · same key HARDWARE-GATED RESTRICTED DEPRECATED RETIRED (⚠ = the family offers it, the kit isn't using it — address in the kit)

OpenAI — GPT / Codex

5 wired · 14 total

ROLE IN KIT the roster's BUILDER — single-writer owned-file edits, bounded well-specified coding

KIT CONFIDENCE BUILDER. Reliably ships correct-STRUCTURE with wrong-DETAIL — numeric/unit-scaling bugs slip past its own selftest gate (Observify 2026-07-18: Hz->GHz off 1000, milli-C unconverted, `free -h` Gi/Mi regex). PAIR EVERY codex build with a read-only agy correctness review (+ grok security on a sensitive surface) before ship. Commit-incapable by default (L127 — cannot create .git/index.lock; the conductor adopt-commits).

LANE codex exec -m <id> (ALWAYS pass -m explicitly on detached launches)

gpt-5.6-sol

WIRED

Flagship builder — hard/core rounds; PILOTS a new lane (establishes the procedure + judgment calls).

COST$5.0in /1M $30.0out /1M $0.5 cachedbatch: 50% off ($2.50/$15) · overage: >272K input tokens billed 2x in / 1.5x out
CONFIDENCE
◆ KIT — measured

Highest codex tier. Use to PILOT a lane, then hand the crank to Luna. Structure-solid, audit the numeric detail.

▲ BENCHMARK — external
GPQA Diamond94.6%SWE bench Pro64.6%Terminal bench 2.188.8% (Ultra mode 91.9%)FrontierMath v2 T1-389%ARC AGI 292.5%AA Coding Agent Index80
exhaustive specs ▾
context window1050000
max output128000
knowledge cutoff2026-02-16
modalities
  • input text, image
  • output text
reasoning
  • is_reasoning_model yes
  • effort_tiers none, low, medium, high, xhigh (API), Ultra runs 4 parallel agents (product-level)
featuresfunction/tool calling, structured outputs / JSON schema, streaming, vision (image in), prompt caching, Batch API, Responses-API tools (web/file search, code interpreter, hosted shell, apply patch, computer use, MCP, programmatic tool calling in a V8 sandbox)
no supportfine-tuning, audio/video
rate limitsTier 5: 15,000 RPM / 40M TPM / 15B batch-queue tokens (scales down by tier)
releaseGA 2026-07-09 (preview 2026-06-25)
notesBare alias `gpt-5.6` routes to Sol. System card: High capability in Cybersecurity + Bio/Chem risk (not Critical).

gpt-5.6-terra

WIRED

Value tier — bulk, fully-specified features; ~half Sol's cost, ~GPT-5.5-competitive.

COST$2.5in /1M $15.0out /1M $0.25 cachedbatch: 50% off
CONFIDENCE
◆ KIT — measured

The everyday workhorse for well-specified feature tickets. Same build-audit pairing as Sol.

▲ BENCHMARK — external
GPQA Diamond92.9%SWE bench Pro63.4%Terminal bench 2.187.4%AA Coding Agent Index77.4
exhaustive specs ▾
context window1050000
max output128000
knowledge cutoff2026-02-16
modalities
  • input text, image
  • output text
reasoning
  • is_reasoning_model yes
  • effort_tiers configurable none..max
featuressame tool/feature surface as Sol
no supportfine-tuning
rate limitsTier 5: 15,000 RPM / 40M TPM
releaseGA 2026-07-09
provenancesource ↗

gpt-5.6-luna

WIRED

Fast/floor tier — crank-turning in an ESTABLISHED lane (review cycles, checklist passes, extraction, classification, routing).

COST$1.0in /1M $6.0out /1M $0.1 cachedbatch: 50% off
CONFIDENCE
◆ KIT — measured

A REAL executor tier for pattern-following inside a piloted lane: +30% throughput vs Sol at flat error rate (terminus 2026-07-18). NOT for establishing a lane. If output drifts from the piloted shape, that IS its ceiling — step back up.

▲ BENCHMARK — external
GPQA Diamond92.3%SWE bench Pro62.7%Terminal bench 2.184.7%long context MRCR~41.3% (3rd-party, weaker at long lengths — flagged)
exhaustive specs ▾
context window1050000
max output128000
knowledge cutoff2026-02-16
modalities
  • input text, image
  • output text
reasoning
  • is_reasoning_model yes
  • effort_tiers configurable
featuressame tool/feature surface as Sol/Terra
no supportfine-tuning
rate limitsTier 1-5: 500-30,000 RPM / 500K-180M TPM
releaseGA 2026-07-09
provenancesource ↗

gpt-5.3-codex

WIRED

Agentic-coding specialist — long autonomous coding sessions.

COST$1.75in /1M $14.0out /1M $0.175 cachedbatch: 50% off ($0.875/$7)
CONFIDENCE
◆ KIT — measured

The codex-family coding specialist; same structure-solid / audit-the-detail posture as the 5.6 tiers.

▲ BENCHMARK — external
SWE bench Pro public56.8%Terminal bench 2.077.3%OSWorld Verified64.7%
exhaustive specs ▾
context window400000
max output128000
knowledge cutoff2025-08-31
modalities
  • input text, image
  • output text
reasoning
  • is_reasoning_model yes
  • effort_tiers low, medium, high, xhigh
featuresstreaming, function calling, structured outputs
no supportfine-tuning, predicted outputs
release2026-02-05
notes⚠ verify current ChatGPT gating (was Pro-gated at Feb 2026 launch per secondary sources).

gpt-oss-20b

WIRED

FREE local CODING floor (via `codex --oss`) — bounded, fully-specified coding: boilerplate, mechanical refactor, test scaffold, codegen, tight first-draft. Also the glue tier.

COSTFREE FREE — local compute only. Runs in ~16GB RAM (fits this Mac's 48GB).
CONFIDENCE
◆ KIT — measured

The free floor: every bounded well-specified coding ticket TRIES it before a paid token. NOT for ambiguous/architectural/high-stakes work — fills ambiguity with plausible-wrong; L92/L95 bite HARDEST here. MUST be oriented (seven levers) + audited at the SAME gates. Prove on one low-stakes unit before trust. ⚠ Ollama defaults to a small context window — set num_ctx>=32768 or it loses the spec.

▲ BENCHMARK — external
class~o3-miniMMLU85.3%GPQA Diamond71.5% (no-tools)AIME 2025 with tools98.7%
exhaustive specs ▾
ollama taggpt-oss:20b
params
  • total 21B
  • active_per_token 3.6B (MoE, 24 layers)
quantizationnative MXFP4 (MoE) / BF16
context window128000
knowledge cutoffnot public (model card 2025-08)
modalities
  • input text
  • output text
reasoning
  • is_reasoning_model yes
  • effort_tiers low, medium, high
  • note harmony chat format exposes full CoT
featuresfunction/tool calling, structured outputs, fine-tunable, license: Apache-2.0
release2025-08-05

gpt-oss-120b

⚠ HARDWARE-GATED

Larger free local model (~o4-mini class) — near-parity with o4-mini on core reasoning.

COSTFREE FREE local — but needs a single 80GB GPU (H100/MI300X class).
CONFIDENCE
◆ KIT — measured

⚠ NOT hostable on this Mac (needs ~80GB; the Mac has 48GB). Would be a stronger free coding tier on an 80GB+ host. Available to wire ONLY if a bigger box joins the estate.

▲ BENCHMARK — external
class~o4-miniMMLU90%GPQA Diamond80.1% (no-tools)AIME 2025 with tools97.9%SWE bench Pro16.2% (general model, not an agentic-coding specialist)
exhaustive specs ▾
ollama taggpt-oss:120b
params
  • total 117B
  • active_per_token 5.1B (MoE, 36 layers)
quantizationMXFP4 (MoE)
context window128000
modalities
  • input text
  • output text
reasoning
  • is_reasoning_model yes
  • effort_tiers low, medium, high
featuresfunction calling, structured outputs, fine-tunable on a single H100, license: Apache-2.0
release2025-08-05

gpt-5.4-mini

⚠ AVAILABLE — NOT WIRED

Cheap high-volume subagent/coding model (OpenAI's current 'mini' — the strongest mini for coding/computer-use/subagents).

COST$0.75in /1M $4.5out /1M $0.075 cached
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. The kit's `mini-class` label is AMBIGUOUS: there is NO gpt-5.6-mini (the mini lineage stopped at 5.4). Decide whether the kit means gpt-5.4-mini (a real mini id, 400K ctx, cheaper) OR gpt-5.6-luna (the 5.6 floor). Wire one explicitly.

▲ BENCHMARK — external
noteOpenAI: 'strongest mini model for coding, computer use, subagent deployments'
exhaustive specs ▾
context window400000
max output128000
knowledge cutoff2025-08-31
modalities
  • input text, image
  • output text
reasoning
  • is_reasoning_model yes
  • effort_tiers none (default), low, medium, high, xhigh
releasesnapshot gpt-5.4-mini-2026-03-17
provenancesource ↗

gpt-5.4-nano

⚠ AVAILABLE — NOT WIRED

True floor — cheapest current GPT-family text model for trivial high-volume work.

COST$0.2in /1M $1.25out /1M $0.02 cached
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. A candidate free-adjacent cheap floor below gpt-5.4-mini for glue/classification if the local tier is busy.

▲ BENCHMARK — external
notenot public
exhaustive specs ▾
context window400000
max output128000
knowledge cutoff2025-08-31
modalities
  • input text, image
  • output text
release5.4 generation
provenancesource ↗

openai · image generation

NON-AGENTIC · same key

Text/image → image generation. Callable as a tool for creating/editing visuals.

COSTgpt-image-2 1024²: $0.006 / $0.053 / $0.211 (low/med/high) · gpt-image-1-mini: $0.005–0.052 (cheapest)
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same OPENAI_API_KEY the codex agents use — not an executor itself. Invoke via the Images/Responses API.

exhaustive specs ▾
modelsgpt-image-2 (flagship, snapshot 2026-04-21, $8/$30 per 1M) · gpt-image-1.5 & gpt-image-1-mini (retire 2026-12-01) · dall-e-2/3 SHUT DOWN 2026-05-12
modalitytext + image → image (png/jpeg/webp, ≤50MB in)

openai · video generation (Sora)

NON-AGENTIC · same key

Text/image → video with audio. Callable as a tool for generated clips.

COSTsora-2: $0.10/sec (720p) · sora-2-pro: $0.30–0.70/sec by resolution · batch ~half
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same key. Heavy/slow — batch it.

exhaustive specs ▾
modelssora-2 (720×1280 / 1280×720) · sora-2-pro (up to 1920×1080)
modalitytext / image → video + audio
provenancesource ↗

openai · speech-to-text

NON-AGENTIC · same key

Audio → text transcription. Callable as a tool for voice input / meeting notes / captions.

COSTwhisper-1 & gpt-4o-transcribe: $0.006/min · gpt-4o-mini-transcribe: $0.003/min · realtime-whisper: $0.017/min
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same key. gpt-4o-transcribe-diarize adds speaker labels.

exhaustive specs ▾
modelswhisper-1 (25MB cap, multilingual) · gpt-4o-transcribe · gpt-4o-mini-transcribe · gpt-4o-transcribe-diarize · gpt-realtime-whisper (streaming)
modalityaudio → text
provenancesource ↗

openai · text-to-speech

NON-AGENTIC · same key

Text → speech. Callable as a tool for voice output / narration.

COSTtts-1: $15/1M chars · tts-1-hd: $30/1M chars · gpt-4o-mini-tts: $0.60/1M in + $12/1M audio out
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same key. gpt-4o-mini-tts is steerable (instruction-guided voice). Plus gpt-realtime-translate ($0.034/min) for live speech→speech translation.

exhaustive specs ▾
modelstts-1 · tts-1-hd (9 voices) · gpt-4o-mini-tts (13 voices, 2000-tok cap)
modalitytext → audio (mp3/opus/aac/flac/wav/pcm)
provenancesource ↗

openai · embeddings

NON-AGENTIC · same key

Text → vector. Callable as a tool for RAG retrieval / clustering / semantic search.

COSTtext-embedding-3-small: $0.02/1M · text-embedding-3-large: $0.13/1M · ada-002: $0.10/1M
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same key — a paid alternative to the local nomic-embed-text floor when a hosted embedder is wanted. 3-large supports the `dimensions` param to shorten output.

exhaustive specs ▾
modelstext-embedding-3-small (1536 dims) · text-embedding-3-large (3072 dims) · text-embedding-ada-002
specs8191-token max input; input-only pricing
provenancesource ↗

openai · moderation

NON-AGENTIC · same key

Text+image → safety classification. Callable as a pre/post-flight guard (harassment/hate/self-harm/sexual/violence).

COSTFREE
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same key at $0. A cheap safety gate an agent can call before acting on / emitting content.

exhaustive specs ▾
modelsomni-moderation-latest (snapshot 2024-09-26; images ≤20MB)
modalitytext + image → classification scores
provenancesource ↗

Anthropic — Claude

4 wired · 12 total

ROLE IN KIT in-family implementer / conductor / CEILING — exhausted LAST (most expensive), spent on what weaker vendors cannot do. Commit-capable (unlike codex).

KIT CONFIDENCE The ceiling tier and the conductor's own family (its scarcest pool BY CONSTRUCTION — the orchestrator spends it all day). Commit-capable where codex is not. Buffers -p output, so monitor claude workers via GROUND TRUTH (commits/ledger), not log tail. Route delegable side-tasks to OTHER families first (family_parity + quota thrift).

LANE claude -p "<prompt>" --model <id>

claude-opus-4-8

WIRED

The kit's ceiling / critical tier — hardest core rounds, long-horizon wrong-is-expensive work.

COST$5.0in /1M $25.0out /1Mcache write 5m: 6.25 · cache read: 0.5 · batch: 50% off · fast mode: $10/$50 flat (beta)
CONFIDENCE
◆ KIT — measured

Ceiling. Spent LAST, on what the weaker/cheaper vendors cannot do. ~4x less likely than Opus 4.7 to let code flaws pass — a strong auditor as well as builder.

▲ BENCHMARK — external
SWE bench Verified88.6% (3rd-party)SWE bench Pro69.2%GPQA Diamond93.6% (3rd-party)Terminal bench 2.174.6%/82.7% (sources disagree — flagged)
exhaustive specs ▾
context window1000000
max output128000
knowledge cutoff2026-01 (reliable + training)
modalities
  • input text, image, PDF
  • output text
reasoning
  • adaptive_thinking yes, OFF by default
  • effort_tiers low, medium, high (default), xhigh, max
  • extended_budget_tokens not supported (400)
featurestool use, structured outputs, streaming, vision, prompt caching (incl. automatic), Batch API, citations, computer use (beta), MCP connector (beta), Files API (beta), code execution, web search/fetch, task budgets (beta), fast mode (beta, Opus-only), mid-conversation system messages (Opus 4.8 only)
rate limitsShared Opus 4.x bucket. Scale tier: 10,000 RPM / 10M ITPM / 2M OTPM
release2026-05-28
notesIDs are pinned dateless snapshots (no -latest alias in the 4.6+ scheme).

claude-fable-5

WIRED

This workspace's pinned authoring model (Claude 5 family) — always-on adaptive thinking, SOTA-class reasoning.

COST$10.0in /1M $50.0out /1Mcache write 5m: 12.5 · cache read: 1.0 · batch: 50% off
CONFIDENCE
◆ KIT — measured

The packet-workspace ceiling (CLAUDE.md pins Fable 5). Thinking is ALWAYS ON (adaptive); raw CoT never returned. Its own (smaller) rate-limit bucket.

▲ BENCHMARK — external
SWE bench Verified95.0% (3rd-party)SWE bench Pro80.3%Terminal bench 2.188.0%Humanity Last Exam tools64.5%
exhaustive specs ▾
context window1000000
max output128000
knowledge cutoff2026-01
modalities
  • input text, image, PDF
  • output text
reasoning
  • adaptive_thinking ALWAYS ON (disabled -> 400)
  • effort_tiers low, medium, high, xhigh, max
  • cot never returned raw (summarized/omitted only)
featureseffort, task budgets (beta), memory tool, code execution, programmatic tool calling, context editing (beta), compaction (beta), vision, refusal stop-reason + server-side fallbacks
rate limitsOWN bucket (smaller). Scale: 4,000 RPM / 4M ITPM / 800K OTPM
releaseGA 2026-06-09 (redeployed 2026-07-01 after an export-control pause)
provenancesource ↗
notesRequires 30-day retention (a 'Covered Model') — NOT available under zero-data-retention. Sister model Mythos 5 = same weights, classifiers removed (restricted — see below).

claude-sonnet-5

WIRED

The kit's standard / implementer tier — strong agentic coding at a fraction of Opus cost.

COSTintro $2 in / $10 outintro through 2026-08-31: $2 in / $10 out · standard from 2026-09-01: $3 in / $15 out · cache read intro: 0.2 · batch: 50% off
CONFIDENCE
◆ KIT — measured

The default implementer for standard work; adaptive thinking ON by default; first Sonnet with xhigh effort + hi-res vision. NOTE: substantially WEAKER than Opus 4.8 at cyber-exploit dev (0% on a Firefox exploit set) — route authorized-security work to Opus/Mythos, not Sonnet.

▲ BENCHMARK — external
SWE bench Verifiedbreaks 80%SWE bench Pro63.2%OSWorld Verified81.2%Terminal bench 2.180.4%GPQA Diamond96.2% (3rd-party, flagged high)
exhaustive specs ▾
context window1000000
max output128000
knowledge cutoff2026-01
modalities
  • input text, image, PDF
  • output text
  • vision hi-res 2576px long edge
reasoning
  • adaptive_thinking yes, ON by default
  • effort_tiers low, medium, high (default), xhigh, max
featurestool use, structured outputs, prompt caching (incl. automatic), Batch API, citations, computer use (beta), MCP connector (beta), Files API (beta), task budgets (beta), compaction (beta)
rate limitsOwn bucket, same ceilings as Opus 4.x. Scale: 10,000 RPM / 10M ITPM / 2M OTPM
release2026-06-30

claude-haiku-4-5-20251001

WIRED

Light / mechanical tier — cheap, fast; ~90% of Sonnet 4.5 agentic-coding at a fraction of the cost.

COST$1.0in /1M $5.0out /1Mcache write 5m: 1.25 · cache read: 0.1 · batch: 50% off
CONFIDENCE
◆ KIT — measured

The cheap fast tier for mechanical/light work. Pre-4.6 API shape: supports CLASSIC budget_tokens thinking, NO adaptive/effort. Anthropic's 'safest model yet' by internal metrics.

▲ BENCHMARK — external
SWE bench Verified73.3%Terminal bench40.2% (no thinking) / 41.75% (32K budget)
exhaustive specs ▾
aliasclaude-haiku-4-5
context window200000
max output64000
knowledge cutoffreliable 2025-02 / training 2025-07
modalities
  • input text, image, PDF
  • output text
reasoning
  • extended_budget_tokens SUPPORTED (classic API)
  • adaptive_thinking no
  • effort not supported
featurestool use, structured outputs, streaming, vision, prompt caching, Batch API, citations, MCP connector, Files API, interleaved thinking (legacy beta header)
rate limitsGeneral bucket. Scale: 10,000 RPM / 10M ITPM / 2M OTPM
release2025-10-15
notesThe only currently-wired model below 1M context / 128K output.

claude-mythos-5

⚠ RESTRICTED

Same weights as Fable 5 with SAFETY CLASSIFIERS REMOVED — for vetted cyberdefense/biosecurity work.

COST$10.0in /1M $50.0out /1M
CONFIDENCE
◆ KIT — measured

⚠ RESTRICTED — invite-only via Project Glasswing (~150 orgs). NO self-serve key, no Console toggle, no waitlist. The ONLY Claude with cyber safeguards lifted — directly relevant to Wolf's authorized bug-bounty/pentest work IF a Glasswing partnership is approved. Un-wireable today: no code path makes it selectable.

exhaustive specs ▾
context window1000000
max output128000
knowledge cutoff2026-01
release2026-06-09
provenancesource ↗

claude-opus-4-7

⚠ AVAILABLE — NOT WIRED

Previous-gen Opus flagship (one rung below 4.8).

COST$5.0in /1M $25.0out /1M
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. Same price as 4.8, strictly weaker. Only value: a version-pinned reproducibility fallback if 4.8 regresses.

exhaustive specs ▾
context window1000000
max output128000
release2026-04-16 (Active; retire not before 2027-04-16)
provenancesource ↗

claude-sonnet-4-6

⚠ AVAILABLE — NOT WIRED

Previous-gen Sonnet (predecessor to Sonnet 5).

COST$3.0in /1M $15.0out /1M
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. Currently MORE expensive ($3/$15) than Sonnet 5's intro rate and weaker. No reason to add.

exhaustive specs ▾
context window1000000
max output128000
release2026-02-18 (Active; retire not before 2027-02-17)
provenancesource ↗

claude-opus-4-5

⚠ AVAILABLE — NOT WIRED

Last Opus before the 4.6+ adaptive-only generation — supports BOTH budget_tokens AND effort.

COST$5.0in /1M $25.0out /1M
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. Only relevant if a component needs a fixed thinking-token CEILING (budget_tokens) rather than adaptive/effort. 200K ctx.

exhaustive specs ▾
context window200000
max output64000
release2025-11-24 (Active; retire not before 2026-11-24)
provenancesource ↗

claude-sonnet-4-5

⚠ AVAILABLE — NOT WIRED

Pre-4.6 Sonnet — classic budget_tokens thinking, 200K ctx.

COST$3.0in /1M $15.0out /1M
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED, and the EARLIEST retirement floor of any Active model (not before 2026-09-29). Superseded by Sonnet 5 at a lower intro price.

exhaustive specs ▾
context window200000
max output64000
release2025-09-29 (Active; retire not before 2026-09-29)
provenancesource ↗

claude-opus-4-1

⚠ DEPRECATED

Two-gen-old Opus.

COST$15.0in /1M $75.0out /1M
CONFIDENCE
◆ KIT — measured

⚠ DEPRECATED — RETIRES 2026-08-05 (~2 weeks). 3x the price of 4.8 for less capability. If ANY kit config hardcodes `claude-opus-4-1`, that is a LIVE BUG to fix before it 404s.

exhaustive specs ▾
context window200000
max output32000
release2025-08-05 (retires 2026-08-05)
provenancesource ↗

claude-haiku-3-5

⚠ RETIRED

Old cheap tier.

COST
CONFIDENCE
◆ KIT — measured

⚠ RETIRED on the first-party API (2026-02-19). Listed ONLY so nothing references `claude-3-5-haiku-20241022`. (Still on Bedrock/Vertex with their own schedules.)

exhaustive specs ▾
provenancesource ↗

anthropic · (text-only family)

NON-AGENTIC · same key

Anthropic ships NO non-agentic models — no image, video, audio, or embedding models. Claude is a text/vision-in → text family only.

COSTn/a
CONFIDENCE
◆ KIT — measured

For image/audio/video/embedding TOOLS, an agent on the Claude key must reach ANOTHER family (OpenAI/Google/xAI) or a local model. Claude itself accepts image + PDF INPUT (vision) but only emits text.

exhaustive specs ▾
provenancesource ↗

xAI — Grok

1 wired · 10 total

ROLE IN KIT adversarial SECURITY-review lens — the SECOND lens next to agy; also wide cheap parallel recon / N-variant prototyping.

KIT CONFIDENCE HIGH-RECALL / LOW-PRECISION and inconsistent — treat EVERY output as a LEAD, never a final answer. Adjudicate every grok finding against the actual code; NEVER block a ship on an unverified grok flag. It earns its seat by occasionally catching the real one nobody else did (Observify 2026-07-18: 2 rounds ALL false-positives, 1 round 2 REAL findings — a TOCTOU + a tail-path arbitrary-read the builder AND agy both missed, 1 round empty). SuperGrok is the DURABLE review pool — its window outlasts agy's; when agy caps for days, grok survives.

LANE grok -p "<prompt>" -m <id> (OAuth ties to SuperGrok; --effort low|medium|high)

grok-4.5

WIRED

Flagship — adversarial security/recon lens; coding + agentic work; mandatory reasoning.

COST$2.0in /1M $6.0out /1M $0.3 cachedover 200k: $4 in / $12 out
CONFIDENCE
◆ KIT — measured

The kit's grok default. Security lens: lead-only, adjudicate every finding. Reasoning is MANDATORY (cannot disable). Vision + live web/X search.

▲ BENCHMARK — external
SWE bench Pro64.7%Terminal bench 2.183.3%DeepSWE 1.153%AA Intelligence Index54 (rank ~#4)AA Coding Agent Index76hallucinationAA-Omniscience 54% (up from 25% on 4.3 — flagged)
exhaustive specs ▾
context window500000
max outputnot public
knowledge cutoff2026-02-01 (single-source, flagged)
modalities
  • input text, image
  • output text
  • live_search web + X real-time
reasoning
  • is_reasoning_model yes
  • mandatory yes
  • effort_tiers low, medium, high (default)
  • note presence/frequency penalty + stop are rejected
featuresfunction/tool calling, structured outputs, vision, live web/X search, native subagents (Grok Build harness), --worktree per-agent
rate limits150 req/s, 50M TPM (API tier); regions us-east-1/us-west-2 (+EU)
release~2026-07-16 (x.ai/news; 3rd-party says Jul 8 — flagged)
notesAliases grok-4.5-latest, grok-build-latest.

grok-4.3

⚠ AVAILABLE — NOT WIRED

General-purpose flagship-adjacent — cheaper than 4.5, TWICE the context (1M).

COST$1.25in /1M $2.5out /1M $0.2 cachedover 200k: $2.50/$5 · batch: -20%
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. A candidate BIG-CONTEXT + cheaper grok lane (1M ctx vs 4.5's 500K). Also the silent redirect target for all retired grok-4/grok-3 slugs. Benchmarks mostly 3rd-party/low-confidence.

exhaustive specs ▾
context window1000000
modalities
  • input text, image
  • output text
reasoning
  • is_reasoning_model yes
release~2026-04-30 (AA article; no x.ai/news post)
provenancesource ↗

grok-build-0.1

⚠ AVAILABLE — NOT WIRED

Dedicated CODING/agentic-engineering model (successor to grok-code-fast-1) — the cheapest current xAI text model.

COST$1.0in /1M $2.0out /1M $0.2 cachedover 200k: $2/$4
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. A cheap dedicated CODING lane worth piloting against codex/gpt-oss on bounded tickets. Predecessor grok-code-fast-1 scored 70.8% SWE-Bench Verified, 190 TPS (xAI's own claim, old slug).

exhaustive specs ▾
context window256000
modalities
  • input text, image
  • output text
reasoning
  • is_reasoning_model yes
featuresfunction calling, structured outputs
release~2026-05 (day unconfirmed)
provenancesource ↗
notesAliases: grok-code-fast-1, grok-code-fast (old names redirect here).

grok-4.20-multi-agent-0309

⚠ AVAILABLE — NOT WIRED

Parallel multi-agent 'deep research' mode (successor to consumer 'Heavy') — reasoning.effort controls agent count (4 or 16).

COST$1.25in /1M $2.5out /1M $0.2 cachedover 200k: $2.50/$5
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. A native parallel-agent deep-research lane; overlaps what the kit does with its own fan-out, so lower priority — but a candidate for single-call wide recon. Lower rate limits (9 req/s).

exhaustive specs ▾
context window1000000
reasoning
  • is_reasoning_model yes
  • effort_controls_agent_count 4 or 16 agents
rate limits9 req/s, 2.5M TPM (multi-agent fan-out cost)
release~2026-03
provenancesource ↗

grok-4.20-reasoning-and-non-reasoning

⚠ AVAILABLE — NOT WIRED

grok-4.20-0309-reasoning (reasoning-on) and -non-reasoning (lowest hallucination, strict prompt adherence) — 1M context.

COST$1.25in /1M $2.5out /1M $0.2 cached
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. The non-reasoning variant's 'lowest hallucination / strict adherence' profile is interesting for schema-forced extraction; the reasoning variant is a cheaper 1M-ctx alternative to 4.5.

exhaustive specs ▾
context window1000000
reasoning
  • note one variant reasoning-on, one reasoning-off

grok-4-and-grok-3-legacy

⚠ RETIRED

grok-4, grok-4-fast, grok-4.1-fast, grok-4-0709, grok-3, grok-code-fast-1, grok-3-mini.

COST
CONFIDENCE
◆ KIT — measured

⚠ RETIRED 2026-05-15 — these slugs silently redirect to grok-4.3 / grok-build-0.1 and bill at the successor's rate. Listed so nothing hardcodes a dead grok-4/grok-3 id. (grok-3-mini fully retires 2026-08-15.)

exhaustive specs ▾
provenancesource ↗

xai · image generation

NON-AGENTIC · same key

Text/image → image generation.

COSTgrok-imagine-image: $0.02/image · grok-imagine-image-quality: $0.05/image (flat)
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same xAI key. Up to 10 images/request, 1K/2K, many aspect ratios. (grok-imagine-image-pro + grok-2-image retired 2026.)

exhaustive specs ▾
modelsgrok-imagine-image · grok-imagine-image-quality (default 'best')
modalitytext + image → image
provenancesource ↗

xai · video generation

NON-AGENTIC · same key

Text/image/video → video.

COSTgrok-imagine-video: $0.050/sec (≤720p) · grok-imagine-video-1.5: $0.080/sec (adds 1080p)
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same key. 1–15s generation.

exhaustive specs ▾
modelsgrok-imagine-video · grok-imagine-video-1.5
modalitytext / image / video → video
provenancesource ↗

xai · speech (TTS + STT)

NON-AGENTIC · same key

Text→speech and audio→text, via /v1/tts and /v1/stt endpoints.

COSTTTS: $15/1M chars · STT: $0.10/hr (REST) / $0.20/hr (streaming)
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same key. TTS voices (eve/ara/leo/rex/sal + cloning); STT does word-timestamps + diarization, 25 languages.

exhaustive specs ▾
modalityTTS text→audio · STT audio→text
provenancesource ↗

xai · embeddings — NONE

NON-AGENTIC · same key

xAI exposes NO standalone embeddings model (confirmed negative).

COSTn/a
CONFIDENCE
◆ KIT — measured

⚠ For RAG/embeddings, an agent on the xAI key must use ANOTHER family's embedder (OpenAI text-embedding-3, Google gemini-embedding-2) or the local nomic-embed-text. xAI's Collections API embeds internally with no selectable model/price.

exhaustive specs ▾
provenancesource ↗

Local open-weight (Ollama) — free floor

2 wired · 3 total

ROLE IN KIT the free FLOOR — zero quota, spend before any API token. Glue + retrieval; Holt's on-device brain runs here.

KIT CONFIDENCE Free and unlimited, but WEAKER — fills ambiguity with plausible-wrong and emits confident, fluent, WRONG output with NO cap signature to warn you (L92/L95 bite hardest). Use for glue and bounded work; orient it (seven levers) and audit at the SAME build-failing gates. ⚠ Ollama defaults to a SMALL context window — a silently-truncated weak model is worse than a weak model; set num_ctx explicitly. (gpt-oss:20b / :120b live in the OpenAI family above — they are OpenAI open-weights also run here via `codex --oss`.)

LANE ollama run <model> "<prompt>" (needs `ollama serve`; model pulled first)

qwen3:8b

WIRED

Holt's on-device BRAIN — closed-book graduate-exam answering on the Raspberry Pi (the Pi's model ceiling; 14b times out on the Pi). Hybrid thinking.

COSTFREE FREE local. q4_K_M 5.2GB / q8_0 8.9GB / fp16 16GB. Runs on a Raspberry Pi 5 (kit's own claim, not a vendor spec).
CONFIDENCE
◆ KIT — measured

The Pi's ACTIVE brain for Holt closed-book answering. Certifies Holt-as-deployed, so it is the HONEST exam model (using a bigger Mac model would inflate the grade). Instruct/thinking benchmarks NOT public at 8B; base-model figures only.

▲ BENCHMARK — external
note8B thinking-checkpoint benchmarks NOT public. Qwen3-8B-Base: MMLU 76.9, GPQA 44.4, GSM8K 89.8, MATH 60.8
exhaustive specs ▾
params
  • total 8.2B (dense, 36 layers, GQA 32Q/8KV)
  • non_embedding 6.95B
context windowconfig max_position 40960 (Ollama shows ~40K); native 32768 → 131072 via YaRN (Qwen sources disagree — flagged)
modalities
  • input text
  • output text
reasoning
  • is_reasoning_model yes
  • hybrid thinking / non-thinking toggle (/think, /no_think, enable_thinking)
featurestool calling (Qwen-Agent + MCP), license: Apache-2.0
release2025-04-29 (Qwen3)

nomic-embed-text

WIRED

Embedding backbone — text→vector for retrieval/RAG, clustering, classification (NOT a chat model).

COSTFREE FREE local. 274MB — runs on anything incl. a Pi.
CONFIDENCE
◆ KIT — measured

Pulled on the Mac; the retrieval/embedding tier. ⚠ REQUIRES task-instruction prefixes (search_document: / search_query: / clustering: / classification:) or quality degrades. ⚠ Ollama defaults its context to 2K vs the model's native 8192 — raise num_ctx for long-doc RAG.

▲ BENCHMARK — external
MTEB avg62.28 (768-dim); beats OpenAI ada-002 & text-embedding-3-small on MTEB + long-context LoCo
exhaustive specs ▾
params
  • total 137M (nomic-bert-2048 encoder)
context window8192
embedding
  • dimensions 768 (Matryoshka truncatable to 512/256/128/64)
  • max_seq_tokens 8192
modalities
  • input text
  • output dense vector
featuresMatryoshka dims, task-prefix required, license: Apache-2.0
release2024-02-14 (v1.5)

qwen3:1.7b

⚠ AVAILABLE — NOT WIRED

Tiny nano glue tier — hybrid thinking at 1.7B for the cheapest local classification/labeling.

COSTFREE FREE local. q4_K_M 1.4GB / q8_0 2.2GB / fp16 4.1GB — Pi-class.
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED (not confirmed pulled). A candidate nano glue model below gpt-oss:20b for trivial classification/triage on tight-RAM hosts. Base figures only: MMLU 62.6, GPQA 28.3.

exhaustive specs ▾
params
  • total 1.7B (dense, 28 layers, GQA 16Q/8KV)
  • non_embedding 1.4B
context window32768
modalities
  • input text
  • output text
reasoning
  • is_reasoning_model yes
  • hybrid thinking/non-thinking toggle
featurestool calling (Qwen-Agent), license: Apache-2.0
release2025-04-29
notesQwen3 also ships 0.6b / 4b / 14b / 32b dense + 30B-A3B / 235B-A22B MoE — other unused sizes if a bigger local tier is wanted.

Google — Gemini

1 wired · 13 total

ROLE IN KIT the big-window READ-ONLY CORRECTNESS / architecture auditor — whole-monorepo + full external docs in one context.

KIT CONFIDENCE The MOST RELIABLE correctness/architecture reviewer in the roster — clean, line-cited findings (Observify 2026-07-18: caught 4 real unit/logic bugs + a request-path issue codex shipped). READ-ONLY BY CONVENTION — enforce at the LAUNCH POSTURE (`--mode plan`), not the prompt: given write access it EDITED the single-writer file it was reviewing (correct fix, but an L46 breach). The big-window seat. LIGHT subscription — caps FAST (multi-day / ~6.5-day reset); a green 1-token probe does NOT prove run-room. Pin explicit model ids, not `-latest` aliases — a silent repoint mid-audit is a reproducibility risk. ⚠ the kit does not record which Gemini tier agy uses — confirm with `agy models`.

LANE agy -p "<prompt>" --mode plan --model <id> (READ-ONLY posture: --mode plan, NOT --dangerously-skip-permissions)

gemini-3.1-pro-preview

WIRED

Flagship Pro — the correctness/architecture auditor's tier: 1M context, strongest reasoning, best for whole-repo + docs audits.

COST$2.0in /1M $12.0out /1M $0.2 cachedover 200k: $4 in / $18 out · batch: $1/$6 · cache storage: $4.50/1M/hr · search grounding: 5k/mo free then $14/1k
CONFIDENCE
◆ KIT — measured

The likely agy auditor tier (big window + top correctness). ⚠ Still labeled 'Preview' (Gemini 3.5 Pro delayed). Deep Think mode is GATED (AI Ultra / early-access) — do not assume agy reaches it. Confirm the exact tier with `agy models`.

▲ BENCHMARK — external
GPQA Diamond94.3%SWE bench Verified80.6%SWE bench Pro54.2%Terminal bench 2.068.5%ARC AGI 277.1%MMLU92.6%HLE44.4% (51.4% w/ tools)
exhaustive specs ▾
context window1048576
max output65536
knowledge cutoff2025-01
modalities
  • input text, image, audio, video, PDF
  • output text
reasoning
  • is_reasoning_model yes
  • thinking_level minimal, low, medium, high (default)
  • deep_think gated (AI Ultra / early-access only)
  • note keep temperature at default 1.0
featuresfunction/tool calling, structured outputs (incl. with tools), streaming, code execution, Google Search + Maps grounding, URL context, Batch API, context caching, File Search / Files API, flex + priority inference
no supportaudio/image generation, Live API, computer-use tool
rate limitsno static table — AI Studio dashboard, account-specific (tier by spend)
releasePreview 2026-02-19
notesFree tier NOT available for Pro-class (pulled ~Apr 2026). Also `-customtools` variant that de-prioritizes bash/shell in favor of declared tools.

gemini-3.6-flash

⚠ AVAILABLE — NOT WIRED

Current Flash — fast/mid, cheaper than Pro, FRESH knowledge cutoff (Mar 2026), computer-use tool.

COST$1.5in /1M $7.5out /1M $0.15 cachedbatch: $0.75/$3.75
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. A cheaper, FASTER auditor/reviewer tier with the NEWEST cutoff (Mar 2026 vs Pro's Jan 2025) — attractive for high-volume review passes or when agy's Pro quota is scarce. Free tier available (rate-limited).

▲ BENCHMARK — external
SWE bench Pro58.7%OSWorld Verified83.0%Terminal bench 2.178.0%MLE Bench63.9%MRCR 128k91.8%
exhaustive specs ▾
context window1048576
max output65536
knowledge cutoff2026-03
modalities
  • input text, image, audio, video, PDF
  • output text
reasoning
  • is_reasoning_model yes
  • thinking_level minimal, low, medium, high
featuresfunction calling, structured outputs, code execution, Search + Maps grounding, URL context, Batch API, context caching, computer-use (preview)
releaseGA 2026-07-21

gemini-3.5-flash-lite

⚠ AVAILABLE — NOT WIRED

Cheapest current Gemini — fast (~350 tok/s), Mar-2026 cutoff, computer-use built-in.

COST$0.3in /1M $2.5out /1M $0.03 cachedbatch: $0.15/$1.25
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. A very cheap Gemini glue/triage/classification tier ($0.30/$2.50) with a fresh cutoff — a candidate free-tier-eligible reviewer for high volume when the local floor is busy.

▲ BENCHMARK — external
SWE bench Pro54.2%Terminal bench 2.154.0%OSWorld Verified74.0%
exhaustive specs ▾
context window1048576
max output65536
knowledge cutoff2026-03
modalities
  • input text, image, audio, video, PDF
  • output text
reasoning
  • is_reasoning_model yes
  • thinking_level minimal, low, higher
featuresfunction calling, structured outputs, code execution, Search grounding, Batch API, context caching, computer-use tool
releaseGA 2026-07-21

gemini-3.5-flash-cyber

⚠ RESTRICTED

Gemini fine-tuned for VULNERABILITY DISCOVERY & fixing — the CodeMender pilot model.

COST
CONFIDENCE
◆ KIT — measured

⚠ RESTRICTED — governments + trusted partners only (CodeMender pilot); NOT generally available, no public spec sheet. Directly relevant to Wolf's authorized bug-bounty/pentest work IF partner access is granted — un-wireable to `agy` today.

exhaustive specs ▾
release2026-07-21
provenancesource ↗

gemini-3.5-flash

⚠ AVAILABLE — NOT WIRED

Previous Flash (predecessor to 3.6) — still GA.

COST$1.5in /1M $9.0out /1M $0.15 cached
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. Superseded by 3.6 Flash (cheaper output, fresher cutoff); no reason to pick over 3.6. Jan-2025 cutoff.

exhaustive specs ▾
context window1048576
max output65536
knowledge cutoff2025-01
releaseGA 2026-05-19
provenancesource ↗

gemini-2.5-pro

⚠ AVAILABLE — NOT WIRED

Previous-gen Pro flagship — still GA as a fallback tier.

COST$1.25in /1M $10.0out /1M $0.13 cachedover 200k: $2.50/$15
CONFIDENCE
◆ KIT — measured

⚠ AVAILABLE — NOT WIRED. Older-gen fallback if the 3.x Pro quota is exhausted; classic thinking_budget (128-32768, no full disable). Weaker than 3.1 Pro.

exhaustive specs ▾
context window1048576
max output65536
knowledge cutoff2025-01
reasoning
  • thinking_budget 128-32768 or -1 dynamic (cannot fully disable)
benchmarks
  • GPQA_Diamond 86.4%
  • SWE_bench_Verified 59.6% (single) / 67.2% (multi)
  • AIME_2025 88.0%
  • Aider_Polyglot 82.2%
releaseGA 2025-06-17
provenancesource ↗
notesGemini 2.5 Flash ($0.30/$2.50, thinking_budget 0-24576) and 2.5 Flash-Lite ($0.10/$0.40) also still GA — cheap legacy fallbacks.

gemini-3-pro-preview-and-2.0-legacy

⚠ DEPRECATED

gemini-3-pro-preview (Nov 2025), gemini-2.0-flash, gemini-2.0-flash-lite.

COST
CONFIDENCE
◆ KIT — measured

⚠ RETIRED — gemini-3-pro-preview shut down 2026-03-09 (migrate to 3.1 Pro Preview); Gemini 2.0 Flash/Flash-Lite shut down 2026-06-01. Listed so nothing targets a dead Gemini id.

exhaustive specs ▾
provenancesource ↗

google · image generation

NON-AGENTIC · same key

Text/image → image generation & editing (Nano Banana / Imagen).

COSTgemini-3.1-flash-image (Nano Banana 2): $0.045–0.151/img by res · gemini-3-pro-image (Pro, 4K): $0.134–0.24 · flash-lite-image: $0.034/img
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same Gemini key. Nano Banana 2 is strong at layout/text-in-image. (Imagen 4 deprecated — shuts down 2026-08-17.)

exhaustive specs ▾
modelsgemini-3.1-flash-image · gemini-3.1-flash-lite-image · gemini-3-pro-image · gemini-2.5-flash-image (legacy)
modalitytext + image → image (up to 4K)

google · video generation (Veo)

NON-AGENTIC · same key

Text/image → video WITH native audio.

COSTveo-3.1-generate-preview: $0.40/sec (720p/1080p), $0.60 (4K) · veo-3.1-fast: $0.10–0.30/sec · veo-3.1-lite: $0.05–0.08/sec
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same key. Native audio (Veo 3.x). ⚠ Veo 2/3 past their announced shutdown but docs still list them — smoke-test before depending.

exhaustive specs ▾
modelsveo-3.1-generate-preview · veo-3.1-fast · veo-3.1-lite · gemini-omni-flash-preview (conversational video edit)
modalitytext / image → video + audio
provenancesource ↗

google · text-to-speech & music

NON-AGENTIC · same key

Text→speech (30 voices, 90+ langs) and text→music (Lyria).

COSTTTS gemini-3.1-flash-tts: $1/1M in + $20/1M audio out · Lyria 3 Pro: $0.08/song · Lyria 3 Clip (30s): $0.04/song
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same key. TTS single/dual-speaker, expressive tags. Lyria makes full songs from a prompt + lyrics.

exhaustive specs ▾
modelsgemini-3.1-flash-tts-preview · gemini-2.5-{flash,pro}-preview-tts · lyria-3-pro-preview · lyria-3-clip-preview · lyria-realtime-exp
modalitytext → audio (PCM 24kHz) · text → music (mp3/wav)

google · embeddings (multimodal)

NON-AGENTIC · same key

Text/image/audio/video/PDF → unified embedding vector. The first MULTIMODAL embedder — a strong RAG tool.

COSTgemini-embedding-2: text $0.20/1M · image $0.45/1M · audio $6.50/1M · video $12/1M · gemini-embedding-001: $0.15/1M text-only
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same key. gemini-embedding-2 embeds ALL modalities into one space (128–3072 flexible dims, 8192-tok input) — the hosted upgrade over local nomic-embed-text for multimodal retrieval. (text-embedding-004 shut down 2026-01-14; -001 and -2 spaces are incompatible.)

exhaustive specs ▾
modelsgemini-embedding-2 (multimodal) · gemini-embedding-001 (text, 2048-tok)
modalitytext/image/audio/video/PDF → vector
provenancesource ↗

google · Live API / realtime voice

NON-AGENTIC · same key

Bidirectional real-time audio dialogue + live speech-to-speech translation.

COSTgemini-3.1-flash-live: audio $3/1M in ($0.005/min), $12/1M out ($0.018/min) · gemini-3.5-live-translate: ~$0.0053/$0.0315 per min
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same key (WebSocket). A2A voice + 70-lang live translation. Also fills Gemini's STT gap — there is NO standalone Gemini transcription model; for pure STT use Google Cloud Speech-to-Text (a SEPARATE product/key) or send audio to a chat model.

exhaustive specs ▾
modelsgemini-3.1-flash-live-preview · gemini-3.5-live-translate-preview · gemini-2.5-flash-native-audio-preview
modalityaudio+text+image+video → audio+text
provenancesource ↗

google · specialist (computer-use, robotics)

NON-AGENTIC · same key

Purpose-built action models: screen control and embodied/robot reasoning.

COSTgemini-2.5-computer-use-preview: $1.25/$10 per 1M · gemini-robotics-er-1.6-preview: $1/$5 per 1M
CONFIDENCE
◆ KIT — measured

TOOL-CALLABLE with the same key, but lean agentic (they emit ACTIONS): computer-use returns click/type/navigate for browser automation; robotics-ER does spatial/physical task planning. Borderline agentic — flagged.

exhaustive specs ▾
modelsgemini-2.5-computer-use-preview-10-2025 · gemini-robotics-er-1.6-preview
provenancesource ↗
updated just nownext 3m 00s