Concilium — cross-model adversarial review
A second, different model reviews your (or the user's) claims adversarially. Different model lineage means different blind spots — that's the value. The reviewer PROPOSES; the calling session RATIFIES. Never let either side's confidence substitute for evidence.
Designed to be orchestrated from Claude Code — Opus 5 and Fable 5 are both first-class ratification seats (measured at chair parity on a blind outcome-prediction benchmark; any Claude model can drive the loop, but the ratifier should be one of the two). The GPT side (sol/terra/5.5 via codex) does the independent probing and mechanical execution — and that cross-family seat is load-bearing: it measurably catches what same-family chairs jointly miss.
Prerequisites (check once per environment)
-
codex login status→ must say "Logged in using ChatGPT" (subscription OAuth — an API key is NOT needed and a subscription can NOT be used as one; don't attempt proxy/router bridges). -
Discover available models:
codex debug modelsor~/.codex/models_cache.json. If a model errors "requires a newer version of Codex", runcodex updateand retry. -
First time in a new environment, run the calibration bootstrap (references/setup.md) before trusting verdicts: a known-truth reasoning test, then one simple real task, then (optionally) a head-to-head to pick tier models.
-
Kimi seat — EXPERIMENTAL and opt-in; never part of a default round. Add it only when the user explicitly asks for a third family. Two transports, each with its own wrapper here:
- Kimi Desktop (Windows).
scripts/concilium-review-kimi.ps1drives the app's bundled daimon runner under its own Electron. Requires the desktop app installed and signed in. - Kimi Code CLI (Linux/macOS/Windows).
scripts/concilium-review-kimi.shdrives the cross-platform, MIT-licensed CLI — headlesskimi -pplus device-code OAuth, so it suits a container or a throwaway VM (worked example: references/isolated-guest-vmware.md). ⚠ Print mode auto-approves every tool call by construction, so isolation there is a precondition, not a precaution. ⚠ Always pass--model: the CLI's built-in default is an older generation than the flagship and nothing in the output names it.
Calibrated but flakier than codex, and weaker on isolation: paths, models, sandboxing and the caveats are in references/setup.md and pitfalls #18–21.
- Kimi Desktop (Windows).
Tier matrix (defaults are current-day models — override per installation)
| Tier | Default | Effort | Use for |
|---|---|---|---|
| Research | flagship (e.g. gpt-5.6-sol) | max | open review rounds, adversarial verification |
| Mechanical | prev flagship (e.g. gpt-5.5) | medium | verify a known claim with one probe |
| Runner | cheap tier (e.g. gpt-5.6-terra) | low | execute-and-report: run a script, babysit an import |
Research-tier wrappers default to max on every seat. Effort vocabularies differ and are worth
knowing exactly: codex accepts none · minimal · low · medium · high · xhigh · max (measured —
the API rejects anything else and names the enum), the kimi seat accepts low · high · max.
Effort is not a substitute for a second family — measured, with the control. Six runs of one
seat (gpt-5.6-sol) over the same 14-item prediction packet, at low/medium/high/xhigh/max plus a
same-effort replicate: two runs at the SAME effort covered 10/14, while cross-effort pairs
averaged 8.9 and never beat that; the full six-run ensemble also reached only 10/14, equal to the
two-run noise floor and equal to one codex+kimi pair. Accuracy was not monotonic in effort
(high 9, max 9, xhigh 7, medium 7, low 7) while wall time grew 8× (21s → 166s). Four items were
missed by every run at every effort — and the cross-family seat got two of them. Varying effort
resamples the same blind spots; a different lineage is what moves them. If you want coverage,
add a family, not compute. -Mechanical is the deliberate opt-down and stays at medium.
Runner tasks are NOT reviews — skip the wrapper and call codex directly:
codex exec -m <cheap-model> -c model_reasoning_effort=low [-s read-only unless it writes] "<task>" < /dev/null
Always close stdin on direct non-interactive calls (bash < /dev/null; PowerShell $null | codex …) —
an open non-TTY stdin blocks codex forever on "Reading additional input from stdin...", and the orphaned
process survives the parent shell's timeout (pitfall #10; the wrappers are immune — they pipe via stdin).
Running a review
Use the bundled wrappers. They load the shared review contract from
references/contract.md (single source of truth — falsification probe, alternative explanation,
caveat, verdict-proposal, schema/encoding rules; edit the contract there, never in the
scripts) and add provenance stamping. Both wrappers are functionally identical; pick by platform:
Linux / macOS (bash):
- Claim:
scripts/concilium-review.sh claim "<claim>" - Diff:
scripts/concilium-review.sh diff [base-branch] - Config via env:
MODEL,EFFORT,MECHANICAL=1(mechanical tier),REPO_DIR,PROJECT_RULES(rules file path). - First use after clone:
chmod +x scripts/concilium-review.sh.
Windows (PowerShell 5.1+):
- Claim:
powershell -ExecutionPolicy Bypass -File scripts/concilium-review.ps1 -Claim "<claim>" [-Mechanical] [-RepoDir <path>] [-ProjectRules <file>] - Diff:
... -Diff [-Base <branch>]— reviews the working-tree diff of-RepoDir.
Kimi third-family seat — EXPERIMENTAL, opt-in. Not part of a default round; add it only when
the user asks for a third family. Two wrappers, same contract and same five blocks:
concilium-review-kimi.sh (cross-platform CLI, env-configured — MODEL is mandatory there) and
concilium-review-kimi.ps1 (Windows desktop runner).
The desktop wrapper takes the same -Claim/-Diff/-RepoDir/-ProjectRules/-PriorRounds surface, plus
-Model (default k3-agent), -Effort (low|high|max) and -RawPrompt (no contract, for
calibration probes). Two seat-specific differences, both measured: CLAUDE.md auto-bridging is
OFF by default here — the injection reliably kills the run with a bare Connection error., so
it is opt-in via -AutoRules; use -ProjectRules <file> to give the reviewer context instead.
And never trust its exit code — it exits 0 on that failure, so check for the five blocks.
Multi-line prompts must go through this wrapper; a direct kimi-daimon --prompt from bash gets
its argv mangled (pitfalls #18–19).
This seat has no sandbox — give it a disposable copy. -SandboxFrom <dir> copies the tree
to a throwaway directory (excluding .git, virtualenvs and secret-shaped files), runs the agent
there, prints exactly what it created/modified/deleted, and deletes the copy unless
-KeepSandbox. Use it by default. Be clear about what it buys: workDir is not a boundary —
a canary outside it was read by absolute path and returned verbatim (pitfalls #20) — so this is
blast-radius control and an audit trail, not containment. If material on the machine must not
reach the provider, isolate at the OS level (separate account with ACLs, or a VM).
For a blind round, the risk from that same porousness is contamination rather than damage:
an agent that wanders into the real tree can find the answer, and the round silently stops being
blind. Add -WatchPaths <a,b,c> — it snapshots NTFS last-access times for whatever must stay
unread (the real repo, the results log, the answer key) and reports anything read during the
run. Validated against a known escape in both directions; read pitfalls #21 before trusting it,
including why an enumeration-based version of the same check reported "clean" on a run that had
demonstrably escaped.
Operational rules (each one is a measured failure — the why is in references/pitfalls.md):
- Run in background with a full ~10 min timeout from the FIRST call. Real reviews take 5–15+ min at high effort; a foreground timeout kills them mid-probe.
- Prefer a fresh session over resuming a timed-out one. Long resumed chains hit context compaction — the reviewer's early careful reading gets lossy-summarized before the final, consequential step.
- Never bare-resume.
codex exec resumesilently resets model AND sandbox to the user's config.toml defaults. If you must resume (or want to switch models mid-session), re-pin everything:codex exec resume -m <model> -c sandbox_mode="read-only" -c model_reasoning_effort=<tier> <session-id> -Flags go BEFORE the positional session id. The key issandbox_mode—-c sandbox=...is silently ignored, and there is no-sflag on resume. Cross-model resume retains context. - The reviewer is a full agent, not a chatbot — read-only sandbox blocks file writes, not read commands or DB SELECTs. Everything it reviews goes to the second model's provider.
- Watch progress live, don't wait blind — and monitor the right stream. The contract (rule 9)
makes the reviewer emit
STATUS:one-liners as it works, and codex writes progressively — but the streams split (verified live): with1> out 2> err, the final five blocks land on stdout while the streaming transcript (banner, STATUS lines, tool calls) goes to stderr. Point a tail/monitor at stderr for progress + failure signatures; read stdout for the verdict. Caveat: PowerShell1>/2>redirects write UTF-16 — decode accordingly (or redirect through a UTF-8-forcing step) before grepping. - A blind round needs structural isolation, not an instruction. When the round must be
unprimed (a blind eval, a framing-critical blind-first pass per request-template), run the
reviewer in a clean directory with auto-rules bridging OFF (
-NoAutoRules/NO_AUTO_RULES=1): a model carrying project context and told to "answer from the packet alone" measurably still uses that context (pitfalls #16–17).
Ratification protocol (the calling session's job)
The reviewer returns five blocks: PROBE / ALT / CAVEAT / VERDICT-PROPOSAL / PHASE-LOG.
Before relaying or acting:
- Read the actual probe (the query/commands), not just the prose summary.
- Extremal results are a tripwire: 0% or 100% on a first attempt usually means a wrong join key, wrong scope, or wrong table — not a discovery. Verify the probe's load-bearing step yourself before accepting it.
- Scope-check disagreements: two probes can both be factually right at different scopes (one table vs DB-wide, one source vs all sources). Name the scope before comparing numbers.
- Distinguish refuted / stale / incomplete. "The numbers differ today" does not mean the claim was wrong when written — check history/timestamps before saying "refuted".
- Assign the final verdict tag yourself:
[V-code](verified vs source, cite file:line) /[V-db](read-only query, cite it) /[V-probe](re-runnable script) /[C](unverified) /[X](refuted — name what supersedes it). The proposal is input, not the answer. - Weigh agreement by lineage. Same-family confirmation (a Claude chair agreeing with a Claude orchestrator) is weak evidence — same-lineage chairs measurably share wrong answers, down to independently producing the identical wrong inference. A cross-family confirmation or refutation outweighs any count of same-lineage votes; never settle a dispute by majority across chairs that share a lineage. A third family (Moonshot) is available as an experimental opt-in seat; a third seat buys nothing unless it is independent, so weigh by family, not by headcount — and note that a unanimous panel may simply mean the item was easy (setup.md).
The concilium loop (iterative rounds)
A single review pass is often enough. But when the reviewer's probe has a gap, or you (the orchestrator) disagree with the proposal on defensible grounds, one exchange isn't a concilium — a council deliberates. The loop runs review rounds until the verdict converges or the dispute is proven genuine. This loop is orchestrated by you, the calling Claude session — it is a protocol, not a script (the ratification step is your judgment; nothing can automate it).
Each round:
- Run a review (the wrapper) → get the five blocks → ratify per the protocol above.
- Decide the round's outcome and act:
| Outcome | Condition | Action |
|---|---|---|
| Converged | You verified the probe's load-bearing step and it holds | STOP — emit the final tag. |
| Dispute | The probe has a gap, wrong scope, or you have a specific, evidence-backed objection | Write this round's PROBE + your objection to a rounds file; run the next round with -PriorRounds/PRIOR_ROUNDS pointing at it. |
| Dry | A round adds no new checkable evidence — the reviewer re-asserts, or says (in CAVEAT) it has no new path | STOP — escalate to the owner as [C]/[POLICY] with the open question. This is the anti-oscillation guard. |
| Cap | Round limit reached (default 3) without converging | STOP — present the state and escalate; a real dispute is a finding, not a failure. |
Design rules (they follow directly from the pitfalls):
- Fresh session per round — never a resume chain. The loop is exactly the "long chain"
that pitfall #3 warns about; carry context forward via the
-PriorRoundsfile, notcodex exec resume. Each round starts clean and sees only a compact summary of what was already tried. - Every round must add a NEW evidence path. The contract (rule 8) enforces this on the reviewer side; you enforce it on yours — an objection is only worth a round if it's backed by evidence or points at a concrete, checkable gap. "I'm not convinced" is not a round.
- Ratifier stays fixed (you / Fable); the reviewer can drop tiers as the dispute narrows.
Round 1 on the research tier; once it's down to a mechanical check, run later rounds
-Mechanical. Each round is a real 5–15 min codex call — the cap and the dry-stop are cost controls, not just correctness ones. - Keep the rounds file in durable project storage (not a session temp dir), so the whole deliberation is auditable and the final PHASE-LOG can cite it.
Trigger it when the user asks to "loop", "iterate", "keep going until it's resolved", "have them hash it out", or when a first pass comes back disputed and the stakes justify another round.
Project adaptation
The reviewer sees AGENTS.md, not CLAUDE.md — mind the gap
codex auto-loads AGENTS.md (from the working directory upward), the same way Claude Code
auto-loads CLAUDE.md. They are different files: a project with only a CLAUDE.md gives the
reviewer none of the ground rules Claude has — it reviews half-blind. Three ways to close it,
in order of durability:
- Best (project-level): make
AGENTS.mdexist. Mirror yourCLAUDE.mdinto anAGENTS.md(or makeAGENTS.mda short pointer to it), and keep them synced. This helps all codex usage, not just this skill, and is codex's own supported convention. - Automatic (built into the codex wrappers): CLAUDE.md bridging. When no
AGENTS.mdis present, the wrapper auto-injects the project'sCLAUDE.md(root or.claude/CLAUDE.md) into the contract and prints a notice, so the reviewer isn't missing rules. Disable with-NoAutoRules/NO_AUTO_RULES=1(e.g. a huge, mostly-workflow CLAUDE.md you don't want in every review). The kimi wrapper inverts this — bridging is off unless you pass-AutoRules, because the injection breaks that seat (pitfalls #18); use option 3 there. - Curated (explicit):
-ProjectRules <file>. Point at a short, hand-picked extract of the safety-critical rules — this overrides auto-bridging. Best for large instruction files where only a slice is relevant to review (invariants, "never touch X", schema quirks).
If you keep both files but let them drift, the reviewer sees the AGENTS.md version — sync them.
Other adaptation
- If the project keeps a claims ledger, the PHASE-LOG block is a ready-to-paste line
(
Phase N — <reviewer>(<model>) — <date> — <found> [proposed]); append it only via the project's own hygiene rules (typically: owner or main session, append-only). No ledger → drop the block. - Storage: keep probe outputs and frozen samples in a durable project location, never in session-scoped temp dirs (they die with the session).
References
references/request-template.md— how to construct the REQUEST you hand in (your side, not the reviewer's): confidence-tag facts (never "do not re-derive" over a conclusion), always mount the repo/DB, license rejecting the frame, and run a blind-first pass for framing-critical rounds. Read before writing any non-trivial request.references/contract.md— the review contract the wrappers send (edit it there; both scripts load it at runtime).references/pitfalls.md— known issues and the rules that counter them (read when a rule seems overcautious, or when debugging reviewer misbehavior).references/setup.md— first-time setup, calibration bootstrap, and the head-to-head method for picking tier models.