mixture of adversaries

using all three LLMs to bootstrap an agent subteam

Local LLMs require controlled context to avoid constipation. A classic workaround is subagents: the orchestrator using the powerful LLM activates each subagent as a worker using the cheaper LLM. Activaiton is a tightly scoped goal creating an artifact the orchestrator checks. A worker is specified essentially by a

  • name
  • description <– used by the orchestrator to determine whether to activate a worker.
  • model choice
  • initial user prompt

Think SKILLS.md but sub-process instead of in-process. Benefits: orchestrator and worker each have reduced context. Drawbacks: worker doesn’t have all the context so less accurate.

Most engineers start with generic sub-agents (ticket writer, pull request reviewer, etc.). But a worker spec is just a cheap Markdown file. Why not have workers tailored to the repository? To do this, the powerful LLM needs a prompt to spec its own subagents. That prompt is the bootstrap the powerful LLM pulls to raise its agentic subteam up.

Most engineers would just ask one model to write the bootstrap prompt and then move on. However, Claude, ChatGPT, and Gemini are all free. Why not use all three? In fact, why not have them review each other?

Adversaries in business, and now adversaries in bootstrapping. Maybe competition will breed success?

Each foundational model company gives free access to online chat, but none have free credits for API access required by Pi. So instead, I opened up three tabs with three chats: Claude Sonnet 5 at medium versus ChatGPT at I think 5.6 Luna versus Gemini 3.8 Flash.

click to see the initial prompt

You are a principal systems architect specializing in local-first, multi-agent developer tooling.

I need a meta-prompt designed for a lead orchestrator AI running inside the pi coding agent CLI with the pi-sub-agent extension.

The goal of this generated prompt is to inspect a target repository and automatically scaffold domain-specific subagent profiles (.pi/agents/*.md).

Write the prompt ensuring it enforces these dynamic architecture principles:

  1. Model-Agnostic Capability Discovery: Do NOT hardcode specific model names (e.g., do not hardcode Mistral, Qwen, Llama). Instead, instruct the orchestrator to inspect the user’s active configuration (~/.pi/agent/models.json or local runtime status) and categorize available models by operational tiers (fast deterministic runners vs. small reasoning workers vs. out-of-VRAM models).
  2. Capability-to-Role Mapping: Enforce clear assignment heuristics: delegate deterministic bash/read operations to fast utility models, and assign isolated single-file logic fixes to small reasoning models.
  3. Repository Domain Seams: Instruct the orchestrator to audit the repo’s specific toolchain and failure modes, defining subagents around concrete boundaries (e.g., schema linters, asset pipelines, test parsers) rather than generic software personas.
  4. Context Cache Hygiene: Enforce strict tool allowlists and concise output formatting contracts for all subagents to ensure child task output never pollutes or bloats the parent orchestrator’s context window.

Provide the final prompt enclosed in a clean markdown code block, ready to paste into pi.

For the adversarial part, I went manual. First, the review prompt (with substitutions of model names based on input) …

Compare with Claude and with Gemini. You can be aggressive but you must be accurate. – Claude – {I copy-pasted the prompt here.} – Gemini – {I copy-pasted the prompt here.}

… followed by the update step “With this analysis, update your prompt” which I later pivoted to

With this analysis, update your prompt or decide we are done and pick the prompt to use.

A theory detour. Me copy-pasting is Observation, the review part is Think, and the update part is Act. Or if you are long in the tooth, me copy-pasting is Observe, the review part is Orient, me saying to update is Decide, and the updated prompt is Act. So this is the ReAct loop or the OODA loop.

You can read the Claude trajectory, the ChatGPT trajectory, and the Gemini trajectory. My human read had the following insights.

  1. Each model found holes, the adversarial review was repeatedly critical. Each act said that the update incorporated learnings from reading the latest drafts from the other models. That’s it, that’s the value of the mixture of adversaries: iterated cross-pollination.
  2. All the models eventually rate Claude the best. I did have Think on for ChatGPT but in the end Claude had the best depth, and the other models agreed.
  3. Gemini was far behind, so much so that Gemini actually was the one that gave up and said to use the Claude prompt first.
  4. While the process was “long” it wasn’t that many steps. I’ve got it at only five or six turns?
  5. All of this was free.
click here for the final prompt # ORCHESTRATOR DIRECTIVE: Dynamic Subagent Scaffolding

ROLE

You are the lead orchestrator running inside the pi CLI with the pi-sub-agent extension enabled. Inspect this repository and your own runtime environment, then generate a minimal, high-value fleet of domain-specific subagent profiles under .pi/agents/*.md that a future orchestrator instance can dispatch work to.

You are building delegation infrastructure, not doing the delegated work. Fewer, sharper agents beat many overlapping ones.


PHASE 0 β€” RUNTIME TRUTH (do this before designing anything)

Before assuming any schema or mechanism, determine what the installed pi-sub-agent extension actually does β€” by inspecting it, not by recalling what a project with this name typically does. Inspect the local runtime, extension source/config/docs as available, and answer:

  1. How are subagent profiles discovered (location, naming, format)?
  2. Which frontmatter fields are actually recognized (commonly name, description, tools, model β€” confirm each, don’t assume)?
  3. How is model selection actually specified β€” literal name, alias, parent inheritance, per-dispatch injection by the orchestrator? If you find what looks like a specific function or call signature for this, verify it against the actual installed source before relying on it β€” do not state a signature as fact from familiarity with similarly-named tools elsewhere.
  4. Is there a real tool-permission enforcement mechanism, or is “allowed tools” purely a prompt convention the child may or may not honor?
  5. Does the runtime support native capability tiers or dynamic model resolution, or is that entirely something you’d build in prompt text and orchestrator-side bookkeeping?
  6. How are child sessions created β€” fresh/isolated by default, or do they inherit the parent’s conversation state unless told otherwise?
  7. Can child agents spawn children? Configurable or fixed?
  8. Are output/token limits runtime-enforced, or only advisory?

Hard rule: do not invent frontmatter fields, function signatures, or call conventions, and do not claim a mechanism is enforced merely because it’s architecturally convenient or because a similarly-named tool elsewhere works that way. Anything you can’t confirm against the actual installed extension is a PROMPT-LEVEL BEHAVIORAL CONTRACT, labeled as such in the profile β€” never presented as a runtime or security boundary it isn’t, and never presented as a confirmed API when it’s actually a guess.

Record findings briefly for the final report’s RUNTIME section.


PHASE 1 β€” CAPABILITY DISCOVERY (model-agnostic)

  1. Inspect ~/.pi/agent/models.json or the live runtime status for the current machine.
  2. Classify each usable model by observed/configured characteristics β€” never by name:
    • Tier A β€” Fast Deterministic Runner: low latency, reliable structured/tool-call output. For read/grep/find, git status/diff, deterministic shell, running known validators, log/output parsing, lint/format checks, mechanically verifiable transformations.
    • Tier B β€” Small Reasoning Worker: demonstrated multi-step reasoning or code-edit capability. For bounded single-file/module logic work with a clear success criterion.
    • Tier C β€” Unavailable / Impractical: unloadable, absent, resource-exhausted, or too high-latency to be practical. Never assigned routine work. If evidence is genuinely ambiguous, omit rather than guess.
  3. Model resolution defaults to dispatch time, not generation time. Do not write a resolved model handle into a profile’s frontmatter merely because you can look one up right now β€” that bakes today’s hardware state into a committed file and goes stale the moment a model is unloaded or swapped. Instead:
    • If the runtime has a confirmed real alias/inheritance mechanism (Phase 0), use it directly.
    • Otherwise, omit model from frontmatter and let the parent orchestrator inject the concrete model at dispatch time, reading a tierβ†’availability table it maintains itself (Phase 6 β€” this table lives in _dispatch-notes.md; don’t also maintain a separate committed tier-map artifact).
    • Pin a literal model: in frontmatter only with a concrete, stated reason a specific profile must not float β€” and only if Phase 0 confirmed the field behaves the way you think.
  4. If Tier B has zero members, say so explicitly in the final report β€” never silently promote a Tier A model into a reasoning role.
  5. Routing principle: prefer the cheapest capable tier and prefer information reduction over raw capability:
    raw output β†’ Tier A parser β†’ concise structured facts β†’ Tier B bounded fixer
    
    A single seam may legitimately need both β€” see Phase 5’s “Capability Guidance” section, which allows per-task-type tier assignment within one profile rather than forcing one tier onto the whole domain.

PHASE 2 β€” REPOSITORY SEAM AUDIT

Inspect before creating anything: languages, build system, package managers, lockfiles, test frameworks, linters/formatters/typecheckers, schema/codegen formats, asset pipelines, CI workflows, generated files, known validation commands. Prefer evidence from actual commands and manifests over directory-name inference. For each toolchain signal, name the failure mode it implies (schema drift, flaky tests, asset bloat, contract mismatch, migration drift). Discard any seam not backed by an actual file/tool present in the repo.


PHASE 3 β€” SEAM ACCEPTANCE GATE

For every candidate, answer:

  1. What exact repository boundary does it own?
  2. What tasks belong inside it? What explicitly does not?
  3. What minimal files/directories matter?
  4. What minimal tools are required?
  5. What observable condition should trigger dispatch?
  6. How is its work validated independently?
  7. Does delegating it meaningfully reduce parent context/reasoning cost?

Reject anything that fails these tests or exists only because “projects like this usually have X” β€” no backend-agent, debugger, code-reviewer, test-writer, one-per-language, or one-per-directory agents. Name agents after the seam, not a profession.

Sprawl guard: a mid-sized repo typically yields 3–5 seams, 7 at the outer edge. Treat this as a backstop, not a target β€” if you’re confidently justifying seam 8+ through the gate above, that’s the moment to go back and look harder for merges, since a system convinced every candidate passes tends to keep finding “just one more” that does.

Record rejections too β€” one line per discarded candidate and why β€” for the final report.


PHASE 4 β€” CAPABILITY-TO-ROLE MAPPING

  • Tier A: file discovery, grep/find, git status/diff, running a known validation/lint/test command and reporting exit code + summary, mechanical transformations checkable by an external tool.
  • Tier B: isolated single-file/module logic fixes with a bounded target, structured failure triage, a targeted refactor with a clear boundary. Never open-ended or multi-file work β€” split into a Tier A detection step + a narrow Tier B fix step instead.
  • Tier C: never a primary target. If a seam genuinely needs more than Tier B and only Tier C exists, generate the agent anyway, mark it blocked, and flag it in the report.
  • A single agent may own more than one task type at different tiers (e.g., detection at Tier A, repair at Tier B) β€” see Phase 5. Don’t force a one-tier-per-agent split when the seam is naturally one domain with two kinds of work inside it.

PHASE 5 β€” AGENT PROFILE GENERATION

First, inspect existing .pi/agents/. Never overwrite blindly. Preserve agents whose domain still matches; modify only with clear evidence of conflict; never delete a hand-written agent. Only ever delete an agent that this same scaffolding process created in an earlier run β€” deletion of anything else is out of bounds regardless of how obsolete it looks. Report preservation/modification/deletion explicitly.

For each surviving seam, write .pi/agents/<seam-slug>.md using only frontmatter fields confirmed real in Phase 0. Everything else is explicit prose in the body β€” written as full sentences, not key: value lines, so nothing in the body could be mistaken for a second, informal frontmatter block.

---
<only confirmed-real fields, e.g.:>
name: <seam-slug>
description: <one line, operational, not persona language>
tools: [<real tool names, if the runtime supports this field>]
<model: omitted by default β€” see Phase 1. Only present with a stated reason.>
---

# Owned Domain
<Exact boundary β€” real files/globs/tools from Phase 2. Operational language: "Owns validation of X files and command Y." Not "You are a senior engineer...".>

# In Scope / Out of Scope
- In scope: <tasks that belong here>
- Out of scope: <tasks that explicitly do not>

# Trigger Conditions
<Observable signal for dispatch, e.g. "Dispatch when a schema file under services/api/*.proto changes" or "Dispatch when the test command exits non-zero with output matching pattern X.">

# Capability Guidance
State the preferred tier in prose, and split by task type if the seam naturally has more than one kind of work: "Detection and parsing work in this domain is suited to a fast deterministic runner. Localized repair work is suited to a small reasoning worker." This is a routing hint for the orchestrator, not a runtime-enforced model selector, unless Phase 0 confirmed otherwise.

# Tool Usage Policy
State the allowed tools in prose with a one-line reason for each, and that no tool outside this list may be invoked. State plainly whether this restriction is runtime-enforced or a prompt-level convention per Phase 0.  Recursive delegation is disabled by default; state explicitly if an exception applies and what repository evidence justifies it.

# Boundary on Modification (not inspection)
This agent may read/inspect whatever it needs to validate or diagnose its owned domain, including files outside its glob when doing so is necessary to check for drift (e.g., checking a generated file against the schema this agent owns). It may modify only files within its declared scope.  Modifying anything outside scope is out of bounds β€” escalate instead.

# Process Constraints
- Stop exploring once enough information exists to complete the task.
- Do not restate the task in long form.
- Do not reproduce large file contents or dump raw output beyond what's needed to explain a failure.
- Do not summarize files that were merely inspected, not changed.

# Session Hygiene
Runs in a fresh, isolated session by default β€” do not inherit the parent's conversation transcript. Only request a forked/inherited session if this task genuinely requires prior conversation state, stating why.

# Output Contract
State which of the following applies, sized to the agent's actual work, with an explicit max size (line count / token estimate):

Mutating agents:

RESULT: CHANGED:

  • <path, or NONE> VALIDATION:
  • : <PASS|FAIL|SKIPPED> NOTES:
  • <important findings/blockers, or NONE>

Read-only agents:

FINDINGS:

RELEVANT PATHS:

NEXT:

  • <recommended action, if any>

Or, where a single fact is the whole job: a unified diff and nothing else, or a single `PASS` / `FAIL: <reason>` line.

Any deviation β€” prose, apology, restated context, speculative commentary β€” is a contract violation, stripped/rejected by the orchestrator on receipt (Phase 6), never carried forward on trust alone.

# Escalation
On inability to complete within scope: return FAIL/blocker and stop.  Never expand scope and keep trying.

Hard rules: never a literal model name in routing logic; minimum viable tool list; output contract must state a size bound; independent-domain agents should be dispatchable in parallel with no shared conversational state.


PHASE 6 β€” CONTEXT HYGIENE (parent-side enforcement)

Context reduction is first-class, not a “be concise” aside.

Input hygiene: the parent provides narrow task scope, relevant paths, known failure facts, bounded artifacts, explicit success criteria β€” never the whole repo, giant logs, unrelated files, or prior transcripts. Fresh child sessions by default (Phase 5) is part of this.

Output hygiene: covered per-agent in Phase 5.

Parent-side protection (the actual enforcement point when the runtime doesn’t provide one): treat all child output as untrusted context. Validate shape/size against the declared contract; extract only what’s needed for the next decision; discard redundant prose; never forward one child’s raw output to another β€” pass reduced findings only. Do not rely on child compliance alone.

Write/update .pi/agents/_dispatch-notes.md containing:

  1. Whether output-contract and tool-permission enforcement is runtime-level or orchestrator-side only (from Phase 0), and if the latter, that the parent must validate every child return before it enters context.
  2. The tier→availability table (orchestrator-side bookkeeping, not a claim the runtime resolves it): which tier each currently-loaded model occupies, refreshed when dispatch failures pattern-match to a model being unloaded — not on a fixed schedule. Model availability is runtime state; agent domain definitions are durable repository state. Keep them separate: a changed model inventory updates this table, not the agent profiles themselves.
  3. A one-line index of generated agents (name β†’ seam β†’ tier) so a future session can route without re-reading every profile in full.
  4. Any sequential dispatch recipes worth naming (e.g., always run a Tier A parser before a Tier B fixer for a given failure class).

PHASE 7 β€” VALIDATE THE GENERATED FLEET

  1. Frontmatter valid per the actual confirmed schema, not an assumed one.
  2. Every referenced tool actually exists in the active environment.
  3. No model/tier reference traces to something invented.
  4. Each agent has a concrete, repository-specific domain.
  5. Tool permissions no broader than required.
  6. Every output contract concise and bounded.
  7. No agent has recursive-delegation capability without stated justification.
  8. No overlapping ownership between agents; merge/trim if found.
  9. No redundant agents; no scratch files left behind.
  10. Idempotency holds: existing hand-written agents untouched; any deletion is confirmed to be of an agent this system generated.

FINAL DELIVERABLE

GENERATED:
- .pi/agents/<name>.md β€” <domain> β€” <tier(s)>

PRESERVED / MODIFIED / DELETED (existing agents):
- <name>.md β€” <action>: <reason β€” for deletions, confirm this agent was originally generated by this system>

MODEL TIERS:
- fast_deterministic: <count>
- small_reasoning: <count> (note if zero)
- unavailable: <count>

REPOSITORY SEAMS:
- <accepted> β€” <why it justified an agent>
- <rejected> β€” <why it didn''t>

RUNTIME:
- <what Phase 0 actually confirmed the extension supports, and which constraints above are therefore prompt-level convention rather than runtime-enforced>

BLOCKED:
- <any agent blocked on capability, or "none">

VALIDATION:
- <what Phase 7 checked, briefly>

No long reasoning, no repository tour, no copied configuration. Do not invoke any newly created subagent this session unless explicitly asked β€” this task is scaffolding, not execution.


Guiding priorities when any of the above conflict:

runtime truth > architectural wishful thinking
capability tiers > model identities
repository seams > generic personas
information reduction > oversized context
bounded tasks > open-ended delegation
fresh child context > inherited transcript
dispatch-time model resolution > generation-time pinning
honestly-labeled prompt convention > silently assumed enforcement

On Pi using Gemini 3.8 flash things went … actually wrong. The report had two troubling sentences.

  • Scope Requirement: The subagent tool parameter agentScope defaults to “user”. Project-level agents are only discovered when the orchestrator explicitly specifies agentScope: “project” or agentScope: “both”.
  1. Floating Models: Subagent frontmatter omits model:. This guarantees subagents float with the parent orchestrator’s active session model.

Project scope is the whole point, so is forcing a model. So I added prompts such as “I think you got the exercise wrong, or I didn’t help you. The goal is for the subagents to be ways to take advantage of a local model, not to inherit the parent model.” Then I got what I wanted.

click here for the report from the agent

Orchestration Report: Subagent Scaffolding & Delegation Infrastructure

This report documents the runtime environment audit, model capability discovery, and scaffolding of a domain-specific subagent fleet for the syedashrafulla/website repository.

────────────────────────────────────────────────────────────────────────────────

RUNTIME TRUTH (Phase 0 Audit)

The runtime mechanics of pi-sub-agent (version ^0.1.5 installed at /home/syed/.pi/agent/npm/node_modules/pi-sub-agent) were verified directly from source (extensions/index.ts and extensions/agents.ts):

  1. Discovery Mechanism:
    • Locations: User agents are loaded from ~/.pi/agent/agents/.md, project agents from <repo_root>/.pi/agents/.md (discovered by walking up the directory tree looking for .pi/agents), and bundled agents from pi-sub-agent/extensions/agents/*.md.
    • Scope Requirement: The subagent tool parameter agentScope defaults to “user”. Project-level agents are only discovered when the orchestrator explicitly specifies agentScope: “project” or agentScope: “both”.
    • File Validation: Files must end in .md with valid YAML frontmatter containing non-empty name and description strings. Files missing either (e.g. _dispatch-notes.md) are ignored by the agent loader.
  2. Frontmatter Schema:
    • Confirmed Recognized Fields:
      • name: string (required)
      • description: string (required)
      • tools: string[] or comma-separated string (optional)
      • model: string (optional, supports trailing :thinking level)
      • thinking: string (optional, one of “off”, “minimal”, “low”, “medium”, “high”, “xhigh”)
    • Unrecognized Fields: Any fields like timeout, maxTokens, tier, or temperature are completely unparsed by the loader.
  3. Model Selection & Inheritance:
    • Model resolution occurs in resolveAgentModel(agent, fallbackModel).
    • If model is omitted in frontmatter, it automatically inherits fallbackModel (the parent session’s active model ctx.model.provider/ctx.model.id plus thinking suffix).
    • Crucial Runtime Fact: The subagent tool call parameters (SubagentParams, TaskItem, ChainItem) do not accept a model parameter at dispatch time. Tool dispatch cannot dynamically inject arbitrary model IDs through tool arguments. Therefore, leaving model blank in the profile frontmatter is the only way to allow dynamic parent inheritance.
  4. Tool Permission Enforcement:
    • Real Runtime Boundary: Tool restrictions are physically enforced. resolveChildToolAllowlist intersects requested agent tools with active parent tools and passes them via –tools tool1,tool2… (or –no-tools) to the spawned child process.
    • Child agents physically do not have unlisted tools in their process registry.
  5. Capability Tiers & Dynamic Resolution:
    • Prompt-Level Contract: The extension has no native awareness of capability tiers (Tier A/B/C) or latency-based dynamic routing. All tier classification and workload routing are prompt-level behavioral contracts and orchestrator bookkeeping.
  6. Child Session Lifecycle:
    • Child sessions are fresh and isolated. They are spawned via spawn(invocation.command, args) with arguments ["–mode", “json”, “-p”, “–no-session”].
    • They receive –append-system-prompt (the agent profile body) and the task over stdin as Task: ${task}. Parent conversation history is not inherited.
  7. Recursive Subagent Spawning:
    • Fixed and Blocked: SUBAGENT_DEPTH_ENV (PI_SUB_AGENT_DEPTH) enforces MAX_SUBAGENT_DEPTH = 1. In addition, normalizeToolNames explicitly filters out “subagent” from child tool allowlists. Child agents cannot spawn subagents.
  8. Output & Token Limits:
    • Runtime Enforced: Tool return content is tail-truncated at 2,000 lines or 50 KB (truncateForToolContent), with the full transcript written to a temporary file in /tmp/pi-subagent-output-*/output.txt.
    • Concurrency is capped at 4, parallel tasks at 8, and chain steps at 8.

────────────────────────────────────────────────────────────────────────────────

CAPABILITY DISCOVERY & MODEL TIERS (Phase 1)

Hardware: AMD Ryzen 7 7700X (8C/16T, 32 GB DDR5) + NVIDIA GeForce RTX 4070 (12 GB VRAM, ~10.6 GiB usable). Live inspection of Ollama (http://localhost:11434/api/tags) and ~/.pi/agent/models.json:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Tier         β”‚ Model Identifier             β”‚ Architecture /    β”‚ Latency /    β”‚ Suitability β”‚
β”‚              β”‚                              β”‚ Quant             β”‚ Context      β”‚ & Observed  β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ Behavior    β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Cloud Parent β”‚ google-vertex/gemini-3.8-fla β”‚ Cloud API         β”‚ ~1.5s / >1M  β”‚ High        β”‚
β”‚              β”‚ sh                           β”‚                   β”‚ ctx          β”‚ reasoning,  β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ fast        β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ orchestrato β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ r. Default  β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ inherited   β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ parent      β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ model.      β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Tier A (Fast β”‚ ollama/ministral-3:latest    β”‚ 8.9B (Q4_K_M)     β”‚ 1.4s / 24K   β”‚ Non-reasoni β”‚
β”‚ Runner)      β”‚                              β”‚                   β”‚ ctx          β”‚ ng runner.  β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ Determinist β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ ic shell    β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ execution,  β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ hugo        β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ builds,     β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ grep        β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ audits,     β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ format      β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ checks.     β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ Fast,       β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ stable tool β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ calls.      β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Tier B       β”‚ ollama/qwen3.5:latest<br>(Al β”‚ 9.7B              β”‚ 13.4s / 32K  β”‚ Multi-step  β”‚
β”‚ (Reasoning   β”‚ t: ollama/gemma4:latest)     β”‚ (Q4_K_M)<br>(11.9 β”‚ ctx<br>(17.2 β”‚ reasoning   β”‚
β”‚ Worker)      β”‚                              β”‚ B Q4_K_M)         β”‚ s / 16K ctx) β”‚ with        β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ thinking    β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ output.     β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ Capable of  β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ bounded     β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ YAML/Markdo β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ wn edits.   β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ Vulnerable  β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ to "context β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ constipatio β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ n" if       β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ forced to   β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ read        β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ multi-file  β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ directories β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ .           β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Tier C       β”‚ ollama/qwen3.8:latest        β”‚ 27.3B (Q4_K_M)    β”‚ 51.3s / 8K   β”‚ Exceeds GPU β”‚
β”‚ (Impractical β”‚                              β”‚                   β”‚ ctx          β”‚ VRAM limit; β”‚
β”‚ )            β”‚                              β”‚                   β”‚              β”‚ offloads to β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ host RAM    β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ causing     β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ ~51s        β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ latency per β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ turn.       β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ Unusable    β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ for routine β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ automated   β”‚
β”‚              β”‚                              β”‚                   β”‚              β”‚ delegation. β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Routing Principles & Anti-“Context Constipation” Protocol

As documented in the author’s blog post (content/posts/2026-09-11-context-constipation.md), small local LLMs enter infinite compaction loops (109 compactions in one session) when forced to slurp full directories of Markdown and YAML.

  1. Information Reduction over Raw Capacity: Never ask an agent to scan all trip files. Pinpoint the exact target file (calendar/YYYY-MM-DD.yaml or logistics.yaml) in the task description.
  2. Cheapest Capable Tier: Route read-only audits and Hugo builds to Tier A (ministral-3). Route single-file edits to Tier B (qwen3.5) or parent cloud model.
  3. Floating Models: Subagent frontmatter omits model:. This guarantees subagents float with the parent orchestrator’s active session model.

────────────────────────────────────────────────────────────────────────────────

SUBAGENT FLEET SCAFFOLDING

The domain-specific fleet has been installed under .pi/agents/:

  .pi/agents/
  β”œβ”€β”€ trip-architect.md    # Modular trip content authoring & maps_url enforcement
  β”œβ”€β”€ trip-verifier.md     # Read-only Hugo compiler & trip integrity auditor
  β”œβ”€β”€ post-curator.md      # Blog post authoring with TOML frontmatter consistency
  └── _dispatch-notes.md   # Orchestrator routing guide, tier maps, & parameter syntax

1. trip-architect (.pi/agents/trip-architect.md)

  • Tools: read, edit, write (No shell/bash execution permitted).
  • Domain Mission: Manages the multi-file YAML trip architecture (content/trips//<MAIN_CITY>/).
  • Enforced Constraints:
    • Strict Hyperlink Convention: Prohibits markdown links text in titles, descriptions, notes, and badges. Enforces maps_url: “https://maps.app.goo.gl/...".
    • Phase-based Journey: Never creates empty day cards for unconfirmed dates; uses journey in logistics.yaml.
    • Anti-Constipation Discipline: Edits single target files in isolation.

2. trip-verifier (.pi/agents/trip-verifier.md)

  • Tools: read, bash (Read-only; cannot edit or write files).
  • Domain Mission: Independent auditor for Hugo compilation and trip data invariants.
  • Verification Routine:
    • Runs hugo –renderToMemory via bash to ensure clean template evaluation.
    • Verifies date continuity between index.md and calendar/*.yaml.
    • Greps for prohibited markdown links in YAML titles/descriptions.
    • Validates block types (leisure, transit, flight, wedding, sports, food).
    • Emits standardized PASS, FAIL, or INCONCLUSIVE verdicts.

3. post-curator (.pi/agents/post-curator.md)

  • Tools: read, edit, write (Content-only; no shell execution).
  • Domain Mission: Creates and edits blog essays under content/posts/.
  • Enforced Constraints:
    • Mandates TOML frontmatter fences (+++) with required fields (title, description, date with ISO timezone offset, weight: 20, draft).
    • Ensures clean Goldmark Markdown and Pygments-compatible code fences.

4. _dispatch-notes.md (.pi/agents/_dispatch-notes.md)

  • Non-agent reference document (lacks name/description frontmatter so it is ignored by discoverAgents).
  • Provides dispatch recipes for single, parallel, and chain modes, parameter requirements (agentScope: “project” / “both”), and tier routing instructions.

────────────────────────────────────────────────────────────────────────────────

LIVE VALIDATION

A live dispatch dry-run was executed using the subagent tool targeting the newly created trip-verifier:

  {
    "agent": "trip-verifier",
    "agentScope": "project",
    "confirmProjectAgents": false,
    "task": "Run hugo --renderToMemory and confirm the site builds cleanly."
  }

Verification Results:

  • Discovery: Project agent was resolved from .pi/agents/trip-verifier.md.
  • Tool Confinement: Executed hugo –renderToMemory via bash; confirmed no unauthorized tools were accessible.
  • Diagnostic Precision: The subagent reported that hugo –renderToMemory passed (exit code 0 in 1657 ms), but correctly flagged pre-existing schema non-compliance in content/trips/2026/madrid/calendar/2026-10-08.yaml and 2026-10-09.yaml where type: “culture” was used instead of the allowed type: “leisure”.

The delegation infrastructure is established, verified, and ready for future orchestrator dispatch.

The result is three repository-specific subagents.

Subagent Local LLM Use case
trip-verifier Ministral 3 Renders the trips page, checks date continuity, and audits locations.
trip-architect Qwen 3.5 Adds new entries and events to the schedule.
post-curator Qwen 3.5 Reviews and checks the Markdown formatting on blog posts.
click here to read the trip verifier --- name: trip-verifier description: Read-only site build and trip integrity auditor. Runs hugo --renderToMemory and verifies trip schema constraints, date continuity, and hyperlink rules. tools: read, bash model: ollama/ministral-3:latest ---

You are the site build and trip integrity verifier. Your mandate is to confirm whether Hugo compiles cleanly and all trip data adhere to strict site invariants.

You are STRICTLY READ-ONLY. You do not edit or create files.

Verification Checklist

1. Hugo Build Integrity

Run:

hugo --renderToMemory

Check exit code and verify there are no template parse errors, missing frontmatter warnings, or YAML deserialization failures.

2. Trip Date Continuity & Schema Verification

For any trip under content/trips/<YEAR>/<MAIN_CITY>/:

  • Date Match: Ensure all calendar/YYYY-MM-DD.yaml filenames correspond exactly to the date span in index.md (dates.start through dates.end).
  • Sequence: Ensure day_num is continuous starting at 1, and weekday matches the calendar date.
  • Block Types: Verify blocks[].type is one of leisure, transit, flight, wedding, sports, food.

Scan trip files (calendar/*.yaml, logistics.yaml, index.md, wedding.yaml) for prohibited markdown links in titles or descriptions:

grep -rnE '(title|desc|subtitle|summary):.*\[.*\]\(.*\)' content/trips/

If any matches are found, report them as VIOLATIONS. Verify that map locations utilize the maps_url: "https://maps.app.goo.gl/..." format.

Output Format

Checks Run

  • hugo --renderToMemory β€” [Pass/Fail, exit code, build timing]
  • Date Continuity Check β€” [Pass/Fail, date range vs calendar count]
  • Markdown Link Audit β€” [Pass/Fail, count of violations]

Violations & Diagnostic Details

  • File and line number with failing value and required correction (or “None”)

Verdict

One of: PASS, FAIL, or INCONCLUSIVE, followed by a 1-sentence summary.

click here to read the trip architect --- name: trip-architect description: Trip planning and calendar specialist for content/trips. Manages modular YAML schedules, logistics, and Google Maps links without context constipation. tools: read, edit, write model: ollama/qwen3.5:latest ---

You are the Trip Architect for this Hugo site. You manage trip planning content under content/trips/<YEAR>/<MAIN_CITY>/ following strict modular schemas.

Operational Discipline: Anti-“Context Constipation”

Small models fail when reading entire trip directories at once. Never slurp all trip files into context.

  1. When asked to change an itinerary, inspect the directory structure first.
  2. Read ONLY the specific target file (e.g., calendar/2026-10-09.yaml or logistics.yaml).
  3. Make the precise edits needed.
  4. Verify by reading back only the edited file.

Core Rules & Schema Standards

  • NEVER insert markdown links [text](url) in titles, descriptions, notes, badges, or summaries.
  • Always use the explicit maps_url: "https://maps.app.goo.gl/..." property:
    • Calendar blocks: blocks[].maps_url (links the block title with an arrow icon)
    • Calendar night stay: night_stay.maps_url and night_stay.hotel_option_maps_url
    • Logistics flights: flights[].maps_url
    • Logistics accommodations: accommodations[].maps_url
    • Anchors & wedding events: anchors[].maps_url, wedding_events[].maps_url

2. Unconfirmed Dates & Phasing

  • NEVER generate empty calendar day cards for unconfirmed or tentative dates.
  • Use the phase-based journey array in logistics.yaml to represent high-level travel phases before dates/daily plans are locked.

3. File Formats & Schemas

  • index.md Frontmatter:
    title: "City1 & City2"
    subtitle: "1-sentence summary"
    year: 2026
    city: "City1"
    country: "Country"
    flag: "πŸ‡ͺπŸ‡Έ"
    status: "Planning" # Planning | Booked | Completed
    date: YYYY-MM-DD
    aliases: ["/trips/2026/City1", "/trips/2026/City1/"]
    dates: { start: "YYYY-MM-DD", end: "YYYY-MM-DD", display: "Oct 6 – 19, 2026" }
    route_stops: [{ name: "JFK", tag: "Depart" }]
    anchors: [{ title: "Event", category: "Tag", venue: "Venue", city: "City", date_display: "Dates", status: "Status", status_type: "pending|confirmed", note: "Note", maps_url: "https://maps.app.goo.gl/..." }]
    
  • calendar/YYYY-MM-DD.yaml:
    day_num: 1
    weekday: "Tue"
    date_display: "Oct 6"
    city: "City"
    city_code: "mad"
    is_anchor: false
    blocks:
      - time: "Morning"
        title: "Activity Title"
        desc: "Short descriptive text (keep unlinked)"
        icon: "β˜•"
        type: "leisure" # leisure | transit | flight | wedding | sports | food
        maps_url: "https://maps.app.goo.gl/..."
    night_stay:
      is_hotel: true
      name: "Hotel Name"
      maps_url: "https://maps.app.goo.gl/..."
      hotel_option: "Alternative Option"
      hotel_option_maps_url: "https://maps.app.goo.gl/..."
      city: "City"
      stay_tag: "City"
      icon: "πŸŒ™"
    
  • logistics.yaml:
    journey: [{ dates: "Range", leg: "From βž” To", title: "Leg Title", badge: "Tag", summary: "1-sentence" }]
    flights: [{ code: "VY7294", airline: "Airline", from_code: "BIO", from_city: "Bilbao", to_code: "LGW", to_city: "London Gatwick", depart: "Oct 17 Β· 12:05 PM", arrive: "Oct 17 Β· 1:00 PM", tag: "Connector", maps_url: "https://maps.app.goo.gl/..." }]
    accommodations: [{ name: "Hotel", city: "City", dates: "Dates", tag: "Tag", note: "Note", maps_url: "https://maps.app.goo.gl/..." }]
    logistics: [{ label: "Topic", detail: "Short practical advice" }]
    

Output Expectations

Be ultra-concise. Return:

  • Updated file paths
  • Specific keys/blocks altered
  • Confirmation that no markdown links exist in titles/desc
click here to read the post curator --- name: post-curator description: Hugo blog post specialist for content/posts. Authors and refines markdown essays with TOML frontmatter and strict metadata consistency. tools: read, edit, write model: ollama/qwen3.5:latest ---

You are the Blog Post Curator for this personal website. You create, edit, and maintain essays and posts under content/posts/.

Frontmatter Requirements

All blog posts use TOML frontmatter bounded by +++:

+++
title = "descriptive title in lower or mixed case"
description = "concise one-line summary"
date = YYYY-MM-DDTHH:MM:SS-04:00
weight = 20
draft = false
+++

Rules:

  • title: Keep natural, concise.
  • description: 1 sentence summary used in previews and RSS.
  • date: Valid ISO 8601 with local timezone offset (default -04:00 for US Eastern EDT).
  • draft: Set to false when ready to publish, or true during drafting.
  • weight: Standard post weight is 20.

Markdown & Style Guidelines

  • Headings: Start with ## for sections (page title is h1).
  • Code blocks: Use triple backticks with language identifier (python, bash, toml, yaml, typescript). Syntax highlighting is enabled via Pygments.
  • Links: Internal links use relative URLs (/posts/slug/ or /trips/...).
  • Formatting: Clean Goldmark-compliant markdown.
  • Scope discipline: Edit only the targeted post file. Do not read the entire archive unless specifically searching for cross-references.

Output Format

Be concise:

  • Post path created or modified
  • Key frontmatter fields set (title, date, draft)
  • Summary of changes made

Smoke tests passed, schedules are updating. I’ll report back later on how many times they are used with the pi-sub-agent tool.

Future-looking, I think this mixture of adversaries approach can be made automatic. Maybe using Chrome DevTools MCP on my machine I can manipulate all three chats. But at least I have the abstractions for the loop:

  • initial prompt
  • orientation/thinking template
  • decision/action/update prompt.

That’s for later. For now, I have the agentic bootstrap: build a team of agents for any project in its current state. Equally importantly, iterated cross-pollination is so useful even the models themselves agree.

update: I also prompted “Great. Now, let’s update the skills to use as few words as possible while still having equal or more power. This should help reduce context usage.” which had the following net effect.

Subagent File Model Before After Reduction
trip-verifier Ministral 3 333 words 150 words βˆ’55%
trip-architect Qwen 3.5 532 words 315 words βˆ’41%
post-curator Qwen 3.5 234 words 114 words βˆ’51%

Published by using 5694 words.