AI Agent Orchestration 101A Deep Dive Covering Local LLMs
Eight Platforms, Eight Tasks, Eight Agent Teams, What Each One Costs, and How to Remove that Cost using Local LLMs
Section I. Introduction to Agent Orchestration and the Power of Local LLMs
There has been a statistic doing the rounds recently.
Job listings for AI orchestrators rose 1,294% in a single year, from February 2025 to February 2026, according to Lightcast data reported by IT Brew.
You could have thought, “Ah, running these agents costs a lot of money, especially if you are in Asia or Africa!”.
I am here to give you a loophole that removes that cost entirely –
Local LLMs that fit in 64 GB of dedicated VRAM or even Unified Memory can perform orchestration for you!
That removes the cost problem – by effectively taking AI inference costs to zero.
Then you might have thought, “But AI Agent Orchestration must be a deep skill requiring an extremely in-depth knowledge of coding! After all, they pay so much!”.
Again – not so much.
The two skills you need across AI Agent Orchestration?
And the real depth – practice and experience.
You can pick up AI Agent Orchestration in a month with dedicated practice.
And costs can be zero for AI (apart from the hardware) with Local LLMs.
Don’t believe me?
Read till the end!
Do not run multiple agents on Claude everyday – unless you can pay $1,000+ per month for the API or a $200 Max subscription every month.
Obviously, this does not apply if you are running your own local LLM.
That is how you get over the cost problem.
You can definitely orchestrate multiple agents with local LLMs!
Local LLMs like Gemma 4 or Qwen 3.8 are the best option for 20 USD AI subscribers.
Anthropic measured the cost of agents carefully in its write-up on how it built its multi-agent research system, and the numbers are sobering.
- A single agent uses roughly four times the tokens of an ordinary chat conversation.
- A multi-agent system uses roughly fifteen times the tokens of an ordinary chat conversation.
- Token usage alone explained about 80% of the performance difference in Anthropic’s browsing evaluation.
More agents means more tokens, and more tokens means more money.
So the honest starting point of AI agent orchestration is not spinning up ten agents at once—it is setting a budget and deciding what each agent is worth to you.
However, with local LLMs, that problem disappears.
Read to the end to find out everything!
What Orchestration Actually Means
Orchestration is simple to describe, even though it takes practice to do well. In every setup in this article, the roles break down the same way.
The developers who learn to conduct agents, rather than simply chat with them, are quietly becoming 10X engineers. They are not typing any faster than before. They are delegating far better than before.
The Entry Tickets In October 2026
These are list prices taken from official pricing pages and recent reporting. Regional pricing, including pricing in India, can differ from the figures below.
| Provider | Entry Plan | Heavy-Use Plan | Pay-As-You-Go |
|---|---|---|---|
| Anthropic (Claude Code) | Pro, $20/month | Max, from $100/month | Sonnet 5.5: $2 / $10 per M tokens |
| OpenAI (Codex) | Go $8, Plus $20 | Pro: $100, $200, $500 | Credits and API keys |
| xAI (Grok Build) | SuperGrok or X Premium+ | Higher SuperGrok tiers | xAI API |
| Google (Antigravity CLI) | Free tier with AI credits | Enterprise via Google Cloud | Paid API keys |
| Microsoft (GitHub Copilot) | Pro $10 ($15 in credits) | Pro+ $39, Max $100 | $0.01 per extra credit |
| OpenCode | Free client, some free models | — | Zen per token, or your own key |
| Pi | Free client | — | Your own key |
| DeepSeek | No subscription | — | V4-Pro: $0.66 / $1.98 per M (off-peak) |
How To Read The Bills In This Article
Every provider section ends with an estimated bill for its task, and you should read those bills with three caveats in mind.
- Each bill uses list API prices and token counts that I assumed for a mid-sized repository, not measured usage.
- Prompt caching usually reduces the real figure, sometimes quite sharply, because repeated context is billed at a discount.
- On a subscription such as Pro, Max, Plus, or SuperGrok, the same work draws down your plan allowance instead of charging your card.
My confidence in the list prices is high, because they come from official pages as of early October 2026. My confidence in the token counts is moderate, because they are working assumptions rather than measurements.
Section II. How To Stay Under Budget
The single most important idea in this entire article is that context is the meter. Every turn re-sends the conversation, every tool call adds its output to the pile, and every new agent starts a pile of its own.
Anthropic’s guide to managing Claude Code costs puts real numbers on this, and they are worth memorising.
- Average enterprise spend is about $13 per developer per active day.
- Monthly spend typically runs between $150 and $250 per developer.
- Ninety percent of users stay below $30 per active day.
- Agent teams use roughly seven times the tokens of a standard session when teammates run in plan mode.
Seven Rules That Keep You Under Budget
/clear in Claude Code costs nothing, whereas /compact is itself a large request because the model must read everything it summarises.CLAUDE.md under 200 lines and moving specialist workflows into skills, which load only when they are needed.gh, gcloud, and aws cost almost nothing in context, while every idle MCP server adds a small cost to every request.--max-budget-usd flag for scripted runs, Copilot lets you keep overages switched off, and DeepSeek charges half price during off-peak hours.The Cheapest Hook You Will Ever Write
The hook below filters test output so that only failures reach the model. It is adapted from Anthropic’s own documentation, and on a large test suite it can save thousands of tokens on every single run.
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{ "type": "command", "command": "~/.claude/hooks/filter-test-output.sh" }
]
}
]
}
}
#!/bin/bash
# filter-test-output.sh — runs BEFORE every Bash call Claude Code makes.
# If the command is a test runner, rewrite it so only failures come back.
input=$(cat) # hook payload arrives as JSON on stdin
cmd=$(echo "$input" | jq -r '.tool_input.command') # the shell command Claude wants to run
if [[ "$cmd" =~ ^(npm test|pytest|go test|cargo test) ]]; then
# Keep only FAIL/ERROR lines plus 5 lines of context, capped at 100 lines
filtered_cmd="$cmd 2>&1 | grep -A 5 -E '(FAIL|ERROR|error:)' | head -100"
# Allow the call, but with the filtered command
echo "$input" | jq --arg filtered "$filtered_cmd" \
'{hookSpecificOutput: {hookEventName: "PreToolUse", permissionDecision: "allow", updatedInput: (.tool_input + {command: $filtered})}}'
else
echo "{}" # any other command passes through untouched
fi
Section III. Claude Code
In Brief
- Claude Code is Anthropic’s agentic coding tool, and it runs in the terminal, in VS Code and JetBrains, in a desktop app, and on the web.
- Every surface shares the same engine, the same
CLAUDE.mdfiles, and the same MCP servers, so your setup travels with you. - It offers the deepest orchestration toolkit available today, including subagents, agent teams, hooks, skills, plugins, and dynamic workflows.
- It is included with Pro at $20 a month and with Max from $100 a month, according to Claude’s official pricing page.
The Perfect Task: Build A Full-Stack Feature End To End
Claude Code is the best tool for tightly managed feature work, because each subagent gets its own model, its own tools, and its own context window. That lets one lead keep the design coherent while specialists do the heavy lifting in parallel. The task here is to add an “export invoices as CSV” feature to a web application, covering the API, the user interface, the tests, and a final review.
The Agent Team
| Agent | Model | Job |
|---|---|---|
| Lead (your session) | Opus 5.5 | Plans, assigns, and integrates |
backend-dev |
Sonnet 5.5 | Builds the API endpoint |
frontend-dev |
Sonnet 5.5 | Builds the export button and flow |
test-runner |
Haiku 5.5 | Runs tests and reports only failures |
reviewer |
Sonnet 5.5 | Reviews the final diff in read-only mode |
Step-By-Step
# macOS, Linux, or WSL — the native installer auto-updates
curl -fsSL https://claude.ai/install.sh | bash
claude --version # a version number means the install worked
cd your-project
mkdir -p .claude/agents # every .md file in here becomes a subagent
backend-dev, which may only touch server-side files.cat > .claude/agents/backend-dev.md <<'EOF'
---
name: backend-dev
description: Implements backend API changes only. Use for server-side tasks.
model: sonnet # mid-tier model for implementation work
tools: Read, Edit, Write, Bash
---
You implement API endpoints. Touch only files under src/api/ and src/services/.
Return a short summary of the files you changed and why.
EOF
frontend-dev, which may only touch client-side files.cat > .claude/agents/frontend-dev.md <<'EOF'
---
name: frontend-dev
description: Implements UI changes only. Use for client-side tasks.
model: sonnet
tools: Read, Edit, Write, Bash
---
You build UI components. Touch only files under src/web/.
Return a short summary of the files you changed and why.
EOF
test-runner, on the smallest model, because reading logs needs no deep reasoning.cat > .claude/agents/test-runner.md <<'EOF'
---
name: test-runner
description: Runs the test suite and reports ONLY failures. Use after any change.
model: haiku # small, fast and cheap — ideal for reading logs
tools: Bash, Read # can run and read, but never edit
---
Run the tests. Return the test name, file:line, and a one-line probable cause. Never paste full logs.
EOF
reviewer, which can read and run git diff but cannot edit anything.cat > .claude/agents/reviewer.md <<'EOF'
---
name: reviewer
description: Reviews diffs for bugs, security issues and missing tests. Read-only.
model: sonnet
tools: Read, Bash # Bash is only for git diff; there are no editing tools
---
Review the current branch against main. Return a numbered list of issues, most severe first.
EOF
claude --model opus
Shift+Tab to enter plan mode, ask for a plan for the CSV export feature, and approve the plan before any code is written.Use backend-dev and frontend-dev in parallel to implement the approved plan.
When both finish, use test-runner. Fix any failures yourself.
Finally, use reviewer and address every issue it raises.
/usage when the work is done, because it attributes your spend to each subagent and shows you where the tokens went.Estimated Bill
| Agent | Tokens In / Out | Cost At List Price |
|---|---|---|
| Lead (Opus 5.5, $4 / $20) | 300K / 30K | $1.80 |
| backend-dev (Sonnet 5.5, $2 / $10) | 400K / 40K | $1.20 |
| frontend-dev (Sonnet 5.5) | 400K / 40K | $1.20 |
| test-runner (Haiku 5.5, $1 / $5) | 300K / 10K | $0.35 |
| reviewer (Sonnet 5.5) | 200K / 15K | $0.55 |
| Total | ≈ $5.10 |
On a Pro or Max plan, this work comes out of your plan allowance rather than your card. On the API, prompt caching usually brings the real figure well below five dollars.
Section IV. OpenAI Codex
In Brief
- Codex is OpenAI’s coding agent, and it is bundled into ChatGPT plans rather than sold as a separate product.
- It runs as a command-line tool, an IDE extension, a web app, and a cloud sandbox that returns results to your repository.
- The
codex execcommand runs Codex non-interactively, which makes it ideal for scripts and continuous integration. - The default model is GPT-6.1 Sol at medium reasoning, according to OpenAI’s Codex CLI documentation.
- Codex paired with GPT-5.5 holds the top Terminal-Bench 2.0 entry, a score of 82 in April 2026, on the Codesota Terminal-Bench 2.0 leaderboard.
The Perfect Task: A Parallel Test-Coverage Sweep
Codex is the best tool for headless, scriptable fan-out, because codex exec turns each agent into a single shell command. When you combine it with git worktree, every agent gets its own isolated copy of the repository and the agents can never overwrite each other. The task here is to write missing unit tests for four packages at the same time, and then to review all four sets of tests.
The Agent Team
| Agent | Where It Runs | Job |
|---|---|---|
tests-auth |
Worktree 1 | Tests for packages/auth |
tests-billing |
Worktree 2 | Tests for packages/billing |
tests-search |
Worktree 3 | Tests for packages/search |
tests-notify |
Worktree 4 | Tests for packages/notify |
reviewer |
Main repository | Audits all four branches |
Step-By-Step
curl -fsSL https://chatgpt.com/codex/install.sh | sh
codex # the first run walks you through ChatGPT sign-in
AGENTS.md, which Codex reads automatically and which Claude Code can also read, so your rules travel between vendors./permissions, or your Codex configuration, to allow codex exec to edit files and run tests for this task, and keep the settings conservative everywhere else.# One branch and one folder per package, so the agents never collide
for pkg in auth billing search notify; do
git worktree add ../wt-$pkg -b tests/$pkg
done
# Each subshell is one Codex agent working inside its own worktree
for pkg in auth billing search notify; do
( cd ../wt-$pkg && \
codex exec "Write missing unit tests for packages/$pkg. Run them until they pass. \
Do not modify any other package. Commit when done." \
> ../log-$pkg.md 2>&1 ) & # run in the background and save each agent's report
done
wait # block until all four agents have finished
codex exec "Review branches tests/auth, tests/billing, tests/search and tests/notify \
against main. Flag weak assertions, flaky tests and missing edge cases. \
Output a numbered list per branch." > review.md
git worktree remove ../wt-<package>.Estimated Bill
This estimate assumes API-equivalent pricing of $1.75 per million input tokens and $14 per million output tokens, which is the GPT-5.3-Codex rate reported in Jetadmin’s Codex pricing breakdown.
| Agent | Tokens In / Out | Cost |
|---|---|---|
| 4 test-writing agents | 500K / 40K each | $5.74 |
| reviewer | 300K / 20K | $0.81 |
| Total | ≈ $6.55 |
On a Plus plan at $20 a month, a run of this size typically fits inside your rolling five-hour usage window.
Section V. xAI Grok
In Brief
- Grok Build is xAI’s terminal coding agent, and its CLI entered early beta on May 25, 2026, for SuperGrok and X Premium Plus subscribers.
- xAI open-sourced the harness on July 15, 2026, according to the official xAI news timeline.
- It offers plan mode, subagents, a headless mode for scripting, and a fast dedicated coding model called Grok Build 0.1.
- An independent research note on Grok Build’s July 2026 upload incident reports that it uploaded repository bundles against users’ wishes until a server-side fix stopped it, so keep it on open or throwaway code for now.
The Perfect Task: Rapid Prototyping With Competing Variants
Grok Build is ideal when speed matters more than secrecy. Prototypes are disposable, they are often public anyway, and they benefit enormously from a fast model. So the best use is to let several agents race each other and let a critic agent pick the winner. The task here is to build three competing landing-page prototypes for a new open-source tool and then judge them.
The Agent Team
| Agent | Job |
|---|---|
| Lead (your session) | Writes the brief, then spawns and judges the agents |
proto-minimal |
Builds a minimal, typography-led design |
proto-bold |
Builds a bold, illustration-heavy design |
proto-docs |
Builds a documentation-first, developer-focused design |
critic |
Scores all three designs against the brief |
Step-By-Step
curl -fsSL https://x.ai/cli/install.sh | bash # macOS / Linux / WSL
# Windows PowerShell: irm https://x.ai/cli/install.ps1 | iex
mkdir landing-race && cd landing-race && git init
grok # the first launch opens browser sign-in
Spawn three subagents in parallel, each in its own git worktree:
1. proto-minimal: a minimal, typography-led landing page.
2. proto-bold: a bold, illustration-heavy landing page.
3. proto-docs: a docs-first page with a live code sample.
Each must follow the approved brief and build as static HTML/CSS.
Spawn a fourth subagent, critic. It must not edit anything.
Score each prototype from 1 to 10 on clarity, speed and fit to the brief, then recommend one.
# A headless one-shot with streaming JSON output, handy for nightly variant runs
grok -p "Re-run the critic on all three worktrees" --output-format streaming-json
Estimated Bill
This estimate assumes $2 per million input tokens and $6 per million output tokens, the Grok 4.6 API price listed in the independent research note above. My confidence in this figure is low, because xAI’s pricing has changed several times this year.
| Agent | Tokens In / Out | Cost |
|---|---|---|
| Lead | 100K / 10K | $0.26 |
| 3 prototype agents | 200K / 30K each | $1.74 |
| critic | 150K / 10K | $0.36 |
| Total | ≈ $2.36 |
On a SuperGrok subscription, CLI usage is covered by the plan rather than billed per token.
Section VI. Google Gemini
In Brief
- Gemini CLI stopped serving free users, Google AI Pro subscribers, and Google AI Ultra subscribers on June 18, 2026.
- Its replacement is Antigravity CLI, a new binary called
agythat Google rebuilt in Go and that shares its agent harness with the Antigravity 2.0 desktop app. - Antigravity is designed for asynchronous, multi-agent work, so agents can run in the background while you keep using the terminal.
- These details come from Google’s official post on moving Gemini CLI to Antigravity CLI.
The Perfect Task: Map A Legacy Monorepo
Gemini models are built for very long context windows, and Antigravity adds background agents that keep working while you do something else. That combination is perfect for reading a huge, unfamiliar codebase without blocking your day. The task here is to produce an architecture brief for a legacy monorepo with four subsystems.
The Agent Team
| Agent | Job |
|---|---|
| Lead (your session) | Splits the repository and merges the reports |
map-api |
Maps the API gateway and its routes |
map-data |
Maps the database layer and its migrations |
map-jobs |
Maps background jobs and queues |
map-web |
Maps the front end and every API call it makes |
Step-By-Step
curl -fsSL https://antigravity.google/cli/install.sh | bash # installs `agy` to ~/.local/bin
# Windows PowerShell: irm https://antigravity.google/cli/install.ps1 | iex
gemini command in your scripts and CI configuration, because those calls stopped working on June 18 if they used consumer credentials.# Search scripts and CI configs for lingering references to the retired CLI
grep -rn --include="*.yml" --include="*.yaml" --include="*.sh" "gemini " .
agy from the root of the repository and confirm workspace trust during the first-run setup.Spawn four background agents, each read-only:
1. map-api: map the gateway, routes and auth flow under services/api/.
2. map-data: map schemas, migrations and data access under services/db/.
3. map-jobs: map queues, workers and schedules under services/jobs/.
4. map-web: map the front end and every API call it makes under apps/web/.
Each writes its findings to docs/arch/<agent-name>.md, max 80 lines.
/agents to watch and manage them.Read docs/arch/*.md and write docs/ARCHITECTURE.md: a one-page overview,
a dependency list between subsystems, and the top five risks.
Estimated Bill
This estimate assumes a Gemini Pro-class rate of $2 per million input tokens and $12 per million output tokens.
| Agent | Tokens In / Out | Cost |
|---|---|---|
| Lead | 200K / 20K | $0.64 |
| 4 mapping agents | 800K / 20K each | $7.36 |
| Total | ≈ $8.00 |
Reading is the expensive part of this task, which is exactly why it belongs on a long-context model with a free tier.
Section VII. Microsoft Copilot
In Brief
- For coding agents, the Microsoft product that matters is GitHub Copilot, while Microsoft 365 Copilot is a separate product for documents and email.
- The Copilot CLI includes
/fleet, which runs parallel subagents, and/delegate, which hands a task to the cloud coding agent. - The plans are Free, Pro at $10, Pro+ at $39, and Max at $100, with extra usage at $0.01 per credit, according to GitHub’s Copilot plans page.
The Perfect Task: A Dependency Upgrade Across Many Packages
Copilot is the best tool for work that lives right next to your pull requests. The /fleet command splits a repetitive job into parallel subagents, and /delegate sends the stubborn leftovers to a cloud agent that opens pull requests for you. The task here is to upgrade a logging library across 12 packages and then hand off the difficult cases.
The Agent Team
| Agent | Job |
|---|---|
| Main agent (your session) | Plans the upgrade and coordinates the work |
12 /fleet subagents |
One per package: upgrade, fix, and test |
2 /delegate cloud agents |
Finish the packages that failed and open pull requests |
Step-By-Step
npm install -g @github/copilot
# or: brew install --cask copilot-cli | winget install GitHub.Copilot
copilot, then sign in once with /login, and note that organisation members need an administrator to enable the Copilot CLI policy first./model to choose a mid-tier model, because mechanical edits do not need a frontier model./fleet instruction./fleet Upgrade our logging library from v3 to v4 in every package under packages/.
Use one subagent per package. Each subagent updates imports, fixes breaking
changes, runs that package's tests, and reports PASS or FAIL with one line of detail.
/delegate, and repeat the command for the second failed package./delegate Finish the logging v4 upgrade in packages/payments. Tests are failing on
structured-field serialisation. Open a pull request when tests pass.
copilot -p "List every package still importing logging v3" -s
Estimated Bill
This estimate assumes that AI credits track the underlying model cost at $0.01 per credit, using a mid-tier model priced at $2 per million input tokens and $10 per million output tokens.
| Agent | Tokens In / Out | Cost |
|---|---|---|
| Main agent | 200K / 20K | $0.60 |
| 12 fleet subagents | 150K / 10K each | $4.80 |
| 2 cloud agents | Rough allowance | $1.00 |
| Total | ≈ $6.40 (about 640 credits) |
Copilot Pro’s $15 of monthly credits covers roughly two runs of this size.
Section VIII. OpenCode
In Brief
- OpenCode is an open-source, provider-agnostic terminal agent that works with almost any model provider.
- It has two primary agents, Build with full access and Plan with editing disabled by default, and you switch between them with the
Tabkey. - It ships with three subagents—General, Explore, and Scout—which you call with
@name. - Every agent can run on a different model, according to OpenCode’s agents documentation, and the client itself is free.
The Perfect Task: Write API Documentation For A Whole Repository
OpenCode is the best tool for mixing cheap and premium models inside one team. Documentation needs a great deal of reading and only a little excellent writing, so cheap models should do the reading and one premium model should do the writing. The task here is to produce reference documentation for every public module and then fact-check it against the code.
The Agent Team
| Agent | Model | Job |
|---|---|---|
| Build (primary) | Sonnet 5.5 | Writes the final documentation |
doc-scout × 5 |
deepseek-flash | Extracts signatures and behaviour, read-only |
doc-checker |
deepseek-flash | Verifies every documented claim against the code |
Step-By-Step
curl -fsSL https://opencode.ai/install | bash
opencode, then use /connect to add two providers: Anthropic for the writer and DeepSeek for the readers./init so that OpenCode analyses the project and generates an AGENTS.md file.doc-scout, as a read-only subagent on a very cheap model.mkdir -p ~/.config/opencode/agents
cat > ~/.config/opencode/agents/doc-scout.md <<'EOF'
---
description: Read-only extractor. Lists public functions, parameters, return values and side effects for ONE module.
mode: subagent # callable by the primary agent via @doc-scout
model: deepseek/deepseek-flash # a very low-cost model for high-volume reading
permission:
edit: deny # cannot change files
bash: deny # cannot run commands
---
For the module you are given, return a compact list of public APIs with one-line behaviour notes.
EOF
doc-checker, which compares the finished documentation against the source code.cat > ~/.config/opencode/agents/doc-checker.md <<'EOF'
---
description: Fact-checks documentation against source code. Read-only.
mode: subagent
model: deepseek/deepseek-flash
permission:
edit: deny
bash: deny
---
Compare each documented claim with the code. Return only the claims that are wrong, with file:line.
EOF
/models to select Sonnet 5.5 as the model for the Build agent, which will do the actual writing.Call @doc-scout five times in parallel, once each for src/core, src/api, src/db,
src/auth and src/utils. Using their output, write docs/reference/<module>.md files.
Then call @doc-checker on every new file and fix each error it reports.
/undo and /redo freely if the output misses the mark, because they make experimentation cheap.Estimated Bill
| Agent | Tokens In / Out | Cost |
|---|---|---|
| 5 doc-scouts (flash, peak $0.30 / $1.20) | 400K / 30K each | $0.78 |
| Writer (Sonnet 5.5, $2 / $10) | 300K / 40K | $1.00 |
| doc-checker (flash) | 200K / 10K | $0.07 |
| Total | ≈ $1.85 |
That is the real power of per-agent models: seven agents for less than the price of a cup of filter coffee at a Chennai café chain.
Section IX. Pi Code
In Brief
- Pi is a deliberately minimal coding agent created by Mario Zechner, the developer behind libGDX.
- It has four core tools—read, write, edit, and bash—along with a short, publicly visible system prompt and support for more than fifteen providers.
- It runs interactively, in print or JSON mode, over RPC, or embedded inside your own application through an SDK.
- It intentionally ships without built-in subagents, and the official Pi website suggests using tmux or an extension instead.
The Perfect Task: Localise An App Into Four Languages
Pi is the best tool for batch jobs that you script yourself, because each Pi process is a small, predictable worker. Run several of them from one shell script and you have an orchestrator with no framework overhead at all. The task here is to translate an app’s English strings into Hindi, Tamil, Spanish, and German, and then verify the translations.
The Agent Team
| Agent | Job |
|---|---|
translate-hi |
Translates English into Hindi |
translate-ta |
Translates English into Tamil |
translate-es |
Translates English into Spanish |
translate-de |
Translates English into German |
verifier |
Checks keys, placeholders, and length limits |
Step-By-Step
curl -fsSL https://pi.dev/install.sh | sh
# npm alternative: npm install -g @earendil-works/pi-coding-agent
pi, connect a provider, and use /model to select a small, inexpensive model such as Haiku 5.5.#!/bin/bash
# localise.sh — creates four translator agents and one verifier agent.
set -e
SRC=locales/en.json
# Agents 1-4: one Pi process per language, all running in parallel
for lang in hi ta es de; do
pi -p "Translate every value in $SRC into language code '$lang'. Keep all keys and \
{placeholders} unchanged. Write the result to locales/$lang.json." \
> logs/translate-$lang.md 2>&1 &
done
wait # block until all four translators have finished
# Agent 5: the verifier checks all four outputs against the English source
pi -p "Compare locales/en.json with hi.json, ta.json, es.json and de.json. Report missing keys, \
broken {placeholders} and button strings longer than 40 characters." > logs/verify.md
mkdir -p logs && chmod +x localise.sh && ./localise.sh
Estimated Bill
| Agent | Tokens In / Out | Cost (Haiku 5.5, $1 / $5) |
|---|---|---|
| 4 translators | 100K / 30K each | $1.00 |
| verifier | 150K / 10K | $0.20 |
| Total | ≈ $1.20 |
Pi will not hold your hand at any point, and that is precisely why experienced developers love it.
Section X. DeepSeek
In Brief
- DeepSeek matters in this article as an engine that powers other agents, rather than as a separate agent of its own.
- It offers two models,
deepseek-flashanddeepseek-v4-pro, both with one-million-token context windows, and off-peak prices are half of peak prices according to the DeepSeek API pricing page. - V4-Pro launched on August 14, 2026, alongside peak price increases of up to 1,100%, as Caixin reported in its coverage of the V4-Pro launch.
- DeepSeek is a Chinese provider, so check your data-residency and contractual obligations before you send it any client code.
The Perfect Task: An Overnight Type-Hint Migration
DeepSeek is the best engine for high-volume, low-risk, well-tested mechanical work. The changes are repetitive, the test suite catches mistakes, and off-peak pricing halves the bill while you sleep. The task here is to add type hints to 200 Python files and make mypy pass across the whole codebase.
The Agent Team
DeepSeek’s Anthropic-compatible API guide maps Claude model names onto its own models, which lets you run this team inside Claude Code.
| Agent | Requested Model | Served By | Job |
|---|---|---|---|
| Lead (your session) | opus | deepseek-v4-pro | Plans the batches and fixes hard cases |
typer-1 to typer-4 |
haiku | deepseek-flash | 50 files each: add hints, run mypy |
Step-By-Step
export ANTHROPIC_BASE_URL=https://api.deepseek.com/anthropic # DeepSeek's Anthropic-compatible endpoint
export ANTHROPIC_API_KEY="sk-..." # your DeepSeek key, not an Anthropic key
mkdir -p .claude/agents
for i in 1 2 3 4; do
cat > .claude/agents/typer-$i.md <<EOF
---
name: typer-$i
description: Adds type hints to batch $i of the Python files and makes mypy pass for that batch.
model: haiku # DeepSeek serves this request as deepseek-flash
tools: Read, Edit, Bash
---
Work only on the files listed in batches/batch-$i.txt. Add precise type hints.
Run mypy on those files until clean. Never change runtime behaviour.
EOF
done
mkdir -p batches
git ls-files '*.py' > batches/all.txt # every tracked Python file
split -n l/4 -d -a 1 batches/all.txt batches/tmp- # four roughly equal file lists
for i in 0 1 2 3; do mv batches/tmp-$i batches/batch-$((i+1)).txt; done
opus request as deepseek-v4-pro.claude --model opus
Run typer-1, typer-2, typer-3 and typer-4 in parallel. When all finish, run the full
test suite and mypy across the repo. Fix any remaining errors yourself.
Estimated Bill (Off-Peak)
| Agent | Tokens In / Out | Cost |
|---|---|---|
| Lead (v4-pro, $0.66 / $1.98) | 500K / 50K | $0.43 |
| 4 typers (flash, $0.15 / $0.60) | 1.5M / 150K each | $1.26 |
| Total | ≈ $1.70 (about $3.40 at peak) |
Two hundred type-hinted files for under two dollars is, quite literally, all in a night’s work.
Section XI. Master Agent Workflow (From Anthropic)
Everything above is tooling, and this section is the method that ties it together. Anthropic has published the clearest public blueprint for multi-agent work in two pieces.
- Its December 2024 essay on building effective agents names and explains the orchestrator-workers pattern.
- Its June 2025 engineering post on the multi-agent research system shows that same pattern running in production.
The Architecture
The Results
The results were not subtle. A system with Claude Opus 4 as the lead and Claude Sonnet 4 workers outperformed a single Claude Opus 4 agent by 90.2% on Anthropic’s internal research evaluation. Parallel tool calling also cut research time by up to 90% on complex queries.
Five Lessons To Apply Daily
The Universal Lead Prompt
You can paste the prompt below into any harness that supports subagents, and it will apply the whole Anthropic method in one go.
You are the LEAD agent. Do not write code yet.
1. Read AGENTS.md and the task below. Write a plan to PLAN.md.
2. Classify the task: SIMPLE (do it yourself), COMPARISON (2-4 workers), or BROAD (5+ workers).
3. For each worker, define: objective, files in scope, files OUT of scope,
output format (max 30 lines), and model tier (small or medium).
4. Run workers in parallel. Collect results in FINDINGS.md.
5. Synthesise and implement on a branch yourself.
6. Spawn ONE reviewer on a different model to try to break your change.
Budget: stop and report if you exceed 15 worker calls.
In short, the lead plans, the workers explore, the lead builds, and a stranger checks the result.
Section XII. When To Use Multiple Agents
Multiple agents pay for themselves when the work is wide, the pieces are independent, and the results are easy to verify. Every task in Sections III to X fits that shape, which is why each of them earns its bill.
The Patterns That Earn Their Bill
- Broad reading of a large codebase, as in the monorepo map from Section VI.
- Parallel, independent edits across many packages, as in the dependency upgrade from Section VII.
- Specialist roles working on one feature, as in the full-stack build from Section III.
- Competing drafts judged by a critic, as in the prototype race from Section V.
- Cheap readers feeding one premium writer, as in the documentation team from Section VIII.
- Batch processing of repetitive work, as in the localisation and type-hint tasks from Sections IX and X.
- Cross-vendor review, where Claude writes the code and Codex reviews it, or the other way around.
What These Patterns Make Possible
- Mapping a 400,000-line monorepo in a single afternoon – easy.
- Upgrading a dependency across 12 packages with tests passing – all in a day’s work.
- Getting a second vendor’s model to tear apart your pull request – no problem.
- Localising an app into four languages before lunch – sure thing!
- Waking up to 200 type-hinted files and a clean
mypyrun – done!
If the work is wide, if the pieces are independent, and if you can check the result cheaply, then more agents genuinely means more work done, faster.
Section XIII. When Not To Use Multiple Agents
Now; for the other side of the coin! Anthropic’s own post says that multi-agent systems are a poor fit where agents must share context or depend heavily on each other, and it names most coding tasks as an example. Cognition, the team behind Devin, makes the case even more bluntly in its essay Don’t Build Multi-Agents.
Hold on, Thomas, you might say. You have just shown eight separate multi-agent setups, the benchmarks favour multi-agent systems, and every vendor is shipping fleets, teams, and swarms. Surely the direction of travel is obvious to everyone by now.
That is true, and I will not pretend otherwise. But the direction of travel across the industry is not the same thing as the task in front of you on a Tuesday afternoon.
Do Not Use Multiple Agents When
- The change is tightly coupled, because shared state across five files needs one mind holding all five files at once.
- The task is small, because a bug with a clear stack trace needs one agent and one test, not a team.
- You cannot verify the output cheaply, because ten unreviewed diffs are technical debt rather than productivity.
- The budget is tight, because the multipliers are roughly 4X for an agent, 7X for an agent team, and 15X for a multi-agent system.
- The code is sensitive, because every extra vendor is one more place your code travels.
My Rule Of Thumb
Start with one agent and a good plan, and add a second agent only when you can name the specific reason—isolation, parallelism, or independent review. If you cannot name the reason, you almost certainly do not need the second agent.
Section XIV. Top Five Best Agent Harnesses
A harness is everything that surrounds the model: the tools, the permissions, the context management, and the subagent machinery. The same model can perform very differently inside different harnesses, which is why rankings such as Terminal-Bench 2.0 score harness-and-model pairs rather than models alone. This ranking reflects daily multi-agent work in October 2026, and it is my opinion, grounded in the evidence presented above.
codex exec makes headless fan-out trivial. My confidence in this ranking is high./fleet and /delegate commands sit right next to your pull requests, which is a real advantage.There are three honourable mentions worth watching.
- Antigravity CLI has a strong architecture, but it is still settling down after the forced migration from Gemini CLI.
- Grok Build is moving quickly, but the July incident keeps it off client work for now.
- ForgeCode and Factory’s Droid both post strong Terminal-Bench 2.0 results and deserve a test drive.
Section XV. Skills Required For Agentic Orchestration
Orchestration looks like a tooling problem, but it is really a skills problem. The tools in this article are all learnable in an afternoon, while the skills below take months to sharpen—and they are what separates a developer who runs agents from a developer who conducts them.
CLAUDE.md or AGENTS.md, what belongs in a skill that loads on demand, and what should never enter the context at all. A lean context is both cheaper and more accurate.git worktree is new to you, learn it before you run your first parallel team./usage output, and choose the cheapest model that can do each job. Every bill in this article is an exercise in exactly this skill.codex exec, claude -p, copilot -p, and pi -p turn agents into building blocks for scripts and CI pipelines. A little Bash goes a very long way in orchestration.The fastest way to build these skills is to start small and deliberate. Anthropic’s free Claude Academy courses, including Claude Code 101 and Claude Code in Action, are an excellent starting point, and every vendor’s documentation linked in the References section is worth reading end to end.
Section XVI. Why Local LLMs Are A Game-Changer For Low Budgets
Every bill in this article so far has been a token bill, and token bills grow with every agent you add. Local models change that equation completely. Once you own the hardware, the marginal cost of a token is a little electricity, so running five agents costs almost the same as running one.
Let me be precise about what “low budget” means here. Local models give you a very low running budget, not a low starting budget, because the hardware is a real one-time purchase. If you cannot make that purchase today, the cloud options in Sections II and X remain the right answer, and nothing in this section changes that.
The Minimum: 64 GB Of Unified Memory
I now treat 64 GB of unified memory as the minimum for serious local agent orchestration. A 32 GB machine can run one good coding model at 4-bit precision with a modest context window, but it cannot hold a proper agent team. At 64 GB, everything that makes orchestration work locally becomes possible.
- Higher-quality weights. The strongest 27-billion-parameter coding models fit at 8-bit precision, which is about 30 GB on Ollama, instead of being squeezed down to 4-bit.
- Room for long contexts. Agents read whole files and long logs, and every extra token of context needs memory of its own on top of the model weights.
- Several agents at once. Each parallel agent keeps its own context cache, so parallelism is a memory question before it is a speed question.
- Two different models side by side. A strong lead model and a fast worker model can stay loaded together, which is exactly the lead-and-workers pattern from Section XI.
macOS normally lets the GPU use about 75% of unified memory, which is roughly 48 GB on a 64 GB Mac. The LLM Configurator guide to raising the Metal memory limit shows how to lift that to about 54 GB with a single sysctl command while leaving at least 10 GB for macOS. That 48–54 GB budget is what every recommendation in Sections XVI and XVII is designed around.
The Cheapest Ways To Reach 64 GB In October 2026
- Mac mini with M5 Pro, 64 GB. This starts at about $2,699, calculated from the base price and the memory upgrade price in AppleInsider’s M5 Pro versus M4 Pro Mac mini comparison.
- Mac mini with M4 Pro, 64 GB. This works out to about $2,199 at current pricing, but AppleInsider notes that the M4 Pro model is now in limited supply.
- Mac Studio with M5 Max, 64 GB. This costs about $3,499 and doubles the memory bandwidth, which makes every token arrive noticeably faster.
The new M6 Mac mini tops out at 32 GB, so it cannot meet the 64 GB minimum at any price. The Appendix lists every current configuration, including the higher memory tiers.
Why Local Changes Everything
- Zero marginal token cost. You can run a five-agent team all day, rerun failed experiments, and let agents read entire repositories without watching a per-token bill climb.
- No usage windows or caps. There are no rolling five-hour limits, weekly caps, AI-credit allowances, or surprise overages, because the only limit is your own hardware.
- Complete privacy. Your code never leaves your machine, which removes the data-residency worries raised in Section X and makes local agents ideal for confidential client work.
- Predictable budgeting. Hardware is a fixed cost you can plan for, and a $2,699 Mac mini that replaces $200 of monthly cloud spend pays for itself in a little over a year—or in under six months if it replaces $500 a month.
- Protection from token price changes. When a provider raises prices, as DeepSeek did in August 2026, a local setup keeps working at exactly the same cost.
- A free orchestration laboratory. Because experiments cost nothing, you can practise every pattern in Section XI—fan-out, competing drafts, adversarial review—until it becomes second nature.
The Honest Trade-Offs
Local models are a game-changer, but they are not magic, and you should go in with clear eyes.
- Quality still trails the frontier. Qwen3.8-27B scores 73.0 on Terminal-Bench 2.1 against 78.2 for Claude Opus 4.6 Max on its official model card, which is remarkably close but still behind.
- Speed depends on memory bandwidth. The M5 Pro Mac mini offers 307 GB/s while the M5 Max Mac Studio offers up to 614 GB/s, and token generation speed on Apple silicon follows bandwidth closely.
- Parallel agents share one machine. Several local agents can run at once, but they split the same memory and compute, so ten local agents are not ten times faster than one.
- Hardware prices are rising. Apple raised Mac prices across the board in June 2026 because of the memory shortage, and the Appendix explains why that trend is likely to continue.
The Hybrid Pattern That Wins On A Budget
The smartest budget setup is usually a hybrid. Let a cloud model act as the lead for planning and final review, and let local models do the high-volume worker jobs such as reading, searching, writing tests, and summarising logs. OpenCode makes this easy, because each agent can point at a different provider, so a single team can mix one paid model with several free local ones.
Step-By-Step: A Fully Local Agent Team In Claude Code
The task here is to refactor a confidential client module and add tests for it, without a single line of code leaving your machine. Ollama exposes an Anthropic-compatible API, so Claude Code can run entirely on a local model, as described in Ollama’s Claude Code integration guide. On a 64 GB Mac, the whole team can run on the 8-bit version of Qwen3.8-27B.
| Agent | Model | Job |
|---|---|---|
| Lead (your session) | qwen3.8:27b-q8_0 (local) | Plans the refactor and integrates the changes |
refactorer |
Same local model | Restructures the module without changing behaviour |
test-writer |
Same local model | Writes tests that pin down current behaviour |
reviewer |
Same local model | Reviews the final diff, read-only |
# Linux install script; on macOS, download the Ollama app from ollama.com instead
curl -fsSL https://ollama.com/install.sh | sh
# Lets the GPU use up to 54 GB (55296 MB) of unified memory until the next reboot
sudo sysctl iogpu.wired_limit_mb=55296
export OLLAMA_CONTEXT_LENGTH=65536 # Claude Code needs a large context window to work well
export OLLAMA_NUM_PARALLEL=2 # two agents can run concurrently; drop to 1 under memory pressure
ollama serve
ollama pull qwen3.8:27b-q8_0 # about 30 GB; leaves roughly 18-24 GB for context on a 64 GB Mac
mkdir -p .claude/agents
# Agent 1: refactorer — may edit only the target module
cat > .claude/agents/refactorer.md <<'EOF'
---
name: refactorer
description: Restructures src/billing/ for readability without changing behaviour.
model: inherit # use the same local model as the lead session
tools: Read, Edit, Bash
---
Refactor only files under src/billing/. Never change public signatures or behaviour.
EOF
# Agent 2: test-writer — writes characterisation tests before and after the refactor
cat > .claude/agents/test-writer.md <<'EOF'
---
name: test-writer
description: Writes tests that capture the current behaviour of src/billing/.
model: inherit
tools: Read, Write, Bash
---
Write tests under tests/billing/ that pin down current behaviour. Run them and report results.
EOF
# Agent 3: reviewer — read-only final check
cat > .claude/agents/reviewer.md <<'EOF'
---
name: reviewer
description: Reviews the diff for behaviour changes and weak tests. Read-only.
model: inherit
tools: Read, Bash
---
Compare the branch with main. List any behaviour change or weak test, most severe first.
EOF
export ANTHROPIC_BASE_URL=http://localhost:11434 # Ollama's Anthropic-compatible endpoint
export ANTHROPIC_AUTH_TOKEN=ollama # required by the client, ignored by Ollama
export ANTHROPIC_API_KEY="" # make sure no cloud key is used
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 # turn off telemetry and other optional calls
claude --model qwen3.8:27b-q8_0
First use test-writer to capture the current behaviour of src/billing/.
Then use refactorer to restructure the module while those tests stay green.
Finally use reviewer, and fix every issue it reports.
git diff and your test suite, and run ollama ps to see which model is loaded and how much memory it is using.Estimated Bill
| Item | Assumption | Cost |
|---|---|---|
| Tokens | All four agents run locally | $0.00 |
| Electricity (64 GB Mac mini) | About 0.1 kW for 2 hours at $0.10–$0.20 per kWh | ≈ $0.02–$0.04 |
| Electricity (64 GB Mac Studio) | About 0.2 kW for 1.5 hours at the same rates | ≈ $0.03–$0.06 |
| Total | Well under $0.10 |
The same team on cloud APIs would cost several dollars, as Section III showed, and every rerun would cost the same again. Locally, a rerun costs a few cents of electricity, and that is what makes local models a genuine game-changer for anyone who wants to keep running costs low.
Section XVII. Best Local LLMs For Agentic Coding With A 64 GB Unified Memory Minimum
Sixty-four gigabytes is the floor for a real local agent team in 2026, not a luxury. It is the smallest memory size that holds the best 27-billion-parameter coding models at 8-bit precision with room to spare, and it is the smallest size that keeps a lead model and a worker model loaded at the same time.
How To Budget 64 GB Of Memory
A model has to fit together with its working memory, not on its own, so plan your 64 GB with these rules in mind.
- Start from about 48 GB, not 64 GB. macOS reserves roughly a quarter of unified memory by default, and you can raise the GPU’s share to about 54 GB with
sudo sysctl iogpu.wired_limit_mb=55296, as Step 2 of Section XVI showed. - Choose 8-bit weights where you can. On Ollama, Qwen3.8-27B is 18 GB at 4-bit and 30 GB at 8-bit, according to its Ollama tag list, and 64 GB is what makes the higher-quality version practical.
- Reserve memory for context and parallel agents. The KV cache grows with context length and with every parallel agent, so a 64K context for two or three agents can take many gigabytes on top of the weights.
- Keep two models loaded when the team needs them. Setting
OLLAMA_MAX_LOADED_MODELS=2lets a lead model and a worker model stay resident together, so Ollama does not have to swap them in and out between turns. - Mixture-of-experts models trade quality for speed. A model with about 3 billion active parameters generates tokens much faster than a dense model, while dense models usually score higher per gigabyte of memory.
The Shortlist
All scores below are results published by each model’s maker or reported in the source linked beside it. SWE-bench Verified, SWE-bench Pro, LiveCodeBench and τ²-Bench measure different skills, so compare scores only within the same benchmark. Always test a model on your own code before you commit to it.
| Rank | Model | Maker | Type | Ollama Tag And Size | Context | Published Coding Score |
|---|---|---|---|---|---|---|
| 1 | Qwen3.8-27B | Alibaba Qwen | Dense, 27B | qwen3.8:27b-q8_0, 30 GB |
256K | SWE-bench Pro 61.7; Terminal-Bench 2.1 73.0 |
| 2 | Qwen3.6-27B | Alibaba Qwen | Dense, 27B | qwen3.6:27b-q8_0, 30 GB |
256K | SWE-bench Verified 77.2 |
| 3 | Qwen3.6-35B-A3B | Alibaba Qwen | MoE, 35B total, 3B active | qwen3.6:35b-a3b-q8_0, 39 GB |
256K | SWE-bench Verified 73.4 |
| 4 | Laguna XS 2.1 | Poolside | MoE, 33B total, 3B active | laguna-xs-2.1:q8_0, 36 GB |
256K | No score in the sources I checked |
| 5 | Muse Glimmer 30B | Meta | Dense, 30B | muse-glimmer, about 17 GB at 4-bit |
128K | SWE-bench Verified 76.0 |
| 6 | Devstral Small 2 | Mistral AI | Dense, 24B | devstral-small-2:24b, 15 GB |
256K | SWE-bench Verified 68.0 |
| 7 | Gemma 4 31B | Dense, 31B | gemma4:31b-it-q8_0, 34 GB | 256K | LiveCodeBench v6 80.0; τ²-Bench 76.9 | |
| 8 | Gemma 4 26B-A4B | MoE, 25.2B total, 3.8B active | gemma4:26b-a4b-it-q8_0, 28 GB | 256K | LiveCodeBench v6 77.1; τ²-Bench 68.2 | |
| 9 | GLM-4.7-Flash | Z.ai | MoE, about 30B total, 3B active | glm-4.7-flash:q8_0, 32 GB | 198K | SWE-bench Verified 59.2; τ²-Bench 79.5 |
| 10 | Gemma 4 12B | Dense, 12B | gemma4:12b-it-q8_0, 13 GB | 256K | LiveCodeBench v6 72.0; τ²-Bench 69.0 |
Two popular models do not make this list for a simple reason: they do not fit. OpenAI’s gpt-oss-120b is 65 GB on Ollama and Qwen3.5-122B-A10B is 81 GB, both above a 64 GB Mac’s GPU budget, so they belong to the higher tiers covered in the Appendix.
What Each Model Is Best At
Step-By-Step: A Two-Model Security-Audit Team In OpenCode
The task here is to audit a private repository for security issues with three parallel scout agents and one verifier, all running locally. This is the setup that 64 GB makes possible and 32 GB does not: a strong dense model leads and verifies, while a fast mixture-of-experts model does the parallel reading.
| Agent | Model | Job |
|---|---|---|
| Build (primary) | qwen3.8:27b at 4-bit, 18 GB | Splits the audit and writes the final report |
sec-scout × 3 |
qwen3.6:35b-a3b at 4-bit, 24 GB | Audits authentication, input handling, and dependencies, read-only |
sec-verifier |
qwen3.8:27b (shared with the lead) | Re-checks every finding against the code, read-only |
sudo sysctl iogpu.wired_limit_mb=55296 # about 54 GB for the GPU, at least 10 GB left for macOS
ollama pull qwen3.8:27b # lead and verifier, about 18 GB at 4-bit
ollama pull qwen3.6:35b-a3b # fast scouts, about 24 GB at 4-bit
export OLLAMA_MAX_LOADED_MODELS=2 # keep the lead and the worker model resident together
export OLLAMA_NUM_PARALLEL=3 # three scouts at once; lower this if memory runs short
export OLLAMA_CONTEXT_LENGTH=32768 # 32K per request keeps the KV cache inside the budget
ollama serve
opencode.json, using Ollama’s OpenAI-compatible endpoint.{
"$schema": "https://opencode.ai/config.json",
"provider": {
"ollama": {
"npm": "@ai-sdk/openai-compatible",
"name": "Ollama (local)",
"options": { "baseURL": "http://localhost:11434/v1" },
"models": {
"qwen3.8:27b": { "name": "Qwen3.8 27B (local lead)" },
"qwen3.6:35b-a3b": { "name": "Qwen3.6 35B-A3B (local workers)" }
}
}
}
}
sec-scout agent on the fast worker model, which the Build agent will call three times with three different scopes.mkdir -p ~/.config/opencode/agents
cat > ~/.config/opencode/agents/sec-scout.md <<'EOF'
---
description: Read-only security scout. Audits ONE area of the codebase and lists concrete findings.
mode: subagent
model: ollama/qwen3.6:35b-a3b # fast MoE model for parallel reading
permission:
edit: deny # scouts never change code
bash: deny
---
Audit only the area you are given. For each finding return: file:line, severity, and a one-line explanation.
EOF
sec-verifier agent on the stronger lead model, so a different model checks the scouts’ work.cat > ~/.config/opencode/agents/sec-verifier.md <<'EOF'
---
description: Re-checks security findings against the source code. Read-only.
mode: subagent
model: ollama/qwen3.8:27b # the stronger dense model verifies the faster model's findings
permission:
edit: deny
bash: deny
---
For each finding, confirm or reject it with evidence from the code. Return only confirmed findings.
EOF
opencode, use /models to select qwen3.8:27b for the Build agent, and give it one instruction.Call @sec-scout three times in parallel: one for authentication and sessions,
one for input validation and injection risks, and one for dependencies and secrets.
Pass all findings to @sec-verifier. Write the confirmed findings to SECURITY-AUDIT.md,
grouped by severity, with a suggested fix for each.
ollama ps while the audit runs to confirm that both models are loaded and that memory use stays within your budget.Estimated Bill
| Item | Assumption | Cost |
|---|---|---|
| Tokens | Five agents on two local models | $0.00 |
| Electricity | About 0.1–0.2 kW for 1 hour at $0.10–$0.20 per kWh | ≈ $0.01–$0.04 |
| Total | A few cents |
My Recommended Setups For 64 GB
- Best quality: Qwen3.8-27B at 8-bit for every agent, with two parallel slots and a 64K context.
- Best speed: Qwen3.6-35B-A3B for every agent, at 8-bit when it runs alone or at 4-bit when you want three or more parallel slots.
- Best two-model team: Qwen3.8-27B as the lead and verifier, with Qwen3.6-35B-A3B or Devstral Small 2 running the worker agents.
- Best agentic experiment: Laguna XS 2.1 at 8-bit, tested against Qwen3.8-27B on a few of your own tasks before you switch.
- Best hybrid: a cloud model as the lead in OpenCode, with Qwen3.6-35B-A3B running every local worker agent.
Sixty-four gigabytes of memory, a free harness, and an open-weight model now add up to a genuine two-model coding team that costs cents to run. If your work grows beyond that, the Appendix shows what each higher memory tier adds and what it costs today.
Section XVIII. Summary
Here is the whole article on one screen, with every task, team size, and rough bill side by side. The last two rows assume a 64 GB Mac, which this article treats as the minimum for local agent teams.
| Platform | Perfect Task | Agents | Rough Bill |
|---|---|---|---|
| Claude Code | Full-stack feature build | 5 | ≈ $5.10 |
| OpenAI Codex | Parallel test-coverage sweep | 5 | ≈ $6.55 |
| xAI Grok Build | Competing landing-page prototypes | 5 | ≈ $2.36 |
| Google Antigravity | Legacy monorepo map | 5 | ≈ $8.00 |
| GitHub Copilot | 12-package dependency upgrade | 15 | ≈ $6.40 |
| OpenCode | Full API documentation | 7 | ≈ $1.85 |
| Pi | Four-language localisation | 5 | ≈ $1.20 |
| DeepSeek (inside Claude Code) | 200-file type-hint migration | 5 | ≈ $1.70 off-peak |
| Claude Code + Ollama (64 GB Mac) | Confidential refactor, fully local | 4 | Well under $0.10 |
| OpenCode + Ollama (64 GB Mac) | Two-model security audit, fully local | 5 | A few cents |
The Method In Five Points
After more than 500 articles on emerging technology, I strongly believe that orchestration is the most important developer skill of 2026. It matters more than prompting tricks and far more than vibe coding, because it turns one developer into the conductor of a small, tireless team.
By God’s grace, that team no longer has to live entirely in someone else’s data centre. A single 64 GB machine on your own desk can now host a lead agent, a pool of workers, and a reviewer, and the cloud can fill in whenever you need more power.
So plan your work, delegate it precisely, verify the results, merge with confidence, and repeat the cycle every day. Then run /usage—or ollama ps—and check what it cost, because a conductor who ignores the meter does not keep the orchestra for long.
All the very best to you.
And if you are choosing a skill to master this year – choose orchestration.
Cheers!
Appendix. Higher Memory Tiers And Why They Matter
Sixty-four gigabytes is where serious local orchestration begins, but it is not where it ends. Every step up in unified memory lets you run a bigger model, more agents, longer contexts, or several specialist models at once. This appendix sets out what each tier allows, what it costs in October 2026, how long each one takes to pay for itself, and why waiting is unlikely to make any of it cheaper.
Why Memory Matters For Agent Orchestration
- Bigger and stronger models. The strongest open-weight models are now huge mixture-of-experts systems, such as DeepSeek V4 Flash with 284 billion total parameters and weights of about 160 GB, as Simon Willison noted in his write-up of the DeepSeek V4 release.
- Frontier-class open weights. Z.ai’s GLM-5.3 has 744 billion total parameters with 40 billion active, and its weights were published on Hugging Face on 29 August 2026, according to AI Understanding’s report on the GLM-5.3 release.
- More agents in parallel. Every parallel agent keeps its own context cache, so the number of agents you can run at once rises directly with memory.
- Several different models at once. With more memory, a lead model, a fast worker model, and an independent reviewer model can all stay loaded together, which is the full lead-workers-verifier pattern from Section XI.
- Speed through bandwidth. The M5 Pro Mac mini offers 307 GB/s, the M5 Max reaches 614 GB/s, and the M5 Ultra reaches 1.2 TB/s, according to FelloAI’s M5 Ultra Mac Studio overview, so higher tiers also generate tokens faster.
What Each Memory Tier Allows
The GPU budgets below use macOS’s default of roughly 75% of unified memory, which you can raise as Section XVI explained. Model sizes come from Ollama’s library pages and the sources linked above, and the 256 GB and 512 GB fits are approximate.
| Tier | Cheapest Mac In October 2026 | Default GPU Budget | What Fits For Agentic Coding | What It Means For Orchestration |
|---|---|---|---|---|
| 64 GB | Mac mini M5 Pro, about $2,699 | About 48 GB | Qwen3.8-27B at 8-bit (30 GB), or two 4-bit models together (about 42 GB) | A real local team: one strong lead plus fast workers |
| 96 GB | Mac Studio M5 Ultra, $5,499 | About 72 GB | gpt-oss-120b (65 GB), or three or four mid-size models side by side | A 120B-class model, or a full team of specialist models |
| 128 GB | Mac Studio M5 Max, about $5,099 | About 96 GB | gpt-oss-120b with long contexts, or Qwen3.5-122B-A10B (81 GB) | A 120B-class lead with room for small workers and long contexts |
| 256 GB | Mac Studio M5 Ultra, about $10,800 | About 192 GB | DeepSeek V4 Flash at its published size of about 160 GB, or GLM-5.3 at 2-bit | A frontier-class open model running entirely on your desk |
| 512 GB | Mac Studio M5 Ultra, not yet priced | About 384 GB | GLM-5.3 at 4-bit (about 372 GB of weights, with a raised limit), or DeepSeek V4 Flash with room for many agents and long contexts | A complete local stack; DeepSeek V4 Pro (about 865 GB) still does not fit |
Two details in this table matter for buyers. The 128 GB M5 Max costs less than the 96 GB M5 Ultra, so it is the better buy unless you need the Ultra’s bandwidth, and the M5 Ultra has no 128 GB option at all, because it jumps from 96 GB straight to 256 GB in FelloAI’s list of Ultra memory options. The GLM-5.3 fits rely on Unsloth’s 2-bit build, which AI Understanding reports runs on a 256 GB Mac, and on the 4-bit size estimate in Spheron’s GLM-5.3 deployment guide.
Current Mac Mini Prices: M3 To M6
Two generations in this range never existed as Mac minis. Apple released no M3 Mac mini, moving from M2 straight to M4 in 2024, and it released no base M5 Mac mini, launching the M5 Pro and M6 models together, with orders opening on 25 August 2026.
All prices are US prices. Where a configuration is marked “calculated”, I added the upgrade prices reported by the linked source to the base price, because Apple’s own configurator prices could not be read directly—so confirm the final figure in the Apple Store before you buy.
| Generation | Chip (CPU / GPU Cores) | Memory | Storage | Bandwidth | Price | Notes |
|---|---|---|---|---|---|---|
| M4 | M4 (10 / 10) | 16 GB | 256 GB | 120 GB/s | $799 | Was $599 before the June 2026 price rise |
| M4 | M4 (10 / 10) | 16 GB | 512 GB | 120 GB/s | $999 | Was $799 before June 2026 |
| M4 Pro | M4 Pro (12 / 16) | 24 GB | 512 GB | 273 GB/s | $1,599 | Was $1,399; limited supply |
| M4 Pro | M4 Pro (12 / 16) | 48 GB | 512 GB | 273 GB/s | About $1,999 | Calculated |
| M4 Pro | M4 Pro (12 / 16) | 64 GB | 512 GB | 273 GB/s | About $2,199 | Calculated; cheapest 64 GB Mac while stocks last |
| M5 Pro | M5 Pro (15 / 16) | 24 GB | 512 GB | 307 GB/s | $1,699 | Base model |
| M5 Pro | M5 Pro (15 / 16) | 48 GB | 512 GB | 307 GB/s | About $2,299 | Calculated |
| M5 Pro | M5 Pro (15 / 16) | 64 GB | 512 GB | 307 GB/s | About $2,699 | Calculated; meets the 64 GB minimum |
| M5 Pro | M5 Pro (18 / 20) | 64 GB | 512 GB | 307 GB/s | $2,899 | Faster chip |
| M5 Pro | M5 Pro (18 / 20) | 64 GB | 1 TB | 307 GB/s | About $3,199 | Calculated; Apple’s ready-made 64 GB model |
| M6 | M6 (12 / 12) | 16 GB | 256 GB | 170 GB/s | $899 | Base model |
| M6 | M6 (12 / 12) | 16 GB | 512 GB | 170 GB/s | $1,199 | As reported by Macworld |
| M6 | M6 (12 / 12) | 24 GB | 256 GB | 170 GB/s | About $1,099 | Calculated |
| M6 | M6 (12 / 12) | 32 GB | 256 GB | 170 GB/s | About $1,299 | Calculated; 32 GB is the M6 maximum |
The M4 prices come from Macworld’s 2026 Mac mini guide, the M4 Pro and M5 Pro prices and upgrade costs come from AppleInsider’s comparison linked in Section XVI, and the M6 upgrade costs come from AiCybr’s Mac mini M6 and M5 Pro price guide.
Current Mac Studio Prices
| Chip (CPU / GPU Cores) | Memory | Storage | Bandwidth | Price | Notes |
|---|---|---|---|---|---|
| M5 Max (18 / 32) | 36 GB | 512 GB | 460 GB/s | $2,499 | Base model; memory fixed at 36 GB |
| M5 Max (18 / 40) | 48 GB | 512 GB | 614 GB/s | About $3,099 | Calculated |
| M5 Max (18 / 40) | 64 GB | 512 GB | 614 GB/s | About $3,499 | Calculated |
| M5 Max (18 / 40) | 128 GB | 512 GB | 614 GB/s | About $5,099 | Calculated; about $5,400 with 1 TB |
| M5 Ultra (30 / 64) | 96 GB | 1 TB | 1.2 TB/s | $5,499 | Base model |
| M5 Ultra (36 / 80) | 256 GB | 1 TB | 1.2 TB/s | About $10,800 | Calculated; 256 GB needs the top chip |
| M5 Ultra (36 / 80) | 512 GB | — | 1.2 TB/s | Not yet priced | Due late October 2026 |
The M5 Max upgrade prices come from Macworld’s M5 Max Mac Studio review, and the M5 Ultra figures come from FelloAI’s overview linked above, which reports that moving from 96 GB to 256 GB adds $4,000. For comparison, the previous M4 Max Mac Studio rose from $1,999 to $2,499 in June 2026, the M3 Ultra rose from $3,999 to $5,299, and Apple withdrew the $9,499 512 GB M3 Ultra configuration in March 2026, as Notebookcheck’s report on the M3 Ultra memory changes explains.
Break-Even By Memory Tier
Break-even is simply the hardware price divided by the monthly cloud spend it replaces, minus the electricity it uses. The table below assumes 120 hours of agent work a month at $0.15 per kWh, with roughly 0.1 kW for a Mac mini and 0.2 kW for a Mac Studio, which works out to $1.80 and $3.60 of electricity a month.
The 512 GB price is my own estimate, because Apple has not published it yet. It applies the roughly $25 per gigabyte implied by Apple’s $4,000 step from 96 GB to 256 GB, which puts the 512 GB model near $17,200, so treat that row with low confidence.
| Tier And Machine | Price | Replacing $200 A Month | Replacing $500 A Month | Replacing $1,000 A Month |
|---|---|---|---|---|
| 64 GB Mac mini (M5 Pro) | About $2,699 | 13.6 months | 5.4 months | 2.7 months |
| 64 GB Mac Studio (M5 Max) | About $3,499 | 17.8 months | 7.0 months | 3.5 months |
| 96 GB Mac Studio (M5 Ultra) | $5,499 | 28.0 months | 11.1 months | 5.5 months |
| 128 GB Mac Studio (M5 Max) | About $5,099 | 26.0 months | 10.3 months | 5.1 months |
| 256 GB Mac Studio (M5 Ultra) | About $10,800 | 55.0 months | 21.8 months | 10.8 months |
| 512 GB Mac Studio (M5 Ultra) | About $17,200 (estimate) | 87.6 months | 34.6 months | 17.3 months |
The pattern is clear. At $200 a month, which is the price of a top individual subscription, only the 64 GB tier pays for itself in a little over a year. The 256 GB and 512 GB tiers make financial sense when they replace a team’s API spend of $1,000 a month or more, or when privacy and regulation rule the cloud out entirely. These figures also ignore resale value, which shortens the real payback period for well-kept Macs.
Why Prices Will Keep Rising
The evidence that hardware with large amounts of memory is getting more expensive, not less, is strong and recent.
- Apple raised prices across the Mac line. On 25 June 2026 Apple raised the M4 Pro Mac mini from $1,399 to $1,599, the M4 Max Mac Studio from $1,999 to $2,499, and the M3 Ultra Mac Studio from $3,999 to $5,299, as MacRumors’ report on Apple’s June price increases shows.
- Apple’s CEO blamed memory directly. In the same report, Tim Cook told The Wall Street Journal that memory and storage costs had risen sharply, describing the shortage as a “hundred-year flood.”
- Memory upgrades doubled in price. The step from 96 GB to 256 GB went from $1,600 to $2,000 on the M3 Ultra in March 2026, and to $4,000 on the new M5 Ultra, while the 64 GB upgrade on the Mac mini rose from $600 on the M4 Pro to $1,000 on the M5 Pro.
- The 512 GB option has been scarce. Apple withdrew the 512 GB M3 Ultra in March 2026, and the 512 GB M5 Ultra still has no published price or ship date beyond “late October.”
- DRAM contract prices are still climbing. TrendForce expects server DRAM contract prices to rise 13–18% quarter on quarter in the third quarter of 2026 and to keep rising every quarter through the second half of 2027, according to its July 2026 DRAM price forecast.
- AI demand is absorbing the supply. TrendForce now expects HBM prices to rise 121% year on year in 2027, and HBM competes with ordinary DRAM for the same limited wafer capacity, as EE Times Asia reported in its coverage of TrendForce’s 2027 HBM forecast.
Now; for the other side of the coin! Memory has always been a boom-and-bust industry, and TrendForce itself expects the pace of increases to moderate through 2027. New fabs and the eventual end of the AI build-out could bring prices down again, so “prices will only ever rise” is not a law of nature.
My verdict is this. I have high confidence that prices stay elevated through 2027, moderate confidence that they keep rising over that period, and low confidence about anything beyond it. For a buyer, the practical lesson is simple: buy the memory tier you actually need now, starting at 64 GB, rather than waiting for a price drop that the evidence does not support.

