Maxxwell by Rindler
Writing

Shell scripts vs a manager for agent fleets

2026-09-28

Past three or four coding-agent sessions, your problem stops being model quality and starts being you. You’re answering prompts, hopping tmux panes, and.


Past three or four coding-agent sessions, your problem stops being model quality and starts being you. You’re answering prompts, hopping tmux panes, and guessing which agent is stuck or quietly off-mission.

This is where most people either grow a pile of “Herd X” shell/Python glue, or adopt a manager like Maxxwell. The core trade-off is simple: scripts buy raw flexibility; a manager buys a control plane.

What matters when you manage a herd of agents

Before comparing Herd X scripts to Maxxwell, it’s worth being explicit about the constraints.

Anthropic’s 2026 trends report is blunt: developers use AI in ~60% of their work, but can fully delegate only 0-20% of tasks. In ~400,000 Claude Code sessions from ~235,000 people, humans made most planning decisions; agents handled execution.

So the questions aren’t “can my agents write code?” but:

A good manager layer solves for:

Herd X scripts: power and maintenance in one bundle

Herd X scripts are the tmux layouts, shell loops, and Python glue you already have:

You get three big advantages.

What Herd X scripts do well

  1. Total flexibility Your scripts embody your workflow:
    • Custom prompts and bootstraps per agent
    • Tight coupling to your repos, build system, and test harness
    • Odd cases: “run agent 3 with a tiny context and weird environment vars”
  1. No new dependency Everything runs:
    • In your terminal
    • Under your process supervisor
    • With your own logging and access patterns
  1. Infinite hackability If you want to wire agents into a homegrown queue or CI step, you can:
    • Add a cron that kills stale sessions
    • Pipe outputs into a review script
    • Invent your own priority scheme

For 1-3 agents, this is enough. Your shell is the orchestration layer.

Where Herd X scripts break down past 5-10 sessions

The failure modes show up as soon as you treat the herd as a system instead of a set of processes.

From multi-agent research (1,600+ traces across 7 frameworks, 14 failure modes) and telemetry work on reasoning loops, the common problems cluster around:

In practice, your Herd X scripts usually fall short on:

  1. Observability You get ps, tmux status lines, and some log files. You don’t get:
    • “this lane is waiting on you” vs “this lane is just slow”
    • “this session hasn’t produced output in 5 minutes” in human terms
    • A clear state model: working, idle, blocked, done
  1. Attention routing When you have eight agents, your terminal is noise:
    • Every pane is scrolling
    • You alt-tab into whatever last screamed at you
    • You answer low-value questions while high-value work waits
  1. Failure diagnosis When something goes wrong, these questions are hard to answer quickly:
    • Did this session stall, or is it on a long-running test?
    • Did it drift, or did my brief suck?
    • Which sessions are safe to kill vs need a human check?
  1. Scaling the workflow to a team Homegrown scripts scale poorly:
    • Every engineer’s “herd” scripts diverge
    • Onboarding means explaining opaque tmux rituals
    • Debugging orchestration bugs becomes a mini-platform project

If you’re already at five or more agents at once, you will eventually end up writing:

At that point you’ve reimplemented half of a manager.

Maxxwell: a manager for the agents you already run

Maxxwell treats this as a control-plane problem. It doesn’t try to be another coding agent. It manages the ones you already run: Claude Code, Codex, Cursor agents, or your own.

Mechanically:

No sign-up, no hosted control plane for normal local use.

Fleet visibility: every session in one window

Maxxwell puts each agent session in a lane with a readable state:

You see, at a glance:

The system follows the brand’s point of view: say what you can’t verify. It uses “not heard from” instead of guessing “working”. A dashboard that is confidently wrong is worse than one that admits a gap.

Orchestrator seat: one brief, one conversation

On top of the lanes sits an orchestrator seat:

The orchestrator:

Anthropic’s studies show humans still own planning, agents own execution. The orchestrator is built around that fact: it raises decisions instead of hiding them.

Drafts rather than acts: you still press enter

One distinctive property: fleet controls draft rather than act.

Any control that would change the herd:

Nothing auto-runs. You are the one who presses enter.

This matters if you care about:

Maxxwell is conducting, not autopilot. It does not claim to:

It gives you:

Local-first: the herd survives your window

Sessions in Maxxwell:

Closing your laptop is not silently throwing away an hour of work.

Herd X vs Maxxwell: side-by-side comparison

Here’s a direct comparison for typical criteria.

CriterionHerd X scriptsMaxxwell
SetupWrite and maintain your own shell/Python scripts; wire into tmux/zellijInstall local desktop app or CLI; point it at your existing agents
AgentsWhatever you script: Claude, Codex, Cursor, customUnmodified coding agents in real terminal sessions; your tools stay yours
VisibilityProcess list, tmux panes, logsSingle window of lanes with readable state (working, idle, waiting on you, blocked, etc.) plus “possibly stalled” overlay
Attention routingManual: whichever pane you last looked atOrchestrator seat summarizes work, calls out what landed vs what needs decisions
Control semanticsWhatever you encoded in scripts; can be fully automatedFleet controls draft messages rather than act; you press enter for changes
Observability depthDepends on your logging; usually low on reasoning/planning spansSession state plus live context-pressure readout; explicit “not heard from” when unsure
Failure mode handlingYou debug scripts and agents manuallyManager surfaces stuck/blocked states; you still decide how to fix them
PersistenceDepends on how you supervise processesSessions outlive the app; quitting detaches instead of killing
Team adoptionEach engineer maintains their own variant; fragile standardsShared manager and shared vocabulary for session state
CostYour time to build and keep scripts runningLocal use free for individuals; paid tiers for teams

For a deeper look at workflow patterns, see the related guide: herding coding agents without adding more noise.

Scaling beyond five or ten sessions

The scaling break-point is not high. Anthropic reports Claude Code users spend an average of ~20 hours per week in agentic workflows, and the share of GitHub projects with coding-agent activity more than doubled since late 2025.

Once multiple people are running 5+ parallel sessions, the coordination cost becomes visible.

How Herd X scripts scale

With scripts, scaling looks like:

Real-world symptoms:

The failure taxonomy work (MAST, AgentTelemetry) shows that multi-agent systems fail mostly in reasoning loops and misaligned delegation. Ad-hoc scripts rarely model those explicitly; they treat everything as “a process with logs”.

How Maxxwell scales

With Maxxwell, scaling looks like:

You get:

You still do the thinking. The manager just compresses the overhead.

When to stick with scripts vs when to adopt Maxxwell

Stay with Herd X scripts if

Adopt Maxxwell if

Most serious users start with scripts, then add a manager when the overhead starts to cost real hours.

FAQ: Herd X scripts vs Maxxwell

Does Maxxwell replace my existing coding agents?

No. Maxxwell manages the agents you already run.

Workers are unmodified Claude Code, Codex, Cursor agents, or your own tools, running as real terminal sessions you can attach to and take over at any moment.

Can I keep my custom scripts and still use Maxxwell?

Yes.

You can:

Maxxwell is a control plane, not a replacement for your shell.

How does Maxxwell help detect stuck or mis-aimed agents?

Maxxwell marks each session with a state like working, idle, blocked, waiting on you, or not heard from, with a “possibly stalled” overlay when something looks off.

It does not silently assume “working” when it has no signal. You still decide whether a session is truly stuck or just slow, but you don’t have to infer that from logs alone.

Does Maxxwell automate agent decisions?

No.

Fleet controls draft changes instead of executing them:

There is no background autopilot that re-aims sessions, restarts work, or runs goal checks on its own schedule.

What does Maxxwell cost and how is it deployed?

For individuals:

For teams:

You bring your own model access: an API key or an existing Claude or ChatGPT/Codex subscription.


If you’re at the point where most of your day is spent answering agents rather than coding, it’s worth treating the herd as a system and giving yourself a manager.