← Back to Feed

How to Work With Agents in 2026

C
By Charlie
How to Work With Agents in 2026

The cover: a Friday night in September. A review pass over eight open pull requests came back with concerns on seven. Three fixer agents are working in parallel, each in its own checkout on its own PR. Two more PRs will be fixed but held from merging until the Unity PRs they depend on land. The other three, including the one that passed review, wait for disk space.

An agent fleet will hand you a month of plausible work every week. The developers getting ahead in 2026 are the ones who can tell quickly whether the work that came back is true.

My notes call the failure a false green: a passing result that never tested your code. The worst one I've had: agents working in an isolated checkout resolved absolute paths back to the main checkout, edited files there, and then ran the suite against the untouched copy. Every run passed for days. The first question I ask of any result now is what exactly it ran, and where.

I first published this in August. In September I started running fixers in parallel across a stack of pull requests, with critic agents attacking their work before I see it and hourly loops that pick up tickets on their own. The checks had to move earlier and become automatic. This version is shorter and adds what I learned from that change.

Principles

Verification is the bottleneck. Generating work is cheap now. Most of what follows is about checking it.

Fan out on reading, keep writing single-threaded. Five agents mapping five subsystems finish in the time of the slowest, and their results merge cleanly. Parallel writers make incompatible decisions that you pay for at integration. When I do write in parallel, each agent gets its own checkout and its own pull request, and no two agents touch the same file.

Check the harness before you believe the result. I once built a results table that agreed with my hypothesis on every row. My shell wasn't splitting an unquoted variable, so every variant I thought I was testing ran as one nonsense argument. Be more suspicious of a measurement that agrees with you on the first try.

Use a cold reviewer, and treat its findings as evidence. The agent that wrote the code shares every assumption that produced the bug. A fresh one with only the diff finds things the author can't. Reviewers are also sometimes wrong, so I check their findings before acting on them.

A planner agent producing six plans, a checker agent verifying them and returning three warnings, then a revision pass A planner produced six plans for 322.7k tokens. A separate checker spent another 202.9k finding three problems, and a revision pass fixed them.

Measure every day. A job grades the previous day's agent traffic each morning. It once reported that a broken feature had recovered. It hadn't: a change in how the scanner walked the message files had made old records visible. I caught it because the number stayed at exactly the same value for three days while sixteen new events should have moved it.

The field manual

I learned these across four unrelated codebases: a co-op game, a client's CMS and localization pipeline, this product, and the journal that runs my day. None of it depends on my stack.

1. Give every writer its own checkout, then check the isolation

Git worktrees make a checkout per agent nearly free, but isolation fails quietly. The false green above came from my own instruction files. They were full of absolute paths to the main checkout, and the agents followed them. The fix was a standing instruction to search from . and report relative paths, verified by re-running the task that had failed.

After any edit, run git status in the copy that runs the tests. If your change doesn't show up there, it landed somewhere else. Give each checkout its own dependency install, because a shared node_modules resolves workspace packages back to the main source. When you prune old worktrees, check each one's HEAD against the integration branch (git log origin/main..HEAD). Comparing remote refs tells you the branch merged, but it never looks at what the worktree has locally.

2. Make verification one command that adapts

The game project has a single gate script. If the editor is open, it builds a throwaway sandbox checkout and clones a 1.8 GB import cache in seven seconds with copy-on-write. If the editor is closed, it runs in place. There is only one command to learn, so people run it.

Inside the gate, a few details decide whether a pass means anything. Print a marker from the entry point, so the gate can prove the code compiled and ran. A suite that discovers its test cases should fail when it discovers none. Don't pipe a gate command: typecheck | tail reports tail's exit code, and I've seen a typecheck with seven errors report success that way.

3. Test the artifact that ships

A game's spawn system had thorough tests, all passing, while the shipping scene had two spawn points on the same coordinates, floating off the ground. Every test built its spawn points in code, and none read the scene file players load.

When the question is about a live system, check the live system. When a bug only reproduces after deploy, prove the fix A/B/A: deploy the broken version and see the failure, deploy the fix and see it pass, then deploy the broken version again to rule out coincidence. Then put the fix back.

4. Attack the work while it's still in the worktree

Every build now runs gauntlet rounds before I see it. Fresh critic agents get the ticket's acceptance criteria and try to break the fixer's work. A finding only counts if it comes with a failing input or a broken invariant. The fixer adds a failing test for each real finding, fixes it, and the next round starts.

In September this ran on a terrain importer that reads heightmaps from an external world-building tool. Three rounds found five input-handling bugs, and each became a failing test before it was fixed. A height range of exactly 147.66 m was refused. A big-endian file that only rose to the north was misread. A 24-bit splat texture painted half of every texel into the wrong layer. 8-bit heights stored in 16 bits were read byte-swapped, which produced a flat world with no warning. And a folder of tiles decided its byte order by looking at one tile.

Three of those five came from byte-order auto-detection, and each round's fix produced a new failing input for the next. My proposed fix is to read the format the tool documents by default and make detection opt-in. The last round's fix hasn't been re-checked yet. When a heuristic keeps failing on new inputs, stop improving it and make it opt-in.

5. Review every open PR in parallel

I run review passes with separate jobs. One looks at quality only: reuse, simplification, and whether each piece sits at the right layer. Another checks correctness only, using a model from a different company. Their findings rarely overlap, and both find real problems.

In September I started running this across every open pull request at once, with three reviewers per PR: those two passes plus a general reviewer. On eight open PRs it returned concerns on seven. One PR had a redaction filter that was case-insensitive, so it mangled ordinary code like apiKey: string in command output and still leaked the end of long quoted secrets. Another let any caller list and revoke a different user's CLI clients by passing that user's name as a parameter, and it didn't compile.

Then I check the reviewers' findings against the code, and re-run the reviewer on my fix before pushing. Merging is the one step I can't undo, so it always waits for my explicit yes, even when every check is green.

6. Put a preflight in every loop

An hourly loop picks up the next ready ticket on the game project. Before doing anything, it checks free disk against a 15 GB floor, looks for a lock written by another session, and confirms the checkout is clean. One night it stopped at 6.2 GB free. An hour earlier it had 15 GB, and about 9 GB went in that hour, around the time a large asset package was imported. The loop did no work and deleted nothing. A lock from another working session was also in place, and the loop left it alone. Later ticks skipped on the lock or on disk until that session ended.

7. Watch what your agents do to your infrastructure

Our team's file hub went down in September. Every API route returned the static dashboard, because the hosting plan's free tier had capped out after 166,752 function requests in 24 hours. I moved the hub to a paid plan to bring it back. My best explanation, not yet measured: the hub's MCP endpoint answered a streaming GET with a normal 200 JSON response, and MCP clients seem to treat that as a dropped stream and reconnect in a loop. I changed that GET to return 405, which tells the client there is no stream to open. I'll know it worked when the next day's request count drops.

8. Default destructive tools to dry-run, and say which "done" you mean

A translation migration I wrote ran dry by default. The dry runs caught a bug that would have overwritten weeks of human corrections with old machine translations, and none of it reached production because the tool never wrote anything until I told it to.

A fix was once announced as done eight minutes before the fixing commit existed. Merged, released, deployed, and reprocessed are four different claims. Say which one you mean and when.

9. Curate memory

A rule only moves into my always-loaded instructions after it has come up more than once, and it carries the date, the incident, and what the person actually said. Agents write provisional notes to a staging file, and a periodic audit checks them against the code before promoting or deleting them.

10. Give the human decisions a real interface

I still make the merge and scope decisions. When I sort the day's screenshots into my journal, the agent builds a small page: one card per image with a preview, a file/maybe/cut toggle, the proposed filename and caption, and a note field. I make my choices and paste the result back. My notes on that page carried instructions the agent couldn't have guessed: upload this one to the team hub, turn this cut into a design reference, consider this one as a blog cover. The image at the top of this post came from that page.

I also label claims by evidence level (Opinion, Signal, Evidence, Validated, Measured) before acting on them. The most useful catch so far was "there is no capacity problem", written as fact. It was an opinion, and a measurement I already had contradicted it.

What I run

A gateway runs a handful of long-lived agents, each with its own memory and tools, and they remember last week and each other. Scheduled jobs work overnight and start specialist agents in isolated checkouts to open pull requests. A pipeline command takes a ticket from wherever it stands through build, review, merge, and post-deploy checks. It stops at every gate a human owns.

The morning brief on a phone: one decision to make, three blockers, the night shift's results, and eight blueprints on deck What I see on my phone in the morning: the one decision that needs me, three things that are stuck and why, what shipped overnight, and who owns what's in flight.

The bet

Everything above works in any harness. I'm building Vibery so these practices come built in. Here is the bet behind it and how much I know.

Run enough agents and you become the channel between them: carrying context, remembering what the overnight job decided, noticing two of them starting the same work. That gets worse with every agent you add. Vibery is local-first, so your agents, keys, and data stay on your machine.

I'm betting a game layer can work as a shared model of the work. The game is a twin: a lossy simulation of the real engineering system. I read the game state to decide what to do next, and the agents read the same state to decide what they do next. I can't keep forty concurrent agents in my head, but I can keep "engineering is busy, ops has a fire."

The station rendering behind the Log panels, with the code red styling applied to the UI only August: the code red turns the panels red while the 3D station behind them carries on as normal.

When I audited this in August, it was mostly not wired. Agent presence reached the 3D world, while pull requests, XP, and rank did not, and agents had a tool that read their own progression with nothing telling them to use it. I haven't re-run that audit since, so treat it as the last measured state.

The hypothesis: wrapping a multi-agent system in a game layer makes managing many agents easier, more enjoyable, and measurably more effective than a terminal or a purpose-built harness. I'm watching how fast I can answer "what changed overnight, who's blocked, what needs me" as the fleet grows, how many agents I can run before I become the bottleneck, and whether agents that treat game state as an objective do better than the same fleet without it. The cheapest test is to turn the game layer off and see what changes. The real competition is a good terminal.

How much do I know? That verification is the bottleneck is Measured: my own numbers keep showing it. That the twin makes agents easier to manage at scale is Opinion, and it stays untested until the loop is connected.

The reading list

Anthropic's Building Effective Agents is still the clearest explanation of when to use a workflow instead of an agent. Effective Context Engineering for AI Agents led to an audit that found one of my agents 24% over its own system-prompt budget. Your agent changed under the model makes the case that a model id is not a version, which is the best argument I know for running your own evals.


Published 2026-08-02. Updated 2026-08-05 with the field manual and 2026-08-06 with a full restructure. Rewritten 2026-09-25 for September's practice: gauntlet rounds, parallel review, loop preflights, and agent-driven infrastructure load.