I Built a Multi-Agent Coding Setup to Keep Me in the Loop

I Built a Multi-Agent Coding Setup to Keep Me in the Loop

# ai# productivity# programming# opensource
I Built a Multi-Agent Coding Setup to Keep Me in the LoopJancer Lima

Why I built this I didn't build this after a big disaster. No agent ever deleted my...

Why I built this

I didn't build this after a big disaster. No agent ever deleted my database or broke a production deploy. The problem was slower and more annoying than that.

Every time I used an AI agent on a real task, I spent too much time managing it. It asked permission for simple things, the implementation often failed when the task was bigger, and even when the code worked, it often ignored the way the project was organized.

Each of these was small, but after a while they made working with AI more tiring than it should be. So I started adding rules, one problem at a time, and the result is Jancera/skills, a set of skills and subagents I use in Google Antigravity.

Built to keep me in the process

Most of the talk about coding agents is about taking the human out of the process. I wanted a setup where I stay in it, but without having to watch every step.

Every rule in the repo comes from that. Agents only ask me about decisions that matter, the work is split into pieces small enough for me to review, and nothing reaches the main branch unless I merge it myself.

Three frustrations that became rules

1. Too many permission prompts

The agent asked me to approve almost everything, and most of it was commands that only read files. I never cared about approving reads. What I wanted to approve were the things that change something.

Now read-only inspection commands are approved by default. Implementation agents run inside an OS-level sandbox, where commands run without prompts and network access is off unless a task needs it. I only get asked when something goes beyond that, like a commit, a git push, installing a package, downloading a file, or editing outside the project.

2. Big tasks, bad implementations

When I gave the agent a large feature, the implementation often failed, and I believe the main cause was poor planning. Matt Pocock's talk on breaking requirements into vertical slices changed how I split the work.

The deepwork skill breaks a feature into phases, and each phase into vertical tasks. A task touches the whole slice it needs (schema, backend endpoint, UI) and produces something I can run and test. Each task runs in its own git worktree from the same base branch, and tasks are never stacked on top of unmerged work.

3. Agents never merge

When a task finishes, I check out the branch, test it, and merge it myself.

After every task in a phase is merged, an oracle subagent reviews the combined diff of the phase for bugs, security issues and regressions. Accepted findings go to a fixer agent for mechanical fixes or to a designer agent for UI work.

A real run: four tasks in parallel

I tested deepwork on shorts-pipeline, a side project I work on now and then. I started four tasks at once:

  • Subtitle settings in the web UI
  • A target number of shorts per project
  • Regenerating the audio of a single video
  • A configurable desired video length

All four branched from the same commit, and each one changed config, the Flask backend, the HTML frontend and the tests within the same task. Together they added about 1,360 lines. The subtitle task alone added 1,068, almost half of them tests.

I started the run and went to do other things, and in about a day all four were ready for review. I didn't reject any of them, and they landed on main.

Then the oracle reviewed the merged phase and found problems that no single task had. The shorts-count task and the video-length task both touched the same settings flow, and each one worked alone. After the merge, the settings route was missing validation for the new length field, the stale-state check ignored changes to the shorts count, and some test mocks still used the old function signature.

I didn't catch any of this in my own review, and the fix landed shortly after the last merge. That's the reason the gate looks at the combined diff of the whole phase. When tasks run in parallel, some bugs only exist after everything is merged, and no per-task review will find them.

Where the time savings come from

Without this setup, my guess is that these four tasks would have taken me several days, though I didn't measure it. It also doesn't mean the code is as good as what I would write myself, which I'll come back to below.

Most of the gain came from where my attention went. I was only needed at the gates, and I could walk away in between because nothing reaches main without me and the oracle checks what I miss once the tasks come together.

What still doesn't work

Project conventions. The agents tend to ignore the structure that already exists in a project, like its layers and where things belong. The code works, but it doesn't follow the patterns I defined or agree with. The repo has an explorer agent that checks module boundaries after a phase, but it's optional and didn't run in the test above, so that's the first thing I'll try.

The lighter workflow. to-spec and implement-spec are meant to be a simpler version of deepwork for medium tasks, and they still need work. My hypothesis is that the planning is too shallow, since implement-spec doesn't split the work into vertical tasks the way deepwork does.

The sandbox. I use Antigravity's default sandbox exactly as it comes. I haven't looked into the other sandbox options, or checked whether the default is enough for this workflow, and for now I'm leaving it that way.

Credits

Most of this is adapted from other people's work. The grilling skill comes from Matt Pocock's grill-me idea, which I picked up from his talk. deepwork started from the structure Fabio Akita uses with OpenCode, which I rebuilt on Google Antigravity's own pieces: one worktree per task, OS sandboxing, and subagents with their own model tiers.

The repo is public at github.com/Jancera/skills. If you try it, I'd like to hear where it breaks.