Billie MWhy I want AI coding checks to run independently, with fallible model reviews, human-owned criteria and room for different workflows.
I have seen a Claude session review a change that another Claude session has already reviewed. A human still has to take responsibility for it. That is uncomfortable when you do not understand the change you are being asked to approve.
In my experience, review has become more hurried as the amount of software we are expected to produce has increased. It has not disappeared, but the depth has struggled to keep up. “Hey, we've got to move. Can you take this?” is a different situation from having time to properly understand the work.
I want checks that run independently of the agent writing the code. But that needs a distinction which is easy to lose: requiring a review to happen does not make its verdict reliable.
Some acceptance conditions can be evaluated mechanically. Others require judgement about the change and the codebase. I want both kinds of checks to happen without depending on the coding agent remembering to run them.
That makes the execution of a required model review more dependable. It does not make the model's conclusions deterministic. The reviewer can misunderstand a requirement, miss a problem, or respond differently after a model update. Another model saying the work looks good does not establish that it is good.
This is why merely adding an automatic reviewer to a pull request does not answer the questions I care about. What is it checking? How closely are its criteria grounded in this particular codebase? What confidence should I put in its judgement?
I am not trying to settle whether the increased pace of development is good or bad. I am encountering it, and I want the acceptance process to account for it. Asking everyone to review as thoroughly as before does not create the time to do that.
A recent, much smaller annoyance brought this into focus. Claude was adding far more code comments than I wanted. I asked it to use fewer, and added the guidance to CLAUDE.md. The code still had too many comments.
That experience did not establish a new comments gate or a solution to excessive comments. It highlighted how weak an instruction can be as a mechanism for controlling output, even when the request seems simple.
Anthropic describes CLAUDE.md as context rather than enforced configuration. I find that a useful way to think about it. The file can explain what matters and help steer the model. Writing “always” inside it does not make the action inevitable.
I have used AI for development since the early Copilot days, then early ChatGPT, followed by Claude Code and a lot of Codex in my own time. Trying different workflows has changed my opinions repeatedly. At this point I put more confidence in checking the completed work than in building an elaborate instruction sequence and hoping the model follows it.
If a requirement matters enough to prevent accepting a change, its check should run whether or not the coding agent remembers it.
That still leaves plenty of useful work for instructions to do.
While talking through the original article, I initially dismissed plugins as silly. That was too broad. I use skills and workflows myself, and plugins can contain tools and hooks as well as instructions.
There are absolute gold mines of human knowledge worth encoding: understanding the problem, knowing its constraints, and recognising what matters through experience. A model can help express that knowledge in a reusable document. It can save me repeating the important context every time.
What I question is generating a large workflow document that contains nothing another person could produce by asking the same model. Producing text is cheap. The knowledge behind it is where I see the value, and the size of the document gives me little reason to trust it.
For a team, I would want qualified people to protect the codebase-specific acceptance criteria. More than one person should discuss changes and understand why each requirement belongs there. Almost a council, although that sounds grander than I mean.
I recognise the habit of accepting a convincing model explanation as authority: “Oh, Claude said the thing. Now this is how the thing is done.” I want friction before that explanation becomes another rule.
Without that discussion, criteria can grow endlessly. Every plausible suggestion gets added, and the review ends up carrying requirements nobody has properly defended. Keeping the set bounded is part of maintaining it.
Independently run checks are increasingly part of my own practice, and for now they seem to be working better for me and giving me more confidence. The group protecting a shared review rubric is how I would like the idea to work in a team. It is a proposal, not an arrangement I am claiming to have solved.
That proposal leaves room for developers to work differently. Someone might prefer Claude, someone else Codex. A detailed workflow can suit one person while a long conversation works better for another. Agreeing on the requirements and checking the result gives us less reason to prescribe every step.
The diagram below separates those responsibilities. You can also view the native HTML diagram on billiem.uk, with readable text and a separate branch showing the proposed human stewardship of model-review criteria.
A static version of the proposed boundary. The checks run independently of the coding agent; model-review judgments remain fallible. Human stewardship of the criteria is a team proposal.
My own preferred interface is voice-to-text. This article grew out of a conversation in which I dictated my thoughts, kept elaborating, and asked the model to interrogate me. Its questions helped me express more of what I meant, including correcting my blanket criticism of plugins.
I read the replies rather than listening to the model talk back. I want to inspect what it has understood and decide what to correct. With enough context, the model can organise the material, prepare tasks for other sessions or delegate to agents.
You should be using voice-to-text. I use Wispr Flow. That is my enthusiastic recommendation, alongside my belief that people and tasks differ; talking is not the right interface for everything.
When I say prompting matters less now, I mean the elaborate ritual around wording and structure. I still need to give the model substantive information and a clear direction. Being able to ramble, answer questions and refine that direction is useful to me. The model can help with organising it.
I do not think anyone has found the optimal AI workflow. New models change what works. Instructions that helped one model may be unnecessary for the next, get in its way, or never have been followed consistently in the first place.
When a new model arrives, I ask it to audit my existing instructions and Markdown files. I expect to keep doing that. I would rather leave room for that experimentation and spend the shared effort on defending the requirements against which the work is accepted. Model-based checks will still need attention too. Changing the way I generate the code should not mean relying on a new set of promises that the agent will remember to check it.
Want to talk about something I’ve written or built? Get in touch.
This article was adapted with AI assistance from an original article on billiem.uk. The original article was reviewed before publication.