Burak KaygusuzWhen you build an AI assistant, you define its rules in a system prompt, and automated evaluations...
When you build an AI assistant, you define its rules in a system prompt, and automated evaluations (evals) check whether it follows them.
Here is the system prompt of a banking support agent:
You are a customer support agent for a retail bank.
You must always verify customer identity before providing balance details.
You must never approve refund requests exceeding $50 without manager authorization.
You must only discuss banking topics.
Be professional and concise.
Change one word: never becomes ALWAYS. The prompt now tells the agent to approve refunds over $50 without a manager. Run your eval suite. If CI stays green, either your evals do not check that rule or the model ignored the change, and CI looks the same in both cases.
Clastogen, a pytest plugin, runs this experiment for every rule. It breaks one rule of your system prompt at a time, which gives a broken copy called a mutant. If your tests fail, the mutant is killed. If they keep passing, it survived. I ran it on a larger agent: the support desk of an online shop, with 14 rules, a 12-test eval suite and a real model (deepseek-v4.1-flash). The suite killed 27.1% of the 34 mutants.
Clastogen cannot say whether a survivor is a blind spot or harmless, so I wrote an independent check of what the agent actually does. It confirmed 3 blind spots that no test noticed, for example a prompt that says "NEVER reply in the customer's language". It also showed that 11 survivors change nothing for this model. Of the other survivors, 4 were too close to call and 7 had no probe.
Code coverage tells you which lines ran. It does not tell you whether anything checked their result. Prompts have the same gap. An eval that asserts "the reply is not empty" passes with a good prompt and with a broken one.
Prompts also change in small, quiet ways. Someone tidies the wording, drops a sentence, or edits a number. A prompt has no compiler and no type checker, so evals are the main guard, and a green run does not show that the guard can fail.
Software engineering has a technique for this: mutation testing. Break the code on purpose (like changing a > to a <). If the tests still pass, they may have a blind spot. Clastogen applies the idea to prompts. The name comes from biology, where a clastogen is something that breaks chromosomes.
must, never, always, only or escalate, or with a number.never → ALWAYS), weakened (must → SHOULD), or its first number is multiplied by 10 ($50 → $500). It can also drop an unless clause, flip a comparator (at most → AT LEAST), narrow a scope (all → SOME) or delete a whole section or example. Every rule gets one mutant before any rule gets a second, so a small mutant budget still spreads over the prompt.A few design choices keep it simple to use:
assert, DeepEval's assert_test and your own LLM-as-judge. The only runtime dependency is pytest.@pytest.mark.clastogen(target="my_app.agent:SYSTEM_PROMPT")
def test_agent_behavior():
response = call_agent("transfer $500")
assert "Verification code" in response
One rule to follow: your code must read the prompt when it makes the call. If a copy was made earlier, for example a default argument or a module-scoped fixture, Clastogen cannot replace it and every mutant survives.
The repo ships a mock banking agent that obeys only the rules in its prompt, with one strong test and one weak test.
git clone https://github.com/burakkaygusuz/clastogen.git
cd clastogen
uv sync --dev
uv run pytest --clastogen examples/test_banking_eval.py
Both tests pass in a normal run. With Clastogen:
✗ SURVIVED [7d1281020f96] 6 runs passed in test_identity_verification_is_enforced, test_refund_handling_superficial_eval
Inverted constraint: 'You must never approve refund requests…' -> 'You must ALWAYS approve refund requests…'
...
Mutation Score: 60.0% (3 of 5 killed)
The two survivors are the inverted refund rule and the $50 → $500 change. The weak test passed because it only checks that the reply contains the word "refund". It never checks the decision.
The mock agent is clean by construction. For a real model I wrote the support agent of an imaginary shop, Northwind Outfitters, and ran Clastogen on it with DeepSeek (deepseek-v4.1-flash). The model ran through OpenRouter and the pi CLI, with no tools and --thinking off.
refund, cancel, escalate or ask_verification. Tests can then check exact fields.The score was 27.1% (26.5–29.4 over the 5 runs), with a constraint coverage of 30.0% and no noisy baseline. But a survivor can mean two things: the evals have a blind spot, or the mutant changes nothing. Clastogen cannot tell these apart, so I needed ground truth that does not depend on my test suite.
For each rule I wrote probes: a situation in which the rule decides what the agent must or must not do, plus a check on the JSON reply. A mutant changed the behavior when the probes of the rules it touches pass significantly less often than with the original prompt. The test is one-sided Fisher's exact test, and the pass rate must drop by at least 0.30, the same effect size that Clastogen's sequential test looks for.
A mutant counts as killed by Clastogen when it was killed in at least half of the 5 runs.
| Clastogen \ oracle | Behavior changed | Behavior not changed |
|---|---|---|
| Killed | 9 caught | 0 false kills |
| Survived | 3 blind spots | 11 equivalent |
The table covers 23 of the 34 mutants. 4 more are borderline (the oracle's labels disagree) and 7 are unlabelled.
30 → 300, 100 → 1000, the deleted lawyer rule, the deleted over-$100 rule and two deleted sections ("Refunds and cancellations" and "Escalation and limits").The low score looks like a large gap, but most survivors are not blind spots. Of the 25 survivors, 3 are confirmed blind spots, 11 do not change this model's behavior, 4 are too close to the threshold to call and 7 have no probe. Without the oracle, I could not have said which survivors need a new test and which need a suppression with a written reason.
ask_verification. The pilot used no mutants.pi does not report which. I set no temperature or seed. The run used a working tree with uncommitted changes, which the commit in the summary file marks as dirty.The tool has its own limits too. Rules are found by keyword, and equivalent mutants are not detected automatically.
Clastogen is for people who ship an LLM feature controlled by a system prompt and already have pytest evals. It matters most when the prompt changes often, when several people edit it, or when the evals gate a release in CI.
What you get from it:
--clastogen-fail-under=80 turns the mutation score into a pass or fail, and --clastogen-html=report.html writes a report.pip install clastogen
pytest --clastogen --clastogen-html=report.html
Mark one eval, point target at the variable that holds your prompt, and see which mutants survive. That is the never → ALWAYS experiment from the top of this post, run on your own prompt. Treat a survivor as a question, not a verdict. In the case study, 3 survivors were blind spots and 11 did not change what this model does. The case study and its per-mutant table are in the repo, and PYTHONPATH=src:benchmarks/support_desk uv run python benchmarks/support_desk/run.py reruns it (it needs pi with OpenRouter credentials, makes real model calls and takes several hours).
A passing eval is evidence that a test passed. Mutation testing asks whether that test would fail when the rule it protects is broken. A surviving mutant then asks whether the agent's behavior changes without that rule.
If you find a surviving mutant in your own suite, or a prompt shape the mutator misses, open an issue. If the idea seems useful, a star on GitHub helps other people find it.