Jordan HuangA prompt change looks small. You swap one verb. You add a constraint. CI stays green. The tests pass....
A prompt change looks small. You swap one verb. You add a constraint. CI stays green. The tests pass. Then a downstream service starts receiving strings where it expected objects. Or the model refuses to answer a class of requests. Nobody notices until a user does.
I've hit this enough times to stop treating prompt edits as config. They are behavior changes. And behavior changes need a regression budget.
The problem isn't the model. The problem is review by reading.
When I read a prompt diff, I can tell if it is grammatical. I cannot tell if it changes the output shape for 200 edge cases. A free model endpoint can make that worse. The same prompt can drift across calls, and rate limits punish the "just run it again" instinct.
That forced me to build a prompt evaluation diff in GitLab CI. It doesn't need a big platform. It needs golden cases, a normalization step, and a merge request artifact.
None of these are "does it sound better?" They are mechanical checks. That is the point.
Here is the tiny repo shape I use:
prompts/
base.yaml
candidate.yaml
cases/
golden.jsonl
scripts/
eval_diff.py
.gitlab-ci.yml
Prompts are versioned YAML, not inline strings in a job. That alone changes the review conversation. The MR diff becomes a real diff.
# prompts/candidate.yaml
name: extract_key_points
system: |
You are a JSON-only assistant.
Return in this exact shape:
{"points": ["..."]}
Each golden case has an id, input, and a small set of assertions. I don't store "the correct answer." I store the contract.
{"id":"case_012","input":"Explain the release process","expect":"json_object","required_keys":["points"],"max_output_chars":800}
One more thing: a prompt diff should not call the model twice for every push. If the base prompt hash already has an evaluation, I reuse it. That is the same content-addressed instinct I use elsewhere: don't spend quota on what hasn't changed.
Exact string matching is useless with a low-cost model. The same meaning can arrive with different whitespace, synonyms, or key order.
So my eval script normalizes before comparing:
I don't claim the model is deterministic. I just measure whether the candidate and base are functionally close on stable cases.
eval_diff.py loads both prompts, calls the model endpoint once per case for the missing side, and writes a report.
import json, hashlib, os, sys
def normalize(value: str) -> str:
return " ".join(value.lower().split())
def content_hash(text: str) -> str:
return hashlib.sha256(text.encode("utf-8")).hexdigest()[:16]
def main(base_path, candidate_path, cases_path, report_path, max_regressions):
base = open(base_path).read()
candidate = open(candidate_path).read()
# In a real client, call the model endpoint for text and parse JSON.
report = {"base_hash": content_hash(base), "candidate_hash": content_hash(candidate), "cases": []}
regressions = 0
for line in open(cases_path):
case = json.loads(line)
# Mocked for illustration: replace with a real model client.
# Actual code should run the prompt and check the returned object.
result = {"shape_ok": True, "required_keys_ok": True, "drift": 0.0}
passed = result["shape_ok"] and result["required_keys_ok"] and result["drift"] < 0.25
if not passed:
regressions += 1
report["cases"].append({"id": case["id"], "passed": passed})
report["regressions"] = regressions
json.dump(report, open(report_path, "w"), indent=2)
if regressions > max_regressions:
sys.exit(1)
This is deliberately small. The real model client belongs in a separate module so you can swap endpoints without touching the test.
I gate the check on merge requests only. The job runs the same script, saves the report, and fails the pipeline when regressions exceed the budget.
stages:
- eval
prompt_diff:
stage: eval
image: python:3.12-slim
variables:
MODEL_URL: $MODEL_URL
before_script:
- pip install -r requirements.txt
script:
- python scripts/eval_diff.py
--base prompts/base.yaml
--candidate prompts/candidate.yaml
--cases cases/golden.jsonl
--report eval_report.json
--max-regressions 0
artifacts:
paths:
- eval_report.json
expire_in: 7 days
rules:
- if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
Now the reviewer doesn't have to trust a comment that says "LGTM." They can open the report and see which cases broke.
The eval runner is small and bursty. It only runs on MR events. I don't want that load on a paid CI runner if I can avoid it.
I can run the same script from a small sidecar on MonkeyCode's free server option, pointed at MonkeyCode's free model access. Disclosure: This article was prepared as part of MonkeyCode's product outreach. That keeps the job isolated without changing the workflow. The script still reads the same prompts, the same cases, and the same report format.
The point is not the provider. The point is that eval belongs next to the model, not inside the pipeline that deploys code. If the endpoint changes, I only change an environment variable.
If you already have a free model endpoint and a free server option, run this report on the next prompt MR and compare it with the existing review thread.
A regression budget won't stop every bad prompt.
If you don't have a small set of stable cases with clear success criteria, skip this. Build that set first. If your prompt changes are rare and you already review each response by hand, the automation may not be worth the maintenance.
I don't merge a prompt change without a report that says what broke. That report has to be a file, not a feeling. It has to be reproducible, and it has to fail the MR when a regression shows up.
That is the same way I treat code. A prompt deserves no less.