Pin the Failure Class. Refuse the Free Retry When the Label Flips.

# tutorial# python# devops# ai
Pin the Failure Class. Refuse the Free Retry When the Label Flips.Dakota Liu

A remote retry is a decision, not a reflex. If two local runs do not share one failure class, I do...

A remote retry is a decision, not a reflex. If two local runs do not share one failure class, I do not boot a box, and I do not send the fixture anywhere.

Did the error change because the code changed, or because the check is unstable? Answer that on disk first. This tutorial builds a small gate you can run from an empty directory. The script is a proposal. I have not timed it against a live service for this draft, and I will not pretend otherwise.

What "working" means here

You finish with a ledger file, not a vibe. The ledger says fix_locally, hold, or eligible_for_isolated_server.

Each step below has a command and a check. If the check fails, you stop. You do not skip ahead to a model just because the call is free.

Free is not the same as harmless. A free call can still leak a path, a token, or a half-written secret into someone else's log. A free server can still see whatever you mount. So the first artifact stays local.

Step 1. Create the scratch tree and prove the interpreter

Start from nothing. A dirty checkout will mix old failures into the new label, and then you will classify a ghost.

mkdir -p "$HOME/failclass/fixtures" "$HOME/failclass/ledger"
cd "$HOME/failclass"
python3 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -c "import sys; print(sys.prefix)"
Enter fullscreen mode Exit fullscreen mode

Verification: the printed prefix must contain failclass/.venv. If it points at /usr or a conda root, deactivate and start over.

Why so fussy? Because a later non-zero exit from the wrong interpreter is not the failure you think you captured.

Write one boring fixture. A missing key is enough. You already know the cause, and that is the point.

# fixtures/case_missing_key.py
def load_budget(doc):
    return int(doc["remaining"])

if __name__ == "__main__":
    load_budget({})
Enter fullscreen mode Exit fullscreen mode

Run it.

python fixtures/case_missing_key.py; echo "exit=$?"
Enter fullscreen mode Exit fullscreen mode

You want a KeyError and a non-zero exit. If the exit is 0, this is not a failure fixture. Replace it. Do not classify a success.

Step 2. Freeze a four-label vocabulary

I keep the vocabulary rude and short. A long taxonomy is how a retry sneaks through while you are still naming things.

Class What you saw locally Send to a model? Boot a server?
local_typo SyntaxError, NameError, missing import No No
contract KeyError, TypeError, assertion on your schema No No
needs_runtime Fails only when a port, service, or path layout is required Only after the class is stable Yes, one fixture, nothing else
unknown The two runs disagree, or you cannot say the cause in one sentence No No

Need a fifth label? Add it in the file, commit the change, and re-capture both runs. Do not invent the label in a comment you will forget.

The classifier runs the script twice, maps the exception, and writes ledger/latest.json. This is unexecuted proposal code. Read it before you trust it.

# classify_failure.py
import hashlib
import json
import re
import subprocess
import sys
from pathlib import Path

CLASSES = {
    "SyntaxError": "local_typo",
    "NameError": "local_typo",
    "ModuleNotFoundError": "local_typo",
    "KeyError": "contract",
    "TypeError": "contract",
    "AssertionError": "contract",
    "ConnectionRefusedError": "needs_runtime",
    "FileNotFoundError": "needs_runtime",
}

def run_once(script: str) -> dict:
    proc = subprocess.run(
        [sys.executable, script],
        capture_output=True,
        text=True,
        timeout=15,
    )
    err = proc.stderr or ""
    match = re.search(r"^(\w+Error):", err, re.M)
    exc = match.group(1) if match else "NoException"
    digest = hashlib.sha256(err.encode()).hexdigest()[:16]
    return {
        "script": script,
        "exit": proc.returncode,
        "exc": exc,
        "class": CLASSES.get(exc, "unknown"),
        "stderr_sha256_16": digest,
    }

def decide(first: dict, second: dict) -> str:
    if first["class"] != second["class"] or first["class"] == "unknown":
        return "hold"
    if first["class"] == "needs_runtime":
        return "eligible_for_isolated_server"
    if first["class"] in {"local_typo", "contract"}:
        return "fix_locally"
    return "hold"

def main() -> int:
    script = sys.argv[1]
    first = run_once(script)
    second = run_once(script)
    action = decide(first, second)
    record = {
        "first": first,
        "second": second,
        "stable": action != "hold",
        "action": action,
    }
    Path("ledger/latest.json").write_text(json.dumps(record, indent=2) + "\n")
    print(json.dumps({"class": first["class"], "action": action}))
    return 0 if action != "hold" else 2

if __name__ == "__main__":
    raise SystemExit(main())
Enter fullscreen mode Exit fullscreen mode

Notice the hash. Same class with a different stderr digest is a warning, not a trophy. I still allow fix_locally in that case, but I want the drift visible.

Would you rather discover that drift after a remote rewrite? I would not.

Step 3. Verify the stable path, then force a flip

Run the boring fixture through the gate.

python classify_failure.py fixtures/case_missing_key.py
python -m json.tool ledger/latest.json
Enter fullscreen mode Exit fullscreen mode

Check three fields. action is fix_locally. Both class values are contract. The process exit code is 0.

If class is unknown, the regex missed your interpreter's traceback format. Fix the regex. Do not "fix" it by hard-coding a pass.

Now a fixture that changes its mind. It uses a flag file under /tmp so the two runs disagree. Delete that flag when you are finished.

# fixtures/case_flip.py
from pathlib import Path

flag = Path("/tmp/failclass-flip")
if flag.exists():
    flag.unlink()
    raise ConnectionRefusedError("simulated runtime miss")
flag.write_text("1", encoding="utf-8")
raise KeyError("remaining")
Enter fullscreen mode Exit fullscreen mode
rm -f /tmp/failclass-flip
python classify_failure.py fixtures/case_flip.py; echo "exit=$?"
rm -f /tmp/failclass-flip
Enter fullscreen mode Exit fullscreen mode

You want exit code 2 and "action": "hold". One run saw contract. The other saw needs_runtime. That disagreement is the result that matters.

If this comes back stable, you did not get two different runs, or the flag path is not writable. Stop and fix the fixture.

Would you paste that flip into a chat and say "make the tests pass"? The model will pick one story. Your ledger says you do not have one story yet.

Step 4. Read the ledger before any network call

python - <<'PY'
import json
from pathlib import Path
rec = json.loads(Path("ledger/latest.json").read_text())
allowed = {"hold", "fix_locally", "eligible_for_isolated_server"}
if rec["action"] not in allowed:
    raise SystemExit("bad action")
print(rec["action"])
PY
Enter fullscreen mode Exit fullscreen mode

The order is the policy.

  1. hold stops the workflow. Capture a smaller fixture, or run a third time and keep the disagreement visible. No server process. No prompt file leaves the laptop.
  2. fix_locally means the patch is yours. Open the fixture. Add the missing key or the missing import. Re-run until the process exits 0 because the bug is gone, not because you relabeled it.
  3. eligible_for_isolated_server is the only branch where a disposable machine is even on the table. Copy that one script. Do not copy your home directory, your .env, or your SSH agent socket.

A model can still be wrong on branch 3. The gate does not make it right. It only stops you from asking a remote box to repair a typo.

Why two runs, not a vibe check

People retry because the button is there. I get it. The red text looks like a task, and a coding model looks like a faster colleague.

But a colleague who has not seen a stable failure will invent a cause. Have you watched a patch "fix" a KeyError by catching Exception? That is what an unstable label invites.

Two runs are cheap. They take seconds. They also create a file you can diff tomorrow. A chat scroll does not.

If the second run is annoying to set up, that annoyance is the signal. Your fixture is still coupled to a clock, a network, or a file you forgot to pin.

I am not telling you to ban models. I am telling you to ban the unordered retry. Fix it locally when the class is contract or local_typo. Use an isolated runtime only when the class stays needs_runtime. Hold when the story changes underneath you.

Step 5. Where a free model and a free server fit

Disclosure: This article was prepared as part of MonkeyCode's product outreach.

The operator supplied two availability claims for this draft: MonkeyCode has free model access, and it has a free server option. I am not adding a token count, a model id, a region, a hardware shape, or an expiry. Those details go stale, and this page will not invent them.

Read the current product docs before you depend on either claim. If the docs conflict with this paragraph, the docs win.

Use the claims only inside the gate.

  • Free model access is for a stable needs_runtime note, or for a review of a patch you already applied locally. Send the ledger JSON and the single fixture. Do not send the repository.
  • The free server option is the isolated runtime for that same fixture, and only after action is eligible_for_isolated_server. If the image cannot run from one copied file and no secrets, it is the wrong server for this workflow.

Keep this prompt as a local file until you choose to send it.

action: eligible_for_isolated_server
constraint: name the smallest runtime assumption. do not edit unrelated files.
fixture:
<one script>
ledger:
<latest.json>
Enter fullscreen mode Exit fullscreen mode

If the client hides the outbound body, do not send. You cannot classify a request you cannot see.

That is the only pitch. If you already have those free options, run this gate once before the next call and count how often the action is hold. I would rather hear that count than a slogan.

Limitations

Two runs can both be lucky. A wrapped exception can hide ConnectionRefusedError inside TypeError, and the table will call it contract. The 15-second timeout is a starting guess, not a measured limit. If the fixture hangs, TimeoutExpired crashes this proposal on purpose so you notice. stderr formats also differ across Python versions.

I have not run this on Windows, so the /tmp flag and the venv activate line are Unix assumptions. Treat that as a gap, not as a portability claim.

Who should not use this approach:

  • Anyone in an incident that can destroy data or expose credentials. Follow the incident process, not a blog gate.
  • Anyone who needs a security review. Class names are not a threat model.
  • Anyone who does not have the right to copy the code onto a third-party server.
  • Anyone hunting a published quota. This draft refuses to print one.

Skip it too if the failure is non-deterministic by design, such as a load test. Stability across two samples is the wrong oracle there. You need a budget and a metric, not a traceback label.

Done looks like this

Re-run the missing-key fixture and read fix_locally. Re-run the flip fixture and read exit 2. Remove /tmp/failclass-flip if it is still there.

If you cannot explain the action field without opening a chat window, you are not done. Name the break twice. If the label flips, the free retry can wait.