Elena RevichevaOriginally published at aideazz.xyz — cross-posted here with canonical link. A field note from the...
Originally published at aideazz.xyz — cross-posted here with canonical link.
A field note from the AIdeazz AI Lab — a real incident on a live production system, written up from the logs. September 28, 2026.
A job-discovery agent read the operator's rejection notes and screenshots every hour, summarised the reasons and put them in its judge's prompt — a loop proven end to end a month earlier. Yet the same kinds of job kept arriving. The lessons reached one reviewer on a side path, were labelled as unable to override its criteria, and were drawn from a window that remembered eight days.
The operator kept rejecting the same shapes of job in the CRM and writing why — a posting that was already closed, a residency list that excluded her country, a role centred on a tool she does not use, a requirement to hand-write production code. The agent's learning loop reported success on every run: it read her notes, read the screenshots she attached with a vision model, and wrote the lessons into the prompt of an AI judge. Every inspection of the loop was green. The pipeline's own precision told a different story — in the week of 21 September she applied to 2 of the 15 jobs it put in front of her and rejected 13.
Three wiring faults, each invisible to a check of the loop itself. First, the main door into her action queue — the Google-Jobs search, run through a scraping provider — sent every job that passed a set of fixed rules straight to her; the judge, the only component that read her lessons, was consulted there only to rescue a job the rules had rejected, never to veto one they had passed. Second, the lessons were presented to the judge as calibration that must not override its base criteria, so even a hard eligibility rule was a hint. Third, the sync read the 400 most recently modified deals and kept twelve examples: with about 150 new deals a week, that window reached back eight days and contained 18 of her 370 rejections. The screenshot reader, meanwhile, asked only what a role demanded, so a location restriction and a closed notice were dropped, and 18 reads had been lost to rate-limit errors with no retry.
The loop was connected to the decision rather than rebuilt. A permanent ledger now keeps every decided deal — 554 at deployment, re-reading notes only when a deal changes. Crisp lessons became enforced rules with provenance, each naming the rejections that taught it: closed postings, stated eligibility such as citizenship or place of birth, country-code rosters that exclude her, tools she does not use, companies she rejected three times in ninety days, and near-duplicates of an out-of-field title. Those rules and the judge now run on every door into her queue. The judge receives a summary of all her lessons and may reject on one only when the listing itself states the fact; the vision prompt now extracts location, closed status and pay, retries rate limits and falls back to a second provider. Two traps surfaced by the new tests were fixed on the way: the agent's own note template was being learned as her reason, and a pre-existing judge bug rejected Latin-America roles as possibly excluding her Latin-American country — which no prompt wording fixed, so the criterion is now enforced in code.
By replay against her real decisions, on the production server. The learned rules catch 12 of the 370 historical rejections and wrongly block 0 of the 35 jobs she applied to. On a sample of 20 rejections carrying her reason and 20 applications, excluding the twelve examples quoted in the prompt, the judge agreed with 19 rejections before the change and 20 after one fix; a later guard released one — rejected by her as closed, which the closed-posting rule catches on the real posting text — for a final 19. It approved 2 of her 20 applications before and 4 after. The replay also caught a regression before it shipped: the lessons and the location-heavy examples, each harmless alone, together turned a listing silent on location into a veto for a job she had applied to, in 2 of 2 runs; a reminder placed nearest the job restored it to 3 of 3. The evaluation suite passed 546 of 547 on the server, the one failure being a provider whose credits are deliberately at zero. The weekly precision is now written to a metrics table that had held no rows since it was created.
Feedback that cannot change a decision is logging, not learning. To tell the difference, trace backwards from the outcome the human keeps rejecting to the exact line that let it through, and ask whether the lesson can reach that line — green checks on the sensor and the memory prove nothing about the switch. Make crisp lessons into rules with provenance, keep memory as a record rather than a recency window, and prove the change by replaying the human's past decisions: count the rejections it now catches and the approvals it now blocks, which must stay at zero. Re-run that replay after every prompt change, because prompt pieces interact — two harmless additions can combine into a new failure.
Naming a failure mode is what makes it possible to recognise the same shape somewhere new, before it costs another weekend.
Feedback that is recorded, summarised and shown to a model — but wired where it cannot change a decision — is logging, not learning.
A self-improving system needs three parts, and it is easy to build two. A sensor that captures the human's verdict and the reason for it. A memory that keeps those verdicts. And a switch — the point in the pipeline where the decision is actually made — that the memory is allowed to move.
Advisory feedback is the failure where the first two exist, work, and are proven end to end, while the third is missing. The verdicts are read, stored, summarised and even injected into a model's prompt, so every inspection of the loop comes back green: the file updates, the prompt contains the lessons, the logs show the sync running. And the outcome does not change, because the component that receives the lessons is not the component that decides.
The usual shapes:
It is a cousin of [[silent-failure]] — nothing errors — and of [[verify-from-logs]]: the tell is that you can show the lesson arriving but not the decision it changed.
The defences. Trace from the outcome backwards, not from the sensor forwards: for the thing the human keeps rejecting, find the exact line that let it through, and check whether the lesson can reach that line. Turn crisp lessons into enforced rules with provenance, and leave only fuzzy judgement to the model. And prove it by replay — run the new logic over the human's past decisions and report two numbers: rejections it now catches, and approvals it now wrongly blocks, which must stay at zero. Re-run the replay after every prompt change, because prompt pieces interact ([[the-prompt-is-a-source]]): two harmless additions can combine into a new failure.
A numbered claim in a system prompt is as much a source as a log line -- and a fail-closed gate will treat it that way.
A verifier that refuses unsourced numbers is doing the right job. The model still has to write from something. If that something -- a topic brief, a few-shot example, a "write about X" paragraph -- contains a leftover figure, the model copies it. The gate then fires on a number that was never in the evidence file, and the pipeline skips even when the day produced plenty of real facts.
The trap is treating the prompt as flavour text. The model does not. To the generator, a brief that says "BrightData $40/run" is a fact. To the gate, $40 is unsourced. Both readings are locally correct. Cadence dies in the gap.
Ordinary-life version: you ask someone to write the minutes from the meeting notes, and you also slide them last year's budget that still says the coffee machine costs forty dollars. They copy the forty. The auditor who only accepted numbers from the notes throws the minutes out. Nobody invented the forty. The briefing packet did.
The defences are structural:
Configuration tells you what somebody intended. Logs tell you what happened.
A setting, an environment variable or a present API key is a statement of intent. It is evidence that somebody meant for a behaviour to occur. It is not evidence that the behaviour occurs.
The gap between the two is where the longest outages live, because reading the configuration feels like verification. It produces confident, wrong statements: the key is set, so the provider works; the schedule says every fifteen minutes, so it runs every fifteen minutes; the file was deployed, so the new code is running.
Each of those has a cheap, decisive check that costs seconds:
The rule this earns: never report a system's behaviour from its configuration. Grep the line that proves the behaviour happened, and quote it.
This note is one entry in a running wiki of production engineering lessons — every concept linked to the incident that taught it — at aideazz.xyz/ai-ops-wiki.html.
No customer data, credentials, hostnames or internal record identifiers appear in these write-ups.