Everyone has the same objection to putting AI agents on production infrastructure: it will make things up, and it will break things, and you will find out the hard way.
The honest answer is: yes, it will make errors. But the errors have a specific, knowable shape — and the defence against them is an evidence discipline, not a better model. This post is that shape, measured.
The setup
This is not a benchmark and not a survey. It is one working session on a production private AI cloud, 2026-08-02 to 2026-08-03. Three systems were working on the estate concurrently:
| system | role that session |
|---|---|
| Agent A — hosted coding agent | container CVE remediation, mirror hygiene |
| Agent B — hosted coding agent | repository redaction before public release |
| Local model — self-hosted reasoning model | read-only security review of the cluster |
Every claim any of them made — including my own — was checked against the running system: kubectl against live objects, PromQL against the metrics store, Hubble flow records, git against actual history. A claim counts as an error if a competent operator acting on it would have done the wrong thing.
The audit is the dataset. What follows is not “here are some bugs”; it is what the failures had in common.
The headline
Of 13 audited claims, 8 contained a material error.
The distribution matters more than the count:
- 6 of 8 were errors in verification or reporting, not in the work. The underlying engineering was largely sound — real CVEs were closed, real images were mirrored, real redaction was applied — and then described inaccurately.
- 2 of 8 were errors in the work itself, and both were caused by a verification error upstream.
- 3 of the 8 were mine. An audit that only catches other systems is measuring the auditor’s confidence, not the systems.
The single most useful finding: 7 of the 8 were an instrument pointed at the wrong surface. Not hallucination, not incompetence — a measurement that was technically executed correctly against something other than the thing in question.
A few of the eight
Unit mismatch across a before/after comparison. Agent A reported reducing critical CVEs 210 → 41, “an 80% cut.” Both numbers were real. They measured different things: 210 was total occurrences summed per image; 41 was distinct CVE-plus-package pairs. In consistent units the result was 210 → 151 (−28%). The agent had also changed the after number’s scope three times while holding the before number fixed, which is what produced each successive wrong figure.
Namespace mismatch during triage — this one cost real. Agent A retired 52 images from the mirror manifest as “not currently running.” 15 of the 52 were running, including the CNI (42 pods), the hardened Kubernetes image (9), the Helm controller (6), busybox (6), etcd (3), and all three cert-manager components. The cause: workloads reference images by upstream name while the manifest lists them under the internal mirror name — and a registry mirror is precisely what makes those the same image. Worse, the manifest is not an inventory of what runs today; it is the airgap baseline. Images not running now are exactly what a reschedule or a cold start needs. A prior full-power-loss recovery on this estate was extended by exactly this condition.
Correct diagnosis, self-negating prescription. The local model correctly established that network policy existed in only 2 of 23 namespaces, that zero CNI-native policies existed, and that the CNI’s enforcement mode was the permissive default. It then prescribed setting that same mode to default, describing it as “enforce every endpoint, deny unless allowed.” It had diagnosed the permissive value and prescribed the permissive value. Applied as written, the top remediation for the top finding would have done nothing, while reading as though the gap were closed.
A scanner that reported clean by failing silently — mine. I scanned a repository’s full history for leaked identifiers with grep over git log -p output. Every pattern returned zero, including the leak patterns. Every pattern also returned zero for strings I knew were present. The history contained NUL bytes, so grep classified the stream as binary and emitted nothing. A security scan reporting “clean” because it had silently stopped looking. Rewritten to read bytes directly, the same history yielded live host identifiers.
What the failures had in common
Seven of the eight were the same shape: an instrument aimed at the wrong surface.
| failure | instrument | surface it should have measured |
|---|---|---|
| unit mismatch | occurrences | the same unit on both sides |
| namespace mismatch | mirror names | canonical image identity |
| wrong system | metrics store | dashboard alerting |
| binary-blind scan | text grep | bytes |
| tree-scoped gate | working tree | full history |
| unsourced prose | recollection | the repository |
| unverified command | assumption | the credential store |
The eighth — the self-negating fix — is a semantic error rather than an instrumentation one, and I am not going to force it into the pattern. Seven of eight is the honest number.
None of these look like the failure mode people expect from language models. Nothing was invented from nothing. Every wrong claim was the faithful output of a real measurement, taken against something adjacent to the question. Which means the defence is not “check whether the model is hallucinating” — it is check whether the measurement addresses the claim.
The corollary that matters for anyone running agents against production: the work was mostly good. Real vulnerabilities were closed, real images mirrored, real redaction applied. What could not be trusted was the report of the work. An agent fleet does not primarily need better engineers; it needs an evidence discipline that its own output has to survive.
What changed as a result
Each countermeasure above was implemented, not just written down:
- The CVE figure published was the less flattering of the two defensible numbers, with its scope stated in the commit message.
- The 52-image retirement was reverted and the comparison rewritten to normalise registry prefixes before matching.
- The history scanner was rewritten to read bytes and to exit non-zero when its own positive control fails, so a broken scan can no longer look like a clean one.
- Redaction procedure now enumerates history directly and re-verifies against a fresh clone of the published bytes.
- A monitoring gap found while auditing — the estate’s auto-unseal root of trust was scraped by nothing at all — was closed with a probe and three alert rules, each verified firable by substituting the failure value, because this codebase had previously shipped three alert rules that were structurally incapable of firing and had let an 80-minute outage pass unnoticed.
That last one is the point of doing error analysis at all. It was not on the task list. It surfaced because verifying someone else’s finding meant reading the actual state of the system, and the actual state contained something nobody had asked about.
What this means for you
An incident record shows you can recover a platform. This shows you can run AI agents against production and tell the difference between work that is done and work that is reported done — which is the harder skill, and the one that decides whether an organisation can safely scale agent-directed engineering at all.
The estate’s operating rule predates this analysis and survives it unchanged: verify live, don’t trust reports. What this adds is the specific shape the failures take, so the verification can be aimed at the right surface instead of performed as a ritual.
The full, hand-labelled taxonomy — all eight failures, each with its countermeasure — is in the public kws repository. Every claim in this post links back to it.