Search
1 results for “phpboyscout”
-
Twenty-one critical and high findings from an architectural review of go-tool-base went out to a fleet of agents. Twenty-one merge requests came back, fast, and all of them green.
Somewhere around the fourth or fifth I realised I had no idea whether any of it was real.
A green suite arriving alongside a fix can mean four things:
1. The fix worked.
2. The bug never existed, and the change does nothing.
3. It fixed something else nearby.
4. The test was written after the change, against the changed code, and was never capable of failing.You can't tell those apart by reading the diff, and certainly not by reading an agent's description of it.
So I stopped the flow and asked: is every agent reproducing its bug before touching it?
What came back was the same test run against unmodified main, failing. A CLI invoked wrongly that printed usage and exited 0: "expected 2 actual -1", because the exit function was never called at all. And alongside it, the two non-fatal subtests passing before the fix.
That last part is the negative control. Two things broken, two closely related things demonstrably fine, a clear line between them. Without it a wall of FAILED is a screenshot rather than a proof.
It's cheap to ask for, cheap to read at a glance, and very hard to fake by accident. Review changes shape: less reading the change, more checking the evidence it turned up carrying.