Ever since I can remember, I've wanted to be an AI engineer. after Goodfellas.
WHAT SURVIVES TEN ROUNDS SHIPS.
TEN TIMES. NO MERCY.
Build to break is a review method. The target is ten rounds. Round one is context. Two through five are check, improve, build-on, benchmark. Six through ten repeat — each round now attacking what the previous round claimed to fix.
Why ten? Because five can still be a demo and twenty can become hiding in the process. Ten is enough pressure to expose the next pass without letting the method become the product. The number isn't sacred. The pressure is.
The five gates map to failure modes I actually hit. CONTEXT_GATE for the run where I answered the wrong question. CHECK for the model agreeing too fast. IMPROVE for the "better on my metric only" trap. BUILD_ON for the restart-instead-of-compose habit. BENCHMARK_SCORE for every claim the thing works.
No mercy. The round is judged against the criterion frozen at the start. Effort doesn't grade. If the output misses the criterion, the miss is the finding.
Pressure can become noise. Late rounds start finding style issues instead of substance. That's a real cost. When it happens, the answer isn't to worship the last round. Go back to the criterion frozen at the start. Ship the strongest round that still satisfies it.
The gates depend on the criterion being honest. If I write one loose enough to pass, the method degrades to theatre. The corrective is to show the criterion to a different model before the round starts. I don't do this every time. When I skip it, the method is only as sharp as I was that morning.
The method doesn't choose what to build. It sharpens execution once a project is chosen. CONTEXT_GATE can catch "answered the wrong question." It doesn't tell me which question was worth asking.