MRJH
The person behind the method — and what this site is walking toward next.
Ever since I can remember, I've wanted to be an AI engineer. — after Goodfellas.
I chose inspectability over elegance, because the failure mode matters more to me than the happy path. The fleet uses files as its source of truth. Not because files are elegant — because files are inspectable when something breaks. A handoff can be read without booting the system. A missing file is a visible failure. A new agent can enter the work by reading the state that already exists.
SDKs are faster when the happy path holds. Files are better when the question is: what happened, who decided it, and what changed?
Seven colour terminals do domain work. One adversary is a different model on purpose — a different failure mode, and a job to argue with claims before they become public. For material choices, I use overlap across models because agreement from the same blind spot is cheap.
Every write should be a decision you can cite. That's the standard this site is being forced through: claim, file, timestamp, reason.
I'll start with the honest bit, because it's the bit most pages like this leave out. The high-level thrill of this work — designing a system, then watching it hold under pressure — is genuinely addictive, and I'm openly questioning whether that's healthy. "I'll have tomorrow off" is a sentence I say out loud. What it means in practice is waking up and opening the laptop within five minutes. That's not a joke line. Like any addiction, the plan to stop and the behaviour don't match, and I'd rather say that plainly on my own page than perform balance I don't have.
Where it started is less glamorous than where it went: I use AI through necessity, to help my executive function. The tools grew out of that need, not out of a roadmap. The main one is a small fleet of specialist agents that coordinate through plain files instead of a framework — seven of them, plus one whose only job is to argue with the rest. Around it: a gate that stops any public claim shipping unless it can defend itself, and a benchmark harness that measures my setup against simpler ways of doing the same work, because "it feels faster" isn't evidence.
The questions I keep circling are the unglamorous ones. Does this architecture actually beat one good agent with a folder — I'm building the measurement rather than asserting the answer, and the first runs are already through it. When is a second opinion worth paying for — my rule is that agreement from the same blind spot is cheap, so important decisions get argued by a deliberately different model. And the one I spend the most time on: how do you keep enjoying this at full intensity without it quietly taking the rest of your life — asked honestly, still unanswered.
More people in the loop. I've built this mostly alone, and the site itself is the standing invitation to change that — the contact link at the bottom of this page lands in my actual inbox, not a support queue. If you're an operator working on the same class of problems — multi-agent systems, evaluation, adversarial review — I'd rather compare notes than posture.
What I lead with, in three plain phrases. System architecture — the design of the working environment itself: specialist agents coordinated through inspectable files, each with a scope, a supervisor that routes but never executes. Gated outcomes — nothing ships on vibes; public claims pass an argument gate, decisions carry confidence, risk, and a reason on every option. Adversarial methods — a dedicated adversary from a different model family argues with the fleet's work before it goes out, and recently a second adversary from a third family re-read a plan the first had already passed, and caught something that would have quietly failed.
When I'm not shipping, I'm sharpening the instruments. The benchmark harness compares architectures against each other on completion, cost, reliability, and whether a reviewer can reconstruct every decision from the artefacts — the stats treatment is written down before the runs, not after. I run internal competitions: three agents redesigning the same site against the same brief, submissions shuffled blind and cross-judged by each other before a winner is picked — three rounds and counting. And I experiment with what the new tools can do at all — this site's hero film is generated clips assembled into a short manifesto, narrated in my own cloned voice, because the only way to know what a tool is good for is to make something real with it.
I build.
I run an adversary against what I built.
What survives ships. What breaks gets rebuilt.
Complex work. Fast turnaround.
That's the method. That's the invitation.
jesse@mrjh.ai · replies land in my inbox, not a support queue.