Grand Coder: bounded results from internal AI engineering evaluations

I’ve spent more than twenty years in systems and infrastructure engineering. Grand Coder grew out of trying to make that experience useful to the AI agents I work with. I don’t primarily think of it as making a model smarter; the aim is to help a capable model reason and work more diligently.

I’m its creator, and our team uses it. What follows is a report on internal historical evaluations—not an independently reproduced public benchmark.

One bounded positive result

A tiered engineering study ran 96 isolated executions: a matched baseline and three Grand Coder configurations. A documented post-scoring audit found invalid hidden requirements in two task families. Those entire families were excluded across all four conditions, leaving 72 runs, or 18 per condition:

  • Baseline: 13/18 passed.
  • Grand Coder configurations: 16/18, 16/18, and 15/18 passed.

The comparison held the shared baseline setup constant and measured Grand Coder’s additional contribution—not our whole setup against a bare model. The gains were concentrated in two retained task families; the other four were already passing in every condition. The exclusions were post-scoring, not preregistered. This is not a universal 16.7-percentage-point improvement.

What did not support a broad win

Other evaluations included saturated tests with no difference, apparent wins that disappeared after scoring corrections, and long-duration failures. In one build study, blinded review scores improved but corrected behavioral tests showed no advantage, at higher cost. Work on an evolving project and a structured decision task found narrower improvements, not a general solution to autonomous engineering.

The practical lesson for us has been to separate three questions: did the artifact look better to reviewers, did it actually pass the behavioral checks, and what did it cost? An improvement on one is not automatically an improvement on the others. The evaluator also needs auditing: a hidden test that imposes an invalid requirement can produce a misleading treatment comparison.

Disclosure limits

I’m keeping the implementation private for now. The field note shares findings and limitations, not the instructions or executable benchmark fixtures. That limits what others can independently verify, and the historical results do not automatically validate every later revision.

My conclusion is that Grand Coder has produced specific measurable benefits and is useful in our ongoing work. It doesn’t need perfect marks for those improvements to matter, but it does need honest reporting about where they occurred and where they didn’t.

Field note: Grand Coder and more diligent AI engineering

For others evaluating engineering agents: how do you report post-scoring test corrections and localized gains without either overstating the result or throwing away a useful finding?

This is how eval reporting should look — bounded claims, negative results included, disclosure limits stated upfront. More teams should do this.

On your question about post-scoring test corrections: report them as a separate line item, never folded into the pass rate. “16/18 passed, of which 2 required post-scoring test corrections” tells a very different story than 16/18 flat — the corrections measure eval harness quality, not model capability. Fold them in and you’re grading your own test-writing, not the agent. I’d also log correction rate per task type; if corrections cluster in one area, that’s where the harness is weakest and where the next round of test-hardening should go.