- What a Two-for-Two Agent Result Actually Proves
A 2/2 completion result is useful evidence of a cohort outcome; it is not automatically proof of a mechanism or a production default.
4 min
- Grounding Is Necessary, Not Sufficient
File-path grounding blocks a class of fabricated plans, but it cannot establish that a cited file supports the claim being made.
4 min
- A Plan Can Validate and Still Be Unsafe to Implement
A controlled indexed-planner study shows why structural validation, path grounding, and semantic review must be reported as separate gates.
5 min
- Designing Bounded Repair Loops for Agent Plans
A bounded repair action makes plan validation observable and safe, but repair telemetry must be tied to the final outcome rather than counted as success by i...
3 min
- Why Build an Agent Harness?
An agent harness turns coding-agent activity into a bounded, inspectable, and repeatable engineering process rather than a stream of opaque terminal outcomes.
5 min
- Experiments Should Be First-Class Product Artifacts
If an experiment cannot be inspected after the run, its result is operationally weaker than it needs to be.
3 min
- Benchmarking Agent Systems Beyond Did It Finish?
A practical scorecard for agent planning experiments separates completion, verification, grounding, repair efficiency, latency, and token cost.
4 min
Back