- What a Two-for-Two Agent Result Actually Proves
A 2/2 completion result is useful evidence of a cohort outcome; it is not automatically proof of a mechanism or a production default.
4 min
- Inside the Sandbox: Shell Execution Kills Reliability
Exploring why allowing agents arbitrary host shell access leads to non-deterministic failures and why strict Docker isolation is mandatory for production.
2 min
- A Plan Can Validate and Still Be Unsafe to Implement
A controlled indexed-planner study shows why structural validation, path grounding, and semantic review must be reported as separate gates.
5 min
- Designing Bounded Repair Loops for Agent Plans
A bounded repair action makes plan validation observable and safe, but repair telemetry must be tied to the final outcome rather than counted as success by i...
3 min
Back