- Why Build an Agent Harness?
An agent harness turns coding-agent activity into a bounded, inspectable, and repeatable engineering process rather than a stream of opaque terminal outcomes.
5 min
- Zero-Spend Concurrency Auditing for Agent Sandboxes
A mock-driven diagnostic method for exercising sandbox resource cleanup without live-model spend.
3 min
- Catching Hallucinated Paths in RLM Trajectories
How pre-finalization trajectory hooks catch ungrounded file paths and trigger self-correction turns before final answer submission.
2 min
- Grounding Is Not Proof: 100% Code, 0% Plan Quality
Why a 100% task completion rate does not guarantee plan quality, and how pre-finalization checks prevent hallucinated path claims.
3 min
- Inside the Sandbox: Unconstrained Shell Execution
Exploring why allowing agents arbitrary host shell access leads to non-deterministic failures and why strict Docker isolation is mandatory for production.
2 min
- The Security Cost of Autonomy: Auditing 29 Runs
Empirical security audit results from 29 benchmark study runs evaluating container mount points, path traversal attempts, and write-tool boundaries.
2 min
- From REPL Loops to Swarms: Leasing Sandboxed Containers
Designing SwarmLeaseManager to lease container worker nodes dynamically for concurrent file inspection and sub-task execution.
3 min
- Targeted Trajectory Repair: Continuations vs Reruns
How provenance-linked continuation child runs inherit parent trajectory history to perform targeted grounding repairs at a 90% reduction in LLM compute costs.
2 min
Back
Blog
Page 1 - Showing 8 of 19 posts
View all posts by years →