2026
19 posts
- Why Build an Agent Harness?
- Zero-Spend Concurrency Auditing for Agent Sandboxes
- Catching Hallucinated Paths in RLM Trajectories
- Grounding Is Not Proof: 100% Code, 0% Plan Quality
- Inside the Sandbox: Unconstrained Shell Execution
- The Security Cost of Autonomy: Auditing 29 Runs
- From REPL Loops to Swarms: Leasing Sandboxed Containers
- Targeted Trajectory Repair: Continuations vs Reruns
- The Token Efficiency Curve: Fences vs. Context Dumps
- Zero Net Access: Hardening LLM Agent Sandboxes
- A Plan Can Validate and Still Be Unsafe to Implement
- Benchmarking Agent Systems Beyond Did It Finish?
- Designing Bounded Repair Loops for Agent Plans
- Experiments Should Be First-Class Product Artifacts
- Grounding Is Necessary, Not Sufficient
- What a Two-for-Two Agent Result Actually Proves
- Strategies, Not Model Ifs
- RLM Is Not Automatically Token-Efficient
- When Smaller Cells Make Agents Worse