Services
Advisory and review work for teams shipping agent systems and production GenAI—classical engineering discipline applied where the runtime is a language model.
Writing on this site is the product of the publication. Services are optional, scoped engagements for teams that want the same operational honesty applied to their stack. Background and credentials live on About.
Offers
Agent runtime review
Teams shipping recursive or tool-using agents under real budgets.
A written audit of cell/tool/history policy, strategy binding, and ledger behaviour—with ranked risks and concrete knobs to change.
Related writing →Evaluation design
Teams whose dashboards look healthy while jobs still fail.
A metric plan that pairs proxy aesthetics (tokens, cells/turn) with deliverable checks (FINAL, artefacts, stop reasons) and clear horizons.
Related writing →GenAI platform architecture
Enterprise conversational, RAG, and multi-model platforms under audit pressure.
Architecture notes on risk vectors, HITL boundaries, operational honesty, and designs that survive operations—not only demos.
Related writing →Speaking & technical writing
Conferences, internal tech talks, and architecture reviews.
Scoped sessions or notes grounded in production agent and platform work—not vendor decks.
Related writing →How engagements work
- Brief — You describe the system shape, what fails (or what you fear will fail), and what success looks like under real constraints.
- Scope — We agree a bounded review: runtime policy, evaluation plan, platform risk, or a time-boxed advisory block—not open-ended “be our AI team.”
- Work product — Written notes you can action: decisions, knobs, failure modes, and limits. No theatre metrics without a deliverable check.
What this is not
- Vendor selection sales or model-leaderboard cheerleading.
- A full-time staff augmentation replacement for your platform team.
- Guarantees of cost or quality without your telemetry, horizons, and stop conditions.
Start a conversation
Prefer a short written brief over a cold pitch call. Include:
- System under review (agent loop, RAG stack, conversational platform, …)
- Failure mode or decision you need to make
- Success criteria and hard constraints (budget, compliance, timeline)