Notes from the work, not the hype.
Practical thinking from the engineers and operators replacing manual work with systems that actually ship in production.

Approve-every-action gating trains reviewers to click yes. How to gate on uncertainty, blast radius, and reversibility instead — with a decision table.
Read more
Enterprise agent projects don't stall on model choice or prompts. They stall on stale ETL, row-level permissions, and no clean API. A CTO's de-risking checklist.
Read more
The chunking strategies we tried and abandoned across messy enterprise corpora — and the structure-aware, metadata-rich approach we default to now.
Read more
How we version, eval-gate, canary, and roll back prompts in production AI systems — and why editable dashboard prompts silently degrade quality across a fleet.
Read more
Teams grade answer quality while ignoring whether the right passage was even retrieved. How to build a retrieval-only eval and why it comes first.
Read more
A candid field guide to the five RAG failure modes we keep diagnosing in client audits — symptom, real root cause, and the cheapest fix for each.
Read more
The harness we use to break an agent before production does: adversarial inputs, failure injection, tail latency, cost per resolved task, and injection probes.
Read more
An anonymised case study: where an agent's tokens actually went, the changes that cut spend ~60%, and the measurement discipline that proved quality held.
Read more
An anonymised case study of a customer-support agent that deflected ~40% of inbound tickets without escalation-rage — the failures, the fixes, and the deflection ceiling.
Read more
How to detect AI agent failure in production before users notice — drift detection, online monitoring, confidence-based fallbacks, and the telemetry that makes it work.
Read more
Most enterprise tasks handed to us as 'agent' projects ship better as constrained workflows. Concrete criteria for deciding when to give a model agency — and when to take it away.
Read more
Production agents balloon their tool-call counts with redundant retrievals and second-guessing loops. How to instrument, budget, and detect thrashing.
Read more
LLM self-reported confidence and logprobs are badly calibrated. Here's how that breaks human-in-the-loop gating in production agents — and the cheap fixes that work.
Read more
Golden eval sets silently rot within months. How we source real edge cases, version evals alongside prompts, and refresh without invalidating history.
Read more
An anonymised case study of an AI agent that worked in pilot but blew the budget at scale. Where the cost went, the fixes, and the before/after numbers.
Read more
AI-generated code ships exploitable flaws with total confidence. Why security review is the one step you can't delegate to the tool that caused the problem.
Read more
DeepMind's AI Control Roadmap signals the shift from hoping models stay aligned to securing agents like insider threats. Here's the executive mental model.
Read more
For agentic systems, the eval harness is what lets you ship with confidence and iterate safely. Here's what goes into a real one — and why teams pay for skipping it.
Read more
The specific failure modes that kill RAG systems after the demo — retrieval drift, chunk boundaries, stale indexes — and how to catch them with evals first.
Read more
A decision framework for CTOs and Heads of AI in 2026. Four paths — Build, Buy, Partner, Wait — with honest costs, timelines, and failure modes.
Read more
The month-four death has remarkably consistent causes. None of them are technical. Four organisational failure modes that kill AI pilots.
Read more
Case study: 100+ vehicles, 2 years, 40% downtime reduction. What we built, why it worked, and the patterns worth stealing.
Read more
There is a specific shape that enterprise AI projects take when they actually ship. Here's the week-by-week breakdown.
Read more
The question is almost never framed correctly. A framework for deciding how to stand up AI capability — honest about where each option breaks.
Read more
Most enterprise AI projects fail. Here's what separates the ones that ship from the ones that stall — based on what we've seen across 6 industries.
Read more
You need senior AI leadership but aren't ready for a full-time hire. Here's how the fractional model works and when it makes sense.
Read more