Agent development
Putting an agent into production
The checks we run before an LLM agent is allowed to touch real users or real data: an evaluation set written before the demo, retrieval measured apart from the model, scoped tool permissions, a reviewer and an audit trail, cost and latency budgets, defined failure behaviour, replayable traces, and a runbook the client team can operate without us.
Read the note