A prototype that answers well in a demo is not an agent in production. This checklist gathers the checks I ask for before a release. The security references are those of the OWASP Top 10 for LLM applications (2025 edition): the main risks include prompt injection (LLM01), excessive agency (LLM06, “Excessive Agency”) and unbounded consumption of resources (LLM10).
The ten checks
- Measurable goal. What must the agent achieve, and how do you measure it (resolution rate, time saved, errors)?
- Evaluation set. A set of real cases, even a small one, run on every change of prompt, model or tools. Without it, every change is a leap in the dark.
- Least privilege. Each tool has only the rights it needs, with separate, short-lived credentials. An agent never has administrative access.
- Human approval. Irreversible or sensitive actions (payments, deletions, external sends) require confirmation.
- Prompt injection defence. Treat all external text (emails, web pages, documents) as untrusted: separate instructions from data, and limit the available tools when reading external content.
- Tracing. Record requests, steps, tool calls and costs. The OpenTelemetry semantic conventions for generative AI are evolving: check their status before adopting them.
- Cost and time limits. A cap on tokens, calls and duration for each run; alerts on consumption.
- Data and privacy. Decide which data goes to the model, where it is processed and for how long it is retained; mask personal data where possible.
- Kill switch. A way to stop the agent immediately and return to the manual procedure.
- Periodic review. Re-run the evaluations and review logs and permissions at fixed intervals.
Tools and protocols
The Model Context Protocol is the standard for exposing tools and data to agents. The latest specification is dated 28 July 2026 and makes the protocol stateless: each request is self-contained. Always refer to the version of the specification that your client and your server actually support.
For a bespoke project, see AI agents and automation with LLMs.