AI agent in production: a security, evaluation and cost checklist

Published 2 October 2026 · Updated 2 October 2026

In short. An agent goes into production when it has a measurable goal, least-privilege permissions, automated evaluations, tracing, cost limits and a kill switch. The main risk is giving it too much power: in OWASP this is Excessive Agency (LLM06).

A prototype that answers well in a demo is not an agent in production. This checklist gathers the checks I ask for before a release. The security references are those of the OWASP Top 10 for LLM applications (2025 edition): the main risks include prompt injection (LLM01), excessive agency (LLM06, “Excessive Agency”) and unbounded consumption of resources (LLM10).

The ten checks

  1. Measurable goal. What must the agent achieve, and how do you measure it (resolution rate, time saved, errors)?
  2. Evaluation set. A set of real cases, even a small one, run on every change of prompt, model or tools. Without it, every change is a leap in the dark.
  3. Least privilege. Each tool has only the rights it needs, with separate, short-lived credentials. An agent never has administrative access.
  4. Human approval. Irreversible or sensitive actions (payments, deletions, external sends) require confirmation.
  5. Prompt injection defence. Treat all external text (emails, web pages, documents) as untrusted: separate instructions from data, and limit the available tools when reading external content.
  6. Tracing. Record requests, steps, tool calls and costs. The OpenTelemetry semantic conventions for generative AI are evolving: check their status before adopting them.
  7. Cost and time limits. A cap on tokens, calls and duration for each run; alerts on consumption.
  8. Data and privacy. Decide which data goes to the model, where it is processed and for how long it is retained; mask personal data where possible.
  9. Kill switch. A way to stop the agent immediately and return to the manual procedure.
  10. Periodic review. Re-run the evaluations and review logs and permissions at fixed intervals.

Tools and protocols

The Model Context Protocol is the standard for exposing tools and data to agents. The latest specification is dated 28 July 2026 and makes the protocol stateless: each request is self-contained. Always refer to the version of the specification that your client and your server actually support.

For a bespoke project, see AI agents and automation with LLMs.

Need a hand?

If you want to apply these points to your case, tell me in a few lines.

Let's talk