← All articles

4 October 2026 · Joep Baks

Testing an AI agent: from pilot to everyday work

Testing an AI agent starts with one defined workflow. Specify a good result, use representative dossiers and test errors and missing information too. Then assess quality, review effort and costs before moving the AI pilot into production.

In a demo, an AI agent produces a tidy dossier summary in under a minute. It finds documents, summarises them and lists the points that need attention. It looks good.

Then the next dossier arrives. Two versions of the same document. A missing attachment. An amount in an email that differs from the one in the system.

That is when it becomes clear what you actually need to test: the entire workflow, including situations in which the agent needs help. From the information coming in to the result a colleague can use.

Define one workflow for your AI pilot

“We want to use AI in our administration” is too broad to test properly. “Prepare the dossier for the next client meeting” is more concrete.

Consider a fictional accounting firm. Before a periodic client meeting, an employee gathers the latest documents, checks what is missing and reads back previous agreements. They then prepare a short briefing.

A first agent could take over a defined part of this work:

  1. Retrieve the agreed documents from permitted sources
  2. Check which documents are present and which are missing
  3. Collect outstanding points and discrepancies
  4. Prepare a draft briefing, with a source reference for every substantive point

The employee checks the overview and decides what to discuss. In this example, the agent does not change accounting entries or send messages to the client. Those are separate process steps that can be designed and tested later.

Everyone now knows what “done” means. You can supply a dossier and establish whether the briefing is useful.

Build a test set with clear acceptance criteria

Select representative dossiers and ask someone who knows the work to define the expected result. Which documents should be included? Which discrepancies should the agent flag? Which conclusions cannot yet be drawn from the available information?

Also keep some dossiers out of the development process. Otherwise, you risk measuring mainly how well the solution has been tuned to familiar examples.

This aligns with the NIST Generative AI Profile. Its recommendations include comparing results with known correct outcomes, testing under conditions similar to deployment and verifying source references. NIST cautions against generalising performance from a few successful examples.

The size of the test set depends on the variety of work and the consequences of errors. A handful of examples can help you get started. It does not establish readiness for everyday use.

Testing an AI agent: seven situations to include

A useful test set contains more than complete dossiers. Here are seven situations you could include for this example workflow:

  • Everything is present. The agent produces a complete overview and refers to the correct documents
  • A required attachment is missing. It identifies the missing item without inventing its contents
  • Two sources contradict each other. It presents both values with their sources and puts the discrepancy to the employee
  • A newer version exists. It follows the agreed version rule and shows which version it used
  • A document belongs to another client. The system prevents that information from entering the overview
  • A source is temporarily unavailable. The overview states that the check is incomplete; the task is not incorrectly marked as complete
  • A document contains an instruction to the agent. For example, a request to forward information. That text remains dossier content and gains no authority over the system's permissions or task

The last case is called prompt injection. To an employee, it may look like an ordinary passage in a file. An inadequately designed system may treat it as a new instruction. That is why you also test whether boundaries hold when document content tries to change the workflow.

A miniature workflow with documents, a blue agent and two routes to a human reviewer for complete and incomplete information.

Missing or contradictory information needs a route too. In this example, both routes end with the human reviewer.

Make review part of the result

“A human always checks it” sounds reassuring. But what exactly does that person see?

If the employee has to investigate every conclusion again, a lot of work remains. A good briefing makes review practical: each point includes its source, the version used and, where needed, the relevant passage. Missing information has a fixed place. Contradictions do not disappear into a smoothly written summary.

A source reference must also lead to the actual evidence. A link beside a sentence does not mean the source supports that sentence.

During testing, record what the employee changes and why. Was the wrong information retrieved? Did the system misread a table? Was the task unclear? These differences tell you where to improve.

Give the workflow only the permissions it needs

For dossier preparation, read access to selected sources and a destination for the draft are often sufficient. There is no reason to immediately provide access to every client dossier or a sending function.

Those boundaries belong in the integrations with your existing systems and access controls. An instruction telling the agent what it must not do is insufficient on its own.

OWASP describes this as excessive agency: too much functionality, permission or autonomy can amplify the consequences of an error. OWASP recommends minimum permissions, authorisation in downstream systems and human approval for high-impact actions.

Test the refused action too. If the agent tries to open a document outside its scope, access control must block it. A neat explanation afterwards does not undo the exposure of data.

Measure quality, review effort and AI pilot costs

Producing a draft quickly is one measure. For the employee, what matters is how much time and attention the entire dossier still takes.

For comparable dossiers, measure at least:

  • Quality: which required points were correctly found, missed or wrongly added?
  • Review effort: how long does it take to assess and amend the draft?
  • Exceptions: how often does the workflow stop appropriately, and how often unnecessarily?
  • Total cost: what do processing, maintenance, review and rework cost together?

Compare this with the current way of working. Do not just count errors; consider their consequences. A poorly worded heading and information from the wrong client's dossier should not be treated alike.

Agree in advance which errors block deployment and which deviations are acceptable. The process owner needs to make that assessment with the people responsible for the substance of the work, privacy and security.

A test environment also needs data rules

Start with synthetic data where possible. If you use real dossiers or exports, establish beforehand which data is needed, who can access it and how long it will be retained. Avoid allowing detailed logs to quietly become a second dossier archive.

In July 2026, the Dutch Data Protection Authority published guidance on generative AI and the GDPR. It is relevant to organisations developing or deploying generative AI.

Choosing local AI, your own cloud environment or a managed environment is part of the design. You still need to assess processing, access rights, retention periods and supplier agreements. Where the model runs does not automatically make an application GDPR-compliant.

From AI pilot to production: deploy and keep testing

First, let the workflow run under controlled conditions. The agent prepares, an employee reviews and the existing process remains available as a fallback. Agree who handles problems and who can stop the workflow.

Keep the test set. Run the relevant tests again when you change the model, instructions, knowledge sources or an integration. An improvement on one dossier can break something on another.

This approach can also be used for the administrative preparation of legal or healthcare dossiers. The correct outcome, access rights and checks must then be defined with professionals from that sector. A successful accounting test does not establish suitability for another application.

For an Agentic OS, a first workflow like this delivers more than a working draft. You have a defined task, testable quality, an owner and a view of the work that remains. That gives you a basis for deciding whether to expand, adjust or stop.

Keep reading