Guide · AI assurance

A team cannot validate its own AI, and whoever holds the release knows it

Independent validation for systems already in front of reviewers, from engineers who have carried reportable systems through audit for 20 years.

What is already in production without a sign-off

53%have had an AI agent exceed the permissions it was given
47%had a security incident involving an agent in the past year
15%can say who owns most of their agents

Cloud Security Alliance, State of AI and Security Survey, April 2026. 445 practitioners. Sponsored by Zenity.

Three routes teams take, and the mechanism that stops each one

The builder is the reviewer

A sign-off exists. It has no independence, so it carries no weight with whoever owns the release. The pilot does not fail review. It never enters one, because nobody will put their name against a test the author wrote and graded.

It waits

No owner kills it and no owner ships it. Waiting has no event attached to it, so nothing forces a decision, and the agent keeps running in the meantime with whatever permissions it was given on day one.

An AI standards team gets stood up

This is the healthiest of the three, because somebody has decided the work needs governing. A standard with nothing auditing against it is a document, and the gap between the document and the running system widens every sprint.

What independent validation produces

A test, evidence, and a stated boundary.

  • A test written by someone who did not build the system. Graded against the standard your reviewers apply, in the form they ask for it, so the result is usable in the review itself, and not a second opinion nobody asked for.
  • Evidence produced while the system runs. An audit trail, versioned prompts and models, and evaluation results recorded as they happen. Assembling that after the fact is slower and it usually arrives past the point where it could change the outcome.
  • A stated boundary. Independent validation applies to systems we did not build. Where we built the thing, we are not the independent reviewer of it, and we say so before anyone asks. That boundary is what makes the first two worth anything.

When independent validation is the wrong spend

  • An AI ask with no workflow behind it. If nobody performs the task today, there is no baseline to validate against and no way to say whether the system is right.
  • Nothing shipping yet. Assurance answers a buyer with a system already in front of reviewers. Earlier than that, the work is design.

One figure worth carrying into the conversation: 60 percent of organizations report knowingly deploying untested code. That comes from Tricentis, surveying 2,501 respondents through Censuswide in April 2026. Tricentis sells test automation, so weigh it accordingly.

Start with a conversation

Bring the system waiting on a review and the question the reviewer keeps asking. You will leave with a read on what clearing it takes, including the parts we would decline.