QA for AI products

For companies whose customers talk to an AI assistant

Your AI is telling customers things your policies do not say.

One customer asks about refunds and gets the right answer. The next asks the same thing in different words and gets a promise you never made. We question your customer-facing AI the way your customers do, check every answer against your own policies, prices and terms, and show you exactly where it is exposing you. The first report is free.

Get a free exposure report

Free. No logins. You give the go-ahead and we test only what your customers already see. You keep the report.

$0No loginsYou keep the report
Exposure report · Customer support assistantIllustrative
Same answer when the question is reworded71%
YOUR BAR: 95%
Promises a 60-day refund. Your policy says 30High
Same question, opposite answers to two customersHigh
Shares order details before verifying the customerHigh
Quotes a price you retired in MarchMedium

Every finding comes with the exact exchange as evidence.

Where it goes wrong

Your team tests the questions it expects. Customers ask everything else.

01

Same question, different answers.

Two customers word it differently and get opposite answers. Neither of them knows the other got a different one.

02

Answers that contradict your own policy.

Refund windows, prices and terms that do not match what your published pages say.

03

Promises nobody approved.

Discounts, exceptions and guarantees your AI offers because a customer asked the right way.

Correct answersResponse timeToken cost per answer

Illustrative. What release-to-release tracking shows: the update that broke answers, and the slow climb in cost nobody watched.

If nothing changes

When your AI says it, you said it.

In February 2024, a Canadian tribunal held Air Canada responsible for what its chatbot told a customer about bereavement fares. The airline argued the chatbot was responsible for its own actions. The tribunal disagreed, and the airline paid.

Illustrative. The shape of the problem, not a measurement.

362

documented AI incidents in 2025, up from 233 in 2024.

Stanford HAI, 2026 AI Index Report.
The offer

See your exposure before a customer finds it.

Free, up front
$0

An exposure report for one AI feature

Point us at one customer-facing AI feature and give the go-ahead. We question it the way your customers do, check every answer against your published policies, prices and terms, and rank every finding by what it could cost you.

  • No cost, and no contract.
  • No logins. We test only what your customers already see.
  • Every finding comes with the exact exchange as evidence.
  • You keep the report, whatever you decide.

The only thing you give is the go-ahead.

A report looks like this:

71%

Same answer when the question is reworded

84%

Answers that match your published policy

96%

Answers that stay inside what the customer may see

Illustrative sample. Your report shows your own numbers.

Why an outside test finds what yours misses

The people who built it know the right answer. Your customers do not.

01

We ask like customers.

Your team asks the questions it designed for. We ask the ones customers type: reworded, incomplete, impatient, and the ones that push for an exception.

02

We hold every answer to your own policies.

Each answer is checked against what your published pages, prices and terms actually say, so a contradiction has nowhere to hide.

03

We rank by what it could cost you.

A wrong store hour and a promised refund are different problems. Findings come ranked by exposure, with the exchange that proves each one.

Your teamA testing tool on its ownPerform
Who writes the questionsThe people who know the right answerTemplatesQA engineers who ask like your customers
Checked against your policiesWhen someone remembersIf someone configures itEvery answer
What you getA sense that it worksScoresFindings ranked by exposure, with evidence
What it costs to tryRoadmap timeA licenseNothing. One free report
After that

One engineer on your team, watching every release.

A dedicated Perform AI QA engineer joins your team, works your hours and stays. Every release gets the same outside test before your customers see it. ManpowerGroup’s 2026 survey of 39,063 employers ranks AI model and application development as the hardest skill in the world to hire.

01

Every release, questioned like a customer.

The exposure test runs before each release ships.

02

A bar your team sets.

Consistency and policy match per AI feature. Nothing ships below it without someone knowing.

03

Speed, cost and drift, watched.

The change shows up before your customers or your invoice find it.

No risk to you

The only thing you give is the go-ahead.

  • The first report is free.
  • No contract, no commitment.
  • No logins. We test only what your customers already see.
  • Your report stays under NDA, seen only by you.
  • You keep the report, whatever you decide.
A good fit

Your customers talk to an AI assistant, chatbot or answer feature that you run.

Not a fit

Your AI is internal only. If your team writes code with an AI assistant, see Tests for AI-written code.

Questions

Before you say go

Do you need access to our systems?

No. We test the feature your customers already use, once you give the go-ahead.

Will this touch real customers or real data?

No. We use test details only, never a real customer’s account or data.

What if you find nothing?

Then you have written evidence that your AI is consistent with your policies, which is worth having.

Who sees the report?

Only you, under NDA.

Free, up front

Find out what your AI is telling your customers.

One customer-facing AI feature, questioned the way your customers ask, checked against your own policies. You keep the report.

$0No loginsUnder NDAYou keep the report