Skip to content
فهد النعيميFahad ALNaimi Entrepreneurship, e-commerce and artificial intelligence
All articles

Before Releasing an AI Assistant, Test the Cost of Its Errors

A brass balance weighing one heavy warning tile against several small blocks to illustrate unequal error consequences
In this article

Do not approve a new AI assistant simply because it answers a higher percentage of questions correctly. Compare it with the current version against three separate conditions: no prohibited failures, lower expected loss under the real workload, and acceptable service quality. A wrong opening time, an unauthorized discount and disclosure of another customer’s information should not carry the same weight just because an accuracy report counts each as one error.

This is a release decision framework for founders and operations leaders, not a staffing model for the review team. The NIST AI Risk Management Framework supports incorporating trustworthiness into system design, use and evaluation. The calculations and acceptance rules below are a proposed management method, not a formula prescribed by NIST or a compliance certification.

Define the business decision before the test

Start with the assistant’s authority. Does it explain a published policy, recommend an offer, or execute a refund? Separate factual correctness from permission to act. A fluent answer can still fail if it promises compensation the assistant cannot authorize. A correct referral to a person is not an error when the request genuinely falls outside its authority.

For each test case, record the input, permitted context, acceptable outcome, prohibited action and reason for the judgment. Do not demand one exact reference sentence when several responses would be valid. Have the process owner resolve disputed labels before comparing versions; otherwise the scoring criteria can quietly shift to favor a preferred model.

Stratify the test, then restore production weights

Sample ordinary requests after removing unnecessary personal data, and add a concentrated set of high-consequence cases. Oversampling rare cases helps expose weaknesses, but it does not make those cases equally common in production. Estimate each category’s production share from operating records, not from a test set deliberately designed to emphasize risk.

Within each category, vary Arabic and English, dialect, wording and missing information. The approach to testing the Arabic your customers actually use helps build realistic inputs. Keep a holdout set that was not used to tune instructions, and refresh it with genuine failure patterns so passing does not become memorizing the test.

A hypothetical comparison: accuracy points the wrong way

Suppose a business tests two versions on 1,000 cases: 900 ordinary inquiries and 100 sensitive commercial approvals. Version A makes 18 ordinary errors and six approval errors. Version B makes 27 ordinary errors and one approval error. Overall accuracy is 97.6% for A and 97.2% for B. Every number in this example is hypothetical, not a result attributed to Fahad ALNaimi or any particular company.

Assume monthly production has 10,000 cases, of which 95% are ordinary and 5% are sensitive. Also assume an average estimable loss of QAR 10 per ordinary error and QAR 1,000 per approval error. Apply each category’s error rate to its production volume:

  • A: 9,500 × 2% × QAR 10 + 500 × 6% × QAR 1,000 = QAR 31,900 estimated monthly error loss.
  • B: 9,500 × 3% × QAR 10 + 500 × 1% × QAR 1,000 = QAR 7,850.
  • If monthly operating costs are QAR 1,000 for A and QAR 2,500 for B, the totals become QAR 32,900 and QAR 10,350.

The estimated difference favors B by QAR 22,550 despite its lower overall accuracy. Avoid counting rework twice, once inside error loss and again in operating cost. This is not a realized saving: it is a decision estimate supporting a limited trial, conditional on assumptions that still need validation.

Small samples are evidence, not certainty

One incorrect approval out of 100 does not establish that the true error rate is 1%. A larger sample or different customer mix may change the result. Show case counts and error counts alongside percentages. Test scenarios in which approval losses double or the workload mix changes. If a modest assumption change reverses the decision, strengthen the evidence before committing.

Run both versions on the same cases under the same conditions, recording the knowledge base, instructions and tools used. If responses vary between attempts, repeat selected cases and record important worst-case failures and their frequency; do not select only the best attempt. Identify gaps in coverage, such as a new enterprise customer type or an updated commercial policy.

Some failures cannot be offset by savings

Define release-blocking events in advance, such as accessing information outside authorization or executing a prohibited action. An occurrence requires root-cause correction and retesting even when average estimated loss looks attractive. Observing none does not prove impossibility, so access and execution permissions must also be constrained outside the model.

Keep this gate separate from service thresholds: response time, appropriate referral rate, unsupported answers and error rate in each language. A version that reduces risk by forwarding every request to a person is not necessarily a business improvement. Check the economics of human review queues before accepting increased referrals, while keeping staffing capacity distinct from the release-quality decision.

Make approval reversible and auditable

Document the two versions, test-set date, production weights, loss assumptions, prohibited failures and thresholds approved by the process owner. Begin in shadow mode without executing actions. Progress to a small share of low-risk cases with a working rollback path, and require review independent of the people tuning the assistant before expanding.

The practical decision is straightforward: reject a version that fails a prohibited-action gate; defer when evidence is too weak relative to potential harm; and permit a controlled trial only when weighted loss improves without unacceptable service deterioration. The objective is not an impressive accuracy score. It is a release decision you can explain, challenge and reverse.

اقرأ المقال بالعربية

Back to article top
  • Artificial intelligence
  • Digital platforms
  • Strategic partnerships

Building opportunities for Qatar and the Gulf

Collaborate & contact

To discuss an opportunity, share the partnership idea, your goal and the proposed timeline.

Contact on WhatsApp