← All Fewer Moving Parts articles

Should We Automate This? · Part 3 of 4

How to Assess the Business Value, Risk and Readiness of an AI Project

A technically suitable AI use case can still make a poor project. Assess value, risk and readiness together before approving a pilot.

Richard SlaterFewer Moving Parts

Technical suitability is necessary for an AI project, but it is not enough to justify investment.

Before approving a pilot, an organisation should assess three connected questions:

  1. What operational or commercial value could the change produce?
  2. What happens when the system is wrong or unavailable?
  3. Can the organisation operate, support and stop it responsibly?

A use case may perform well in a controlled test and still create little net value, unacceptable risk or an operating burden the organisation is not ready to carry.

Separate technical capability from the project decision

An AI system may be able to extract requirements, draft a response or recommend a classification. That evidence answers whether the task is technically possible under the test conditions.

The project decision has a wider boundary.

It includes the current process, human review, integrations, information security, data rights, supplier dependencies, incident handling, user behaviour and the consequences of wrong or delayed output.

The difference matters because much of the cost and risk appears outside the model interaction.

Build the value case around the whole process

Start with the operational baseline established before the technology was chosen:

  • Demand volume and pattern
  • End-to-end elapsed time
  • Active working and review time
  • Waiting and approval time
  • Error, rework and exception rates
  • Cost of the present process
  • Customer or employee outcome

Then describe the mechanism by which the proposed system changes those measures.

If an AI assistant cuts drafting time, what happens to the released time? Does the same person complete more valuable work, does the queue shorten or does a new review step consume the saving?

Time removed from one activity does not automatically become a financial benefit. The organisation needs to identify where capacity can be used or which expense can be avoided without lowering service or control.

The value assessment should also include costs created by the new system:

  • Discovery, evaluation and implementation
  • Integration and data preparation
  • Model, platform and software charges
  • Security, privacy and legal review
  • Human review and correction
  • User training and process change
  • Monitoring, incident response and audit evidence
  • Supplier management and future migration

An early figure will contain uncertainty. Record that uncertainty instead of hiding it inside a single return estimate.

Describe risk as operational failure

Generic AI risk categories need to be translated into events in the workflow.

For a proposal assistant, those events might include an incorrect equipment count, an outdated price, an unauthorised discount, a missing contractual exclusion, confidential information leaving an approved boundary or a reviewer accepting unsupported text.

For each event, establish:

  • Who or what would be harmed
  • The likely scale of the consequence
  • How and when the event would be detected
  • Whether the output or action can be reversed
  • How recovery would work
  • The person accountable for the response
  • The evidence that must be retained

This produces controls that belong to the task. A rate should come from an authoritative pricing system. A discount may need rule-based validation and approval. Extracted facts may need links to their source. A high-consequence action may remain unavailable to the AI system.

The aim is not to claim that risk has been removed. It is to decide whether it can be controlled, observed and accepted by the right owner.

Treat human oversight as work

Human oversight can reduce risk, but it consumes time and can fail.

A useful review design defines:

  • Which cases are reviewed and why
  • Whether review happens before or after an action
  • The facts and sources shown to the reviewer
  • The reviewer's competence and authority
  • The expected review time
  • The escalation and disagreement route
  • The record created by the decision
  • The mechanism for pausing or reversing the system

Review burden needs to be measured during a pilot. If a system saves 30 minutes of drafting but adds 25 minutes of concentrated checking, the headline productivity gain is misleading.

Reviewer behaviour also matters. People may become less attentive when output is usually right, or they may repeat the whole task because they do not trust the system. Either outcome changes the value case.

The NIST AI Risk Management Framework Core asks organisations to consider human oversight, operator competence, expected benefits, financial and non-financial costs, error consequences and third-party components when mapping an AI system's context.

Test organisational readiness

Readiness means more than users agreeing to try a tool. It is the organisation's ability to remain accountable for the service.

Assess whether the project has:

Ownership

An operational owner should be accountable for the business outcome, not only for delivery of the technology. That person needs authority to change scope, require controls and stop the service.

Data and system access

The team needs lawful and practical access to the required information, along with a decision about which source is authoritative. Temporary access created for a demonstration is not an operating design.

Evaluation capability

Subject specialists must be available to identify good, poor and dangerous output. Their time should be included in the project plan and operating cost.

Support and incident handling

Someone must own integration failures, access problems, supplier outages, disputed results and user reports. The route should work during normal operating hours without relying on the original project team.

Change control and records

The organisation should be able to identify which system, prompt, data and policy produced a result, especially when those components change over time.

Exit and recovery

The team needs a safe manual route, a shutdown method and an understanding of how data, prompts, evaluations and workflow logic can be moved if a supplier changes price, terms or capability.

The UK Government's AI Management Essentials guidance provides a self-assessment of organisational processes for using and developing AI. It is intended to help organisations examine responsibility, risk management and communication; it does not certify a product or prove legal compliance.

Where personal data is involved, the ICO AI and data protection risk toolkit can help connect AI risks with people's information rights. The ICO states that this toolkit is being reviewed following the Data (Use and Access) Act, so organisations should also check its current guidance and seek specialist advice where needed.

Use one project decision record

Keep the decision evidence together so that one attractive claim cannot hide cost or risk elsewhere.

AreaEvidence to record
ValueBaseline, change mechanism, expected result, new work, new cost and uncertainty
RiskFailure events, affected people, consequence, detection, recovery, controls and owner
ReadinessOperational owner, data and systems, reviewers, support, records, supplier dependency and stop method

The record should identify evidence gaps and assign them to the pilot. It should also state assumptions that must be resolved before any live use.

Avoid turning the three areas into one weighted score. Some weaknesses are conditions, not points to be offset by a large estimated benefit. If the organisation cannot detect a harmful failure or identify an accountable owner, a strong productivity estimate does not fix the problem.

Make a real decision

The assessment should end with one of four decisions:

  1. Proceed to a bounded pilot. The expected value is credible enough to test, the main risks can be contained and the required people and systems are available.
  2. Narrow or redesign the scope. Remove high-consequence actions, limit the user group, require approval or separate tasks that need different interventions.
  3. Complete prerequisite work. Repair data, document policy, connect systems, establish ownership or design the review process before testing AI.
  4. Stop. The benefit is weak, the risk cannot be accepted or the operating burden outweighs the result.

Stopping at this stage is evidence that the assessment worked. It is much cheaper than learning the same lesson after a demonstration has become a commitment.

One final question exposes many readiness gaps:

If the system produces the wrong result at 8:30 on Monday morning, who notices, who decides what happens next and who can turn it off?

An organisation that cannot answer is not ready to operate the service, even if the model is ready to perform the task.

Part 4 will turn the remaining assumptions into a pilot designed to produce an honest continue, change or stop decision.