Measure the impact of your AI tool to check if it delivers value, meets user needs and justifies its cost. To do this you should consider:
- potential benefit (potential Return on Investment. ROI)
- quality (accuracy and completeness)
- efficiency (how it compares to human performance)Í
- user satisfaction (engagement, trust and experience)
- usage (adoption, new and return users)
- cost, savings and ROI
Find out how to run impact evaluations for AI products on GOV.UK.
How to approach impact evaluation
Make sure impact evaluation is included in the design of your product throughout its life cycle, for example:
- plan your evaluation strategy early so you can align delivery with continuous feedback
- include evaluation metrics in objectives and key results to make sure prioritise performance in delivery
- develop a Theory of Change to identify potential outcomes, risks and potential unintended consequences
- establish a baseline so you can compare information from before the solution was implemented
- design a flexible evaluation approach so you can adapt as your product changes
Find out how BIT evaluated the performance of AI chatbots for public service.
Measure impact during development
Use quick methods during development to understand potential impact and iterate as needed, for example, you can use:
- internal testing to evaluate model accuracy and efficiency
- small-scale pilots and user research to evaluate user satisfaction
- theory-based methods (such as contribution analysis) to understand model quality and performance
Find out about measuring impact during fast-moving projects.
Measure impact after release
After release you must do regular evaluations to understand overall impact make sure the model is continuing to meet user needs, for example, you can:
- get robust evidence of user performance using experimental methods (such as randomised controlled trials)
- compare outputs between those who use your solution and those who don’t, using quasi-experimental methods (such as difference-in-difference)
- identify a baseline and monitor change before and after release
- monitor new and return users to understand adoption
- monitor costs for maintenance and use to calculate ROI
- measure impact on different use groupsto understand if your product meets the needs of all your users
- measure impact on public attitudes and perceptions to understand how, when and why people use your product
Get support when needed
The AI Task Force is developing and embedding best-practise for AI evaluation across government. For support, email: etf@cabinetoffice.gov.uk.
Measuring the impact of AI solutions