i.AI's approach to evaluation

Impact and Evaluation Lead
20 June 2025
How do you know if an AI tool is actually making a difference?
At i.AI, we believe the only way to answer that is to test, measure, and learn — at every stage. Evaluating how our tools perform - and the difference they make - is core to our work at i.AI.
- Improve product functionality by learning what’s working in real-world settings
- Quantify impact, such as time savings, cost reductions, or improved accuracy
- Understand user experience and satisfaction
- Build the case for future investment
- Communicate transparently about how our tools work and what they can do to users and the public
What gets measured, gets done: Our core metrics
We have a standard set of metric themes which we track across all of our tools to help us ensure we are considering a broad suite of benefits. This also allows us to compare them to one another and begin to develop an understanding of where our strengths lie.
These metrics are:
- Efficiency - How much time the tool saves compared to the current process
- Quality - How accurate the tool’s outputs are, compared to a human performing the same task
- Satisfaction - How end-users rate their experience with the tool
- Delivery - Running costs, including infrastructure and support
- Users/Usage - Number of unique users or queries per month
- Size of the prize - The potential reach or value of the tool, i.e. the total addressable market
We won’t limit ourselves to only these metrics if we feel others would add value, however these six are the foundation for how we evaluate all i.AI tools.
One size doesn’t fit all: Tailoring evaluation to each stage
Our products go through several stages of development - Scoping, Incubation, Alpha, Beta and Scaling - and each of these calls for a different approach to evaluation.
Scoping/Incubation:
We start with a Theory of Change to map out the intended impact and the steps needed to achieve it. Alongside this, we test how accurate or relevant the tool’s outputs are through product evaluations.
Alpha:
We run formative evaluations, such as interviews or surveys, to gather feedback on the user experience. This information can provide feedback loops into the product team developing the tool. We also start capturing early data on time savings, usage, and satisfaction.
Beta:
We aim for more robust impact evaluations — including randomised controlled trials (RCTs) or quasi-experimental methods — to estimate the effect of the tool across our key metrics. Where possible, we also capture evidence to inform economic evaluations to assess Value for Money.
Note: We only test tools we have built ourselves. Our colleagues in the AI Security Institute lead on the evaluation of frontier models or advanced AI systems.
Up to the challenge
Evaluating AI products isn’t easy. Here are some common challenges we face — and how we try to solve them.
Challenge 1: The time is…. now?
Our tools are constantly being iterated, and pausing this development to run an evaluation on a static product is not always feasible or desirable. Advances in AI, senior requests or findings from User Research can all also mean swift changes to a tool’s direction.
If the product is still changing, how can we effectively evaluate its impact?
We might:
- Consider what can be done with the proportionate level of rigour at a speed which will mean we don’t miss out on the opportunity to get results
- Conduct a small scale evaluation, perhaps on part of the tool or only covering some of our key metric themes
- Stop and wait for a better time when the product is more stable, not being caught out by the sunk cost fallacy if we have already invested effort in evaluation design
- Consider having a ‘frozen snapshot’ of the tool for the evaluation, which can be merged with a live version later
We're currently thinking about the best way to combine findings from multiple small-scale evaluations through a product's development.
Challenge 2: What does good look like?
Determining what our baselines are before the tool is introduced requires us to consider what we should compare our tools to: The work of an expert? An average worker? What is the level of ambition for the tool, and what is realistic?
When estimating quality, if a human benchmark is the gold standard, i.e. is the ‘ground truth’ or 100% performance, it is impossible for a new tool to be considered ‘better’. Conversely, when we consider efficiency, AI often vastly speeds up the process compared to a human, which may be a larger difference for senior vs junior staff comparisons but those differences are not likely to be meaningful (how much do we care if the product is 100x or 150x faster than a human?).
Quality and efficiency are not generally well-measured in the public sector. In i.AI we are keen to set the bar high on how we are measuring our impact in order to support public sector AI adoption by highlighting what realistic impacts can be expected by using our tools.
What’s the best way to determine a product’s performance?
We might:
- Use a human-in-the-loop model where an ‘operator’ uses the tool and a ‘supervisor’ checks the output — to mimic real-world workflows where senior staff quality assure junior staff output.
- Collect data on a range of metrics (as mentioned above) so that we can see the breadth of impact the tool is having.
- Take the approach of non-inferiority testing, usually seen in clinical research when comparing a new intervention to an existing one, rather than a placebo. Can our product deliver as well as the existing process (within a defined margin), but 10x quicker?
Challenge 3: Comparing apples and pears.
Some of our tools meet a very specific need. However, others are quite broad-use, with different groups of users using them in very different ways. If the tasks users do with the product are very different, we lack consistency in what’s actually being evaluated.
How can we evaluate a product with such a diverse set of users?
We might:
- Evaluate the tool on the more defined use cases, where determining baselines and outcomes is simpler
- Conduct surveys to gather views on the tool, rather than a more robust impact evaluation
- Conduct lab-style research under controlled conditions.
Challenge 4: Controlling the control group.
This isn’t unique to AI evaluation, but is a challenge we face. We have had people in control groups borrowing treatment group members’ laptops so they can use the tool being evaluated! This does give us some encouraging feedback about how much people want access to our products, but it doesn’t help us run rigorous evaluations.
How can we avoid contamination of our control groups?
We might:
- Use a waitlist design (where access to the tool is staggered, meaning everyone gets it eventually, but some just have to wait a little longer),
- Randomise at the team level rather than individual, meaning treatment group members and control group members are not sitting next to one another within the same team, reducing proximity and therefore opportunity for sharing access to the tool.
Always learning
Ultimately, we’re doing this because AI tools can only be useful if they’re also usable, trusted, and genuinely impactful. Evaluation helps us get there - one test at a time. We are still learning our way when evaluating i.AI’s products, and how to best navigate the challenges we face. We will be sharing more blogs on our work over the coming months, so please do check back!
Useful resources
- There’s some really great guidance from Government on evaluating AI, like this addition to the original Magenta Book covering Guidance on the Impact Evaluation of AI interventions.
- Publica’s guide to Evaluating Digital Projects is also full of practical tips when approaching this work.
- We’ve been publishing our work like this: Consult Evaluation report - Scottish Government’s Non-surgical cosmetic procedures consultation and Consult Evaluation blog post
- You can read about Caddy, our copilot for Citizens Advice contact centres, including some details of its evaluation.