Evaluating Consult: An AI Tool for Enhanced Public Consultation Analysis

Head of Impact and Evaluation
14 May 2025
Public consultations are a vital part of the policy process, but they can be overwhelming: a single consultation can receive thousands (or even hundreds of thousands) of responses, take a team of analysts months to analyse, and create time pressure which hampers analytical quality.
We developed a tool called Consult to help support this analysis, using AI to help identify themes and map them to responses. It's really important that Consult leads to good analysis as well as fast analysis, so we’ve been continually evaluating its performance. Today we published an evaluation with the Scottish Government, to see if Consult could streamline consultation analysis while maintaining high-quality insights.
This blog provides a summary of the headlines. You can read the full report here.
What Did We Do?
We tested our current version of Consult on a live consultation about the regulation of non-surgical cosmetic procedures. Consult works in two stages that broadly mirror human analysis: first, it reviews all the consultation responses to identify key themes for each question; and second, it reads through each response and maps it to relevant themes.

What Did We Find?
Consult was generally good at identifying the themes in a given consultation response.
When we compared the themes that Consult identified, to the themes identified by the human reviewers, we found:
- For 60% of responses, the themes identified were exactly the same
- Using an average F1 score (a common measure of accuracy) to credit partial matches (e.g. three out of four themes matched) gives a score of 0.76 out of 1. While there is no set definition of a ‘good’ F1 score, most sources suggest scores above 0.7 or 0.8 are good.
Identifying themes is inherently subjective, so even two human reviewers often disagree. In early stage internal testing, we found that two human reviewers agreed with each other 62% of the time – in this context, we think a 60% agreement rate is positive.
For future evaluations, we will use a double human review to create a human-to-human benchmark for the exact data we are working with.
Reviewers were more likely to add themes than remove them, and use the ‘Other’ label to indicate missing themes.
Reviewers could see the themes that Consult had identified for a given response, and choose to accept them, add new themes, or remove themes.
Overall, reviewers added themes 1,671 times, but only removed them 763 times (across a total of 5,354 question responses). This could suggest that Consult under-identifies themes, or that human reviewers have a bias towards adding themes when they are unsure.

As well as the main themes, both Consult and the reviewers could label responses as ‘Other’ to indicate a relevant response where the reason was not captured in the current theme list. Consistent with reviewers being more likely to add themes, reviewers were 17 times more likely to use the ‘Other’ label (793 times vs. 47).
In follow-up user research, our expert reviewers said that some of these ‘Other’ labels were unnecessary (e.g. they had forgotten or missed a theme in the list), but some were capturing missing themes. Given Consult found more themes initially, which were reduced as part of the theme sign-off process, we will explore whether different ways of running the sign-off process – and a more extensive list for mapping – could reduce dependence on the ‘Other’ theme.
Differences between Consult and the expert reviewers had minimal effects on the overall theme rankings.
We also looked at the overall ranking of themes as a proxy for policy implications, because policymakers often focus on the most prominent themes. Despite some differences in the specific theme mappings, the impact on the ranking of themes was generally small. Across all 6 questions:
- Themes moved less than one position, on average, between the Consult-identified theme ranking and the human-identified theme ranking
- Changes were often driven by a single outlier theme – each question had, at most, one theme that moved more than 2 places
- Movement was often in the lower-ranked themes, as themes with low numbers of labels were more affected by changes to the labels. This means that the top themes were often unaffected.
Reviewing themes is quick and frees up time for analysis
The median time taken to review a response was 23 seconds, with 77% of responses reviewed in less than a minute. Reviewers spent more time on longer responses, taking the time needed to assess them.

By saving reviewers time, they were more able to focus on the analysis:
“Use of AI incredibly useful…sped up this phase, got to the outcome sooner so we can get to the analysis and draw out what’s needed next”
The reviewers believed the process helped guard against their preconceptions, although they still valued an opportunity to influence the themes.
The reviewers appreciated the fact that the themes generated by Consult were solely grounded in the responses to the consultation, rather than any pre-conceived expectations of what themes were likely to appear (or that reviewers wanted to see).
“Takes away the bias and makes it more consistent…not projecting your own preconceived ideas and when lots of humans are also doing the same, it becomes less and less consistent”
However, the reviewers valued the ability to shape the themes, to make sure they captured valuable information. For example, reviewers split out a theme that captured ‘all protected characteristics’ into separate themes for each characteristic (although by the end, they found that these broken-down labels were not always used).
The theme sign-off process was challenging, and could benefit from better guidance.
The expert reviewers were surprised at how long it took to finalise the list of themes, requiring several iterations to narrow-down and edit Consult’s original longlist. However, after completing the theme mapping phase, reviewers felt that the original longlist was better than they had realised, and felt that they should have had more confidence in it.
“Looking back at the initial theme long list, having more confidence in the initial AI proposed themes…may have made the theme checking more logical”
We are looking at how we can improve this process in future, and give reviewers more guidance on how to refine the themes.
Overall, however, experiences of using Consult were positive.
“You’ve done grand…easy to do a sense check…saved us a heck of a lot of time”
What’s next?
These findings are really promising. The quantitative data shows that Consult is doing a good job at identifying consultation themes while saving substantial time for policymakers. Users also identified how Consult could improve the quality of consultation analysis, through more objective identification of themes and by freeing up time for deeper analysis.
However, the evaluation also highlighted some opportunities for improvement, which we’re going to keep working on as we iterate and improve Consult.
Over the next few months, we’ll be testing variations of Consult on more public consultations: trying to improve its performance even further, as well as focusing on the user experience for reviewers.
We’re also developing more features to help policymakers interrogate the findings and dig into the insights Consult provides. We hope to share more findings on that too, as we get there.