
Participant recruitment is one of the most costly (and time-consuming) withdrawals from a research budget. If AI can generate findings by learning about your users, it could save a lot of money and generate study results in milliseconds instead of months.
Synthetic users are the broad name to describe AI’s role in mimicking participants’ attitudes and intentions. But the term is used rather loosely. We’ve earlier defined synthetic users as more of an umbrella term (like a genus) than a type (like a species). The genus encompasses digital twins, research-grounded, persona-based, and demographic-based synthetic users, and AI proto personas.
We’ve written about pro-synthetic and anti-synthetic user attitudes in the UX community, and we published a literature review of experiments with synthetic users. Much of the literature has reported numerous discrepancies between synthetic and human results, including reduced variance, misalignment of means/percentages, distorted correlations, inaccurate regression coefficients, and shallow experiential narratives. Of the different types of synthetic users, digital twins are widely believed to be more accurate than other types of synthetic user despite their accuracy being an open research question.
But what can be helpful is an understanding of how effective synthetic users could be in the context of UX research. UX research has numerous methods and deliverables. One of those methods, the retrospective UX benchmark study, uses the same attitudinal UX metrics as task-based UX testing (such as the SUS, SUPR-Q®, or UX-Lite®), but participants are asked to reflect on their prior experience (as opposed to being presented with tasks), and they answer open-ended questions about their experiences. This sort of study looks a lot like a typical market-research survey, where you collect participant demographic details, attitudes, and intentions. We collect a lot of data with this method, making it a good candidate on which to test the promise of synthetic users.
In this article, we describe our experiment comparing quantitative results of one of our industry benchmark surveys with data generated by two types of digital twins.
The Original Study
In May 2026, we conducted a retrospective benchmark study of four AI-based chat software products with 420 U.S.-based panel participants. This study included the metrics we typically collect in our standard UX and NPS study of consumer software.
There was a roughly equal gender split (52% female, 47% male). Respondents tended to be younger, with 65% under the age of 40. Participants were asked to reflect on their most recent experiences with the software and complete several questionnaires, including the NPS, SUS, UX-Lite, and TAC-10™. The AI-based chat products and sample sizes were:
● ChatGPT: 113
● Claude: 103
● Gemini: 101
● Grok: 103
Replication with Digital Twins
For these experiments, we assessed how well synthetic users could reproduce one of the key metrics collected in the retrospective study, the System Usability Scale (SUS). The SUS is one of the most widely used standardized questionnaires for assessing perceived usability, made up of ten five-point items varying equally in positive and negative tone. The composite SUS score is interpolated to a 0–100-point scale for which there are well-known methods for classifying scores as good or poor, including the Sauro–Lewis Curved Grading Scale (see the Appendix).
We used Claude Sonnet 5 to construct two types of digital twins based on each person’s responses to the original survey (n = 420). We fed the model each participant’s survey responses that included questions on demographics, chatbot use, brand attitude, likelihood to recommend, and open-ended comments about what they like and dislike about chatbots. Most importantly, we left out the participants’ SUS scores.
Claude constructed a 700-word max persona written in the first person in the participants’ voice explaining who they are, based on the “summary agent” approach of Park et al. (2024).
The persona was made up of six sections:
- Who I am
- What matters to me
- How I talk
- How I decide and answer
- What I would have to guess
- What is not known or inferred about me
We built the synthetic users with two levels of fidelity: summary twins or full twins. In our taxonomy, both of these types of synthetic users are digital twins because they were derived from individual respondents’ data. Summary twins were given the 700-word persona summary (no direct access to any of the original survey responses); the full twins were supplied with all the participants’ survey responses (except for their SUS ratings) in addition to the summary. We did this to see whether giving a digital twin access to more specific information about a person would lead to better predictions of the SUS.
Each twin was prompted to role-play as its participant and to respond to the SUS as its human would have. We then matched each twin with its human counterpart and analyzed the differences in SUS scores.
Results
Digital twin means were systematically higher than human means
Both the full and summary twins scored statistically significantly higher SUS scores than the human responses (all p < .05), ranging from about 2.9 to 7.8 points higher when computed by product (Table 1 and Figure 1). The differences by product between full and summary twins were not statistically significant (all p > .36). In other words, the twins were more like each other than they were like the humans on which they were based.
| Chatbot | Human Mean | Summary Twin Mean | Full Twin Mean | Summary – Human | Full – Human | Full – Summary |
|---|---|---|---|---|---|---|
| ChatGPT | 81.5 | 85.7 | 88.3 | 4.2 | 6.8 | 2.6 |
| Claude | 78.9 | 84.6 | 86.2 | 5.7 | 7.3 | 1.6 |
| Gemini | 79.0 | 81.9 | 84.2 | 2.9 | 5.2 | 2.3 |
| Grok | 78.4 | 83.2 | 86.2 | 4.8 | 7.8 | 3.0 |
| Average | 79.5 | 83.9 | 86.2 | 4.4 | 6.7 | 2.3 |
Table 1: Comparison of means and mean differences by product.
Figure 1: Graph of mean SUS by product and source with 95% confidence intervals.
More information led to more discrepancy in means
Contrary to our expectation, giving the full twins more evidence (the actual responses) on which to base their responses did not lead to better predictions. The summary twin SUS scores differed from human scores by 4.4 on average, while the full twin scores differed from human scores by 6.7 points, both statistically significant differences. But is that difference a lot?
There are a few ways to interpret it. First, Table 1 shows that that number of points represents a 4.4% to 6.7% difference on the SUS’s 100-point scale, which can be thought of as an average error. Or those could be rephrased as 95.6% and 93.3% accurate predictions, which doesn’t sound that bad as long as the decisions that need to be made with the data can tolerate that level of precision. It’s also possible that if products have lower SUS scores, then the accuracy of synthetic responses would be lower (something we plan to follow up on).
A second way is to convert the scores into grades. Table 2 shows the SUS grades (see the Appendix for the grading scale) for the means by product and source. The predicted SUS grade was off by at least half a letter grade for all products and in one case off by a full letter grade (Grok).
This illustrates how the upward shift in scores for digital twins distorted the typical interpretation of the SUS, especially for the full twins. A stakeholder looking at these human grades would interpret the human scores as pretty good but not fantastic. The same stakeholder looking at the full twin grades would think the products are all superb.
| Chatbot | Human | Summary Twin | Full Twin |
|---|---|---|---|
| ChatGPT | A | A+ | A+ |
| Claude | A− | A+ | A+ |
| Gemini | A− | A | A+ |
| Grok | B+ | A | A+ |
Table 2: SUS grades according to Sauro–Lewis Curved Grading Scale (see the Appendix).
Small but consistent distortion at the item level adds up
To understand what’s driving the differences in scores, we looked at responses to the ten individual SUS items. Both twin models agreed more with the positive-tone items and disagreed more with the negative-tone items (Figure 2). This could be an artifact of the way that LLMs are trained, with a tendency to be overly agreeable. The mean response to each SUS item (five-point Likert scale) was off by an average of about half a point. Because the distortion was systematic rather than random, the deviations added up across the ten items rather than canceling out.
Figure 2: Mean scores for SUS items by source with 95% confidence intervals (n = 420).
Human-twin correlations were significant but far from perfect
Figure 3 shows the correlations (with 95% confidence intervals) between the human and digital twin SUS scores. The human-twin correlations were lower than you would expect if the digital twins were faithfully reproducing the human scores (although significantly higher than the correlation of .20 reported by Peng et al., 2026). The correlation between the summary and full twins was very high. The difference of .03 in the correlations of summary and full twins with human data was not statistically significant.
Figure 3: Correlations between human and digital twin SUS scores (95% confidence intervals).
Digital twin data was less variable than human data
Figure 4 shows compression of the variability of the digital twin responses relative to the human responses, with the twins’ responses tending to converge more tightly around the mean. Summary twin means ran 2.9 to 5.7 points above human means; full twins ran 5.2 to 7.7 points above. The standard deviations of the individual ratings were 15.1 for humans, 12.6 for summary twins, and 11.6 for full twins.
Figure 4: Distributions of SUS scores by product for Human, Summary Twin, and Full Twin (dots were spread horizontally to separate ties).
Summary and Discussion
We generated digital twins by creating summaries based on data collected from 420 human respondents to a retrospective study on attitudes toward four AI products: ChatGPT, Claude, Gemini, and Grok. Two types of digital twins were created for each respondent: one based only on the summary (summary twin) and one based on the summary plus all the specific data used to create the summary (full twin). Because the goal of this research was to compare human with digital twin SUS scores, the human responses to the SUS were withheld from the data used to create the summaries and from the specific data provided to the full twins.
Our key findings were:
Digital twin means were reasonably accurate, but systematically higher than human means. Averaging across the four products, SUS scores for summary twins were 4.4 points higher than human means (95.6% accuracy); full twin scores were 6.7 points higher (93.3% accuracy). We found this pattern to be statistically significant for all four products, and it was driven by systematic acquiescence by the digital twins at the item level.
Whether these levels of accuracy are good enough largely depends on the research context. We are cautious about generalizing these results beyond the fairly high SUS scores in this one study. It’s possible that because the SUS scores were already very high, a ceiling effect pushed the digital twin means closer to the human means than would be the case if the human SUS scores were lower.
Human responses were more variable than the digital twins. As reported in the literature, we also saw lower variability of the digital twin–generated SUS scores. The standard deviations of the individual ratings were 15.1 for humans, 12.6 for summary twins, and 11.6 for full twins.
Having more information did not improve prediction accuracy. We expected the means for the full twins to correspond closer than the summary twins to the human data. That didn’t happen. At the product level, the means of the full twins were on average about 2.2 points higher than the summary twins (and even farther from the human means).
The inflation of SUS scores affected their interpretation. At higher SUS levels, the deviations of three to eight points across products (the range from Table 1 across both types of digital twins) are enough to change the products’ grades. The grades for the human scores ranged from B+ to A: good but not fantastic. For the full twins, the grades were all A+, leading to a very different interpretation of the perceived usability of the products.
Correlations between human and digital twin scores were significant but far from perfect. The correlations were 0.56 for human/summary twins and 0.59 for human/full twins (no significant difference between these two). The correlation between summary and full twins was 0.92 (significantly higher than their correlations with the human data).
Bottom line: We designed these digital twins to give them the best possible chance to accurately estimate held-out SUS scores relative to other types of synthetic users that are less grounded in specific human data. We were frankly surprised that deviations from human scores (ground truth) were consistently greater for the full twins (more information) than for the summary twins (less information). We were also surprised that this approach generated SUS scores that were within a few points of the human data.
This is the first in a series of similar experiments that we plan to conduct and report using different datasets (with more varied SUS scores) and different ways of creating synthetic users. Stay tuned …
Appendix: The Sauro-Lewis Curved Grading Scale
| SUS Score Range | Grade | Percentile Range |
|---|---|---|
| 84.1–100 | A+ | 96–100 |
| 80.8–84.0 | A | 90–95 |
| 78.9–80.7 | A− | 85–89 |
| 77.2–78.8 | B+ | 80–84 |
| 74.1–77.1 | B | 70–79 |
| 72.6–74.0 | B− | 65–69 |
| 71.1–72.5 | C+ | 60–64 |
| 65.0–71.0 | C | 41–59 |
| 62.7–64.9 | C− | 35–40 |
| 51.7–62.6 | D | 15–34 |
| 0.0–51.6 | F | 0–14 |
Appendix Table 1: The Sauro-Lewis curved grading scale for the SUS.



