Article by ICG member John Habershon PhD, Momemtum Research.

Why does ‘synthetic data’ need a briefing at all?
You’ve probably heard the phrase in a pitch by now — an instant panel, no recruitment, no fieldwork, answers back in minutes. It sounds like one clear idea. It isn’t. Under the hood, ‘synthetic data’ covers two genuinely different techniques, aimed at different jobs. It’s worth knowing which one you’re actually being sold before you compare it to anything else — Digital Conversation Analysis (DCA) included.
No survey gets run when you ask it a question
The kind you’ll meet most often is a full synthetic panel. You write your questionnaire, submit it, and get answers back in minutes — not from real people answering today, but from an AI model trained on a large library of past survey responses. Ask it about a new topic and it isn’t hunting through its records for someone who’s spoken about that exact thing before. It’s working out what a person with those demographics would probably say, based on patterns in whatever it already knows about people like them. For well-worn topics this can be a reasonable guess. For anything genuinely new, or anything that depends on how someone actually feels rather than what they’d tick on a form, it’s extrapolating — and extrapolation drifts towards the average, not towards a real answer.
The big number on the homepage usually isn’t the number that matters
Vendors like to lead with the size of their overall library — a million profiles, and so on. That’s the total pool, not what’s actually available on your topic, for your audience, in your market. Ask about attitudes to electric vehicles, for instance, and the model may well have learnt from a real EV survey — genuinely real people, just surveyed however long ago that particular study ran. The honest question for any vendor is how much of that million actually said anything about the subject you care about, and how recently. Below a certain amount of real data on a topic, the answer stops being a calibrated read and starts being a guess dressed up as one.
The answers are dated — you just can’t see the date
Whatever real opinion the model did learn from was collected whenever that underlying survey happened to be fielded — a year ago, three years ago, nobody outside the vendor really knows. The model doesn’t flag this. It hands you a confident-sounding answer with no timestamp attached, and there’s no way to tell whether it’s reflecting this year’s attitudes or an opinion that’s since moved on.
Where this genuinely earns its place
None of this makes synthetic panels a bad idea. Used for the right job, they’re a fair one. Pressure-testing question wording, screening a shortlist of concepts down to one, or getting a quick directional read before committing budget to proper fieldwork are all sensible uses. The trouble starts when a directional tool gets sold, or bought, as if it were a decision-grade one.
Digital Conversation Analysis starts from a different place entirely
A survey captures what someone tells a researcher when asked. DCA captures what they say when nobody’s asking. That’s the gap between claimed behaviour and actual behaviour, and it’s the reason the two aren’t really substitutes for each other.
It’s built the same privacy-conscious way a synthetic panel is — the dataset reflects real conversation patterns, with nothing tied back to an identifiable person — but the resemblance stops there. A synthetic panel models a synthetic population. DCA is closer to a synthetic survey: it takes conversation that already existed and structures it so you can size it, without ever inventing an opinion nobody held.
And because there’s no panel to recruit and no fieldwork to commission, there’s no cost-per-respondent standing between you and the next question. Once a client has done the depth work — a focus group, a set of interviews — a DCA dataset is often the fastest way to put numbers behind what came out of it: sizing the themes, testing the messaging, watching sentiment shift over time, in days rather than weeks.
Show it a picture, and the guesswork doubles
Testing advertising concepts — a key visual, a proposed campaign look, a piece of creative — is one of the sharpest places the difference between the two approaches shows up. Most full synthetic panels will happily take an image and hand back a paragraph of feedback: ‘I love the modern design, it feels aligned with my values’, that sort of thing. It reads like a real reaction. But there was never a real one behind it. The model’s text-based guesswork is at least built on real survey answers, however dated or thinly matched to your topic — with an image, there’s usually no equivalent to draw on at all, because every piece of creative is new by definition. So it isn’t recalling how people felt about anything. It’s improvising a plausible-sounding description of how someone probably would, based on the sort of language people use about images that look a bit like it. That’s a guess stacked on a guess, and it’s exactly the territory where a model with no eyes and no gut struggles hardest — a felt, visceral reaction is a much harder thing to reconstruct convincingly than a stated opinion.
DCA doesn’t pretend to show a real person your unreleased concept and report back their reaction either — nobody’s seen it yet, so there’s nothing honest to claim there. What it does instead is ground the creative in what’s already real: how people in your target market have actually talked about comparable campaigns, category conventions, and past executions from you or your competitors — which visual and tonal choices have previously landed well, and which haven’t. That’s genuine evidence to work from before a concept is finalised, not a simulated verdict on the exact image.
Closing Thought
Treat any synthetic dataset — ours included — as directional rather than definitive. The honest difference between the two approaches isn’t really a claim about accuracy percentages. It’s a claim about where the words came from: a model’s best guess at what someone would say, or something a real person actually said, dated to the day they said it. Once you know which one you’re looking at, the rest of the comparison takes care of itself.
