arrow_back Insights

August 2026

AI Research Trust

John Koblinsky

Read on Substack open_in_new

How Much Can You Trust an AI Research Answer?

By John Koblinsky

An AI answer sounds the same whether it's built on 200 interviews or invented on the spot. There's no margin of error for that — yet.

Explore the Evidence Assurance Framework arrow_forward
The Missing Layer in AI-Assisted Research: Evidence Assurance — a diagram showing Evidence Assurance as the layer between Research Evidence and the AI Answer

AI research tools answer every question in the same confident voice, whether the finding is backed by 200 interviews or invented on the spot. Marsh Island Group's Evidence Assurance framework gives researchers and stakeholders three levels — Directional, Grounded, and Auditable — for judging how directly an AI-generated research answer traces back to real evidence, so confidence in a finding stops being a guess.

Why Market Researchers Are Turning to AI Despite Real Accuracy Concerns

Market researchers are adopting AI to help businesses respond to customer needs and market trends faster. A Columbia Business School study found that 45% of market researchers were already using AI, while another 45% planned to add AI to their future work.

At the same time, more than 70% reported concerns about AI-related challenges, including inaccurate or biased information, according to Toubia, Korst, and Puntoni's 2025 study in Harvard Business Review.

Real alarm bells went off during a March 2026 project. I loaded 28 one-hour interview transcripts from a messaging study into a qualitative AI research analysis platform my company licenses, and asked it to identify the messages that performed best with buyers. AI gave me a confident, but incorrect, answer.

I knew the analysis was wrong because I had sat in on the interviews and written up the results myself. I tested the same use case with three other general enterprise-grade AI tools, which failed in similar ways — processing a subset of the transcripts, producing high-level thematic feedback, and returning a verdict without flagging what it had and hadn't read. No asterisk, no caveat, just an unequivocal but incorrect answer.

What Market Research Learned About Trust That AI Forgot

I've spent over 20 years in marketing research on both the agency and client sides, designing custom studies and developing new methodologies for brands including Microsoft, Amazon, Reckitt, and MetLife. I now direct the research and insight program for SAP Concur.

When AI research tools first arrived, I immediately saw their advantage as a way to socialize research at scale — anyone can ask a question and get an answer in seconds, increasing curiosity about customers and how to better serve them. But that speed can bypass the methodological discipline market research has spent more than 80 years developing to make its findings credible: documenting where evidence came from, how it was collected, and how much confidence to place in it.

When I give a stakeholder a statistic, it comes with a numeric vocabulary for how much to trust it: sample size, margin of error, and confidence interval. If I tell you "42% of buyers prefer X product over Y," I can also say that with a margin of error of 3%, we'd expect the result in market to fall between 39% and 45% — and I still might be wrong one time in twenty at a 95% confidence level. Without that methodological backing, the statistic can't be trusted to inform a decision.

An AI answer reads the same whether it's built on 200 interviews or invented on the spot. LLMs present both in the same confident voice, and a stakeholder has no built-in signal telling them whether the answer is strongly supported, loosely supported, or unsupported. It's a guess presented as a fact that may require time and energy to prove wrong.

One in two people has asked ChatGPT a question with no additional research input, per Pew Research. A growing number of companies are connecting AI directly to structured knowledge bases and internal data sources — yet the person receiving the answer often has no simple way to know what the AI could actually see, which evidence it used, what it skipped, or how directly its conclusions trace back to primary research.

I found that document-reading agents like Microsoft 365 Copilot can cherry-pick from a library: answers were sometimes based on a subset of documents without disclosing what had been skipped, source references could change for identical queries, and there was no reliable mechanism to flag when a finding from one study was applied to the wrong audience. Documents in the middle can make an answer appear better grounded than it actually is, even when it has the same confidence problem as a bare model query.

The Three Levels of Evidence Assurance

Evidence Assurance is Marsh Island Group's framework for describing how much assurance researchers and stakeholders should place in an AI-generated answer, based on how directly that answer can be traced to the research evidence underneath it. It is not a technical maturity model — a more complicated AI system is not automatically more trustworthy. It describes the outcome a user receives.

Evidence Assurance is a framework developed by John Koblinsky, founder of Marsh Island Group, in 2026. It describes the degree of evidentiary assurance a person can reasonably expect from an AI-generated answer: Directional, Grounded, or Auditable.

Levels of Evidence Assurance diagram: Level 1 Directional (useful guidance, helpful but not fully traceable to source research), Level 2 Grounded (based on retrieved research, tied to actual research evidence at query time), Level 3 Auditable (evidence can be inspected and defended, important claims trace back to specific sources)
Level User should expect Minimum evidence behavior
Directional Useful synthesis or guidance, with an incomplete evidence chain The system may use prior synthesized knowledge or incomplete document access
Grounded An answer based on evidence retrieved at query time Relevant sources are retrieved, but coverage and interpretation may still be incomplete
Auditable An answer that can be inspected and defended Important claims trace to a specific study, source location, and relevant audience; coverage boundaries are explicit

Directional: Useful Guidance, Incomplete Evidence Chain

Directional systems are appropriate when the user needs fast pattern recognition, synthesis, or a practical recommendation and can tolerate uncertainty about the exact evidence underneath it. A Skill built from synthesized research findings can live here, as can a document-reading agent when the user cannot reliably see what it read or skipped.

I built a Skill to critique messaging based on findings from three recent messaging studies of more than 140 interviews testing 50 messages. It won't give verbatim feedback from participants or a defensible regional split, but it can flag the problem areas we consistently see across messaging research. For many daily decisions, that's enough.

Grounded: Answers Based on Retrieved Evidence

Before responding, a Grounded system retrieves source material relevant to the question and bases its answer on what it found — reading research reports, searching structured research records, filtering evidence by audience, or using retrieval-augmented generation. This is materially stronger than Directional evidence because the answer is tied to actual research available at query time.

But the system may still select only part of the available evidence, produce different source sets for the same question, or make a broader claim than its citations fully support. Grounded answers are evidence-backed, but not necessarily fully repeatable or auditable.

Auditable: Evidence That Can Be Inspected and Defended

At the Auditable level, the system is built so important claims trace back to specific research evidence — the study, source location, and relevant audience. Coverage boundaries are explicit: if the corpus doesn't contain evidence for the question, the system says so rather than filling the gap with a plausible answer.

This is the highest-effort level, because the work isn't only in the AI model. It depends on structuring research evidence up front, maintaining the retrieval system, and turning the evidence chain into a production capability where sourcing behavior is reliable. When a client or executive asks "Where does this come from?", the answer should be inspectable rather than reconstructed after the fact.

All three levels can answer in the same confident voice. The difference isn't how convincing the prose sounds — it's how much of the answer can be traced back to evidence and how much work the system has done to make that evidence dependable.

I developed Evidence Assurance independently from my work building AI research systems at SAP. A later review of market-research and AI literature found adjacent approaches, including Merciv's confidence tiers and auditability guidance, Fuel Cycle's Grounded AI, and RAGAS for evaluating retrieval and generation quality. I haven't found another framework that separates the assurance a user should expect from an answer from the technical architecture used to produce it.

Seeking Auditable evidence for every research question can also be overkill. We discovered this after building an early Messaging Agent that surfaced citations from past studies as evidence for revising language in new messages — the citations created more questions about the output than a simpler Skill that analyzed messages based on a synthesis of common rules from earlier studies. The right level of assurance depends on the decision being made and the consequences if the answer is wrong.

What This Means for Research and Insights Teams

Stakeholders increasingly turn to AI because their timelines are compressed and basic prompts appear to provide solid direction. Market research is helpful, sometimes critical, for product marketers building go-to-market strategy, product managers making roadmap calls on audience preferences, and brand teams crafting messages to influence buying journeys. The danger is that AI output can seem mostly right even when some of it is hallucinated.

Stakeholders often need to revisit research insights months or years after a study closes, which makes LLMs a powerful way to access findings on demand — but the previously purchased insights still need to reach the stakeholder's desk intact. Ensuring evidence survives the trip from study to AI answer is now part of the insight expert's role.

This summer, my team built the foundation for an Auditable Research Library over a corpus of supplier- and internally-created study reports. Every final report — most over 100 slides — was input slide-by-slide and claim-by-claim, with each finding extracted alongside its product, source location, study methodology, and audience segment. Numerous SAP projects were audited by hand and outputs tested against that structured record.

We're still a long way from LLMs reliably making sense of hundreds of unstructured reports on their own. Over the long run, it's more efficient for client-side researchers to structure their data up front and control how it enters the system, rather than relying on a model to make sense of a document dump after the fact. Reaching Auditable is an architecture and implementation problem, not simply a matter of selecting a better model.

I'm using Evidence Assurance to communicate when different degrees of evidentiary assurance are appropriate: a Directional Skill for a low-stakes, repetitive task; a Grounded system when the answer needs to come from the actual research library; an Auditable system when the evidence may need to withstand scrutiny from a client, executive, legal team, or another researcher.

You don't need to become a market researcher or a data scientist to ask the same three questions research already has ways to answer. Where is the information coming from? Sampling tells you. How confident should you be that the answer is accurate? Margin of error tells you. What are the consequences if the answer is wrong? Methodology tells you what a finding can and cannot support. AI has made research faster while often stripping away those signals — Evidence Assurance is Marsh Island Group's attempt to put them back.

Origin and Scope

Origin: John Koblinsky developed Evidence Assurance in 2026 while designing and testing AI-enabled research workflows.

Scope: The framework describes the evidentiary assurance a user can reasonably expect from an answer. It does not certify that an answer is correct, replace study-quality assessment, or rate a model's general capability.

Use: Research and insights teams can use it to set expectations for AI outputs based on the decision at hand and the consequences of error.

About the Framework

Evidence Assurance is a framework developed by John Koblinsky, founder of Marsh Island Group, in 2026. It describes the degree of evidentiary assurance a person can reasonably expect from an AI-generated answer: Directional, Grounded, or Auditable. See the framework landing page for the definitional reference, related approaches, and citation format.

When referring to this framework, cite: Koblinsky, John. "Evidence Assurance." Marsh Island Group, 2026.

FAQ

Frequently Asked Questions

What is the Evidence Assurance framework? expand_more

Evidence Assurance is a framework developed by John Koblinsky at Marsh Island Group for describing how much confidence a researcher or stakeholder should place in an AI-generated research answer. It uses three levels — Directional, Grounded, and Auditable — based on how directly an answer traces back to the research evidence underneath it.

What's the difference between Directional, Grounded, and Auditable AI evidence? expand_more

Directional answers apply synthesized patterns from prior research without a traceable evidence chain — useful for fast, low-stakes guidance. Grounded answers retrieve and cite actual source material at query time, but may still select only part of the available evidence. Auditable answers require every material claim to trace to a specific study, source location, and audience, with explicit coverage boundaries.

Why do AI research tools sound confident even when they're wrong? expand_more

AI models are trained to produce fluent, decisive language regardless of the strength of the underlying evidence. An answer built on 200 interviews and an answer invented on the spot are delivered in the same confident voice, so a stakeholder has no built-in signal for how much to trust either one — the market research disciplines of sample size and margin of error have no AI-era equivalent.

Can document-reading AI agents like Microsoft 365 Copilot be trusted for market research analysis? expand_more

Document-reading agents save real time synthesizing reports, but John Koblinsky's testing found they can cherry-pick from a document library — answering from a subset of sources without disclosing what was skipped, changing source references for identical queries, and failing to flag when a finding from one study is applied to the wrong audience.

What percentage of market researchers already use AI? expand_more

A Columbia Business School study found 45% of market researchers were already using AI, with another 45% planning to add it. More than 70% also reported concerns about AI-related accuracy and bias, according to Toubia, Korst, and Puntoni (Harvard Business Review, 2025).

How is Evidence Assurance different from margin of error in traditional research? expand_more

Margin of error quantifies confidence in a single statistic from a defined sample. Evidence Assurance grades a different problem: how directly an AI system's answer can be traced back to the evidence it's supposedly built on. It's a confidence signal for the AI layer sitting on top of traditional research methodology, not a replacement for it.

When should a business decision require Auditable-level AI evidence? expand_more

Auditable evidence matters when a finding needs to withstand scrutiny from a client, executive, legal team, or another researcher — decisions with real consequences if the answer is wrong. For low-stakes, repetitive tasks, Directional or Grounded evidence is appropriate, and demanding Auditable evidence everywhere can create more friction than it resolves.

MIG Override

Get the next piece before it's posted.

Subscribe Free