AI Structured Data Case Study Featured Image

Does AI Actually Use Your Structured Data? We Ran a 20-Week Field Experiment to Find Out

Jordan P. FowlerSEO

Download the full study here:

Download the full pdf here AI and Structured Data Nimbus Case Study.

The Hypothesis

Hypothesis: If a business publishes structured, machine-readable data on its own domain, AI platforms will prefer that structured data over vague marketing copy when answering questions about the business.

Verdict: not supported. In a 20-week field experiment across eight AI platforms, using a fictional business with no prior footprint in any AI model, the structured data was discovered by crawlers within three days, cited seven times, and never once demonstrably used in an answer.

Schema markup, JSON data files, llms.txt, and API endpoints all rest on some version of that assumption. For several years it has quietly served as the foundation of AI visibility advice, largely untested.

The way the hypothesis failed tells you more about how AI answers questions about businesses than a success would have.

How do you test this cleanly? You build a company that doesn’t exist.

Here’s the problem with testing on a real business: if an AI gives a correct answer about an established company, you can’t tell where it came from. The website? A directory? A review? Something absorbed into the model’s training two years ago? Every source is tangled together, and the experiment proves nothing.

So we built Nimbus Ground Effects, a fictional specialty contractor in Argyle, Texas, with a fictional founder, a fictional installation history, a real website, and a real phone number. Because the company had never existed, no AI model had any prior knowledge of it. Every correct, specific fact an AI stated about this company came from something we published. There was no other possible source.

Then we built the measuring instrument:

Eighteen (18) deliberate contradictions planted between the marketing website and a nine-file structured JSON knowledge pack on the same domain. The website said “works with any surface”; the pack excluded concrete, pavers, and artificial turf. The website said “we love pool installations”; the pack prohibited installation within ten feet of a pool.
Tracer values in every data layer. There is a different founding year, different business hours, and a different phone number in each source, every combination unique. Any AI answer containing one of those values identifies which source it came from; the answer attributes itself.

We monitored eight AI platforms weekly for 20 weeks (ChatGPT, Perplexity, Copilot, Gemini, Google AI Mode, Google AI Overviews, Claude, and Meta AI) backed by server logs capturing every bot visit and manual testing sessions where we read full answers word for word. Every deployment date and file state in the study is verified against the site’s commit history.

What 20 weeks of data showed

Before the site existed, AI didn’t say “I don’t know.” All seven platforms monitored at baseline explained that the company wasn’t real and that the name referred to a fog machine sold by another manufacturer. The responses were detailed, well-organized, and invented. When authoritative data is missing, AI fills the gap.

Getting noticed was trivial on some platforms and impossible on others. A one-page site with zero backlinks, zero reviews, and no Google Business Profile was cited by ChatGPT and Perplexity within days. Google’s AI surfaces ignored it for two weeks while Googlebot crawled it daily. Two platforms never cited it once in five months.

Crawler traffic predicted nothing. The single heaviest bot on the domain, at times more than half of all automated traffic, belonged to a platform that cited the site zero times in 20 weeks.

The format works perfectly when handed over directly. We loaded the same files into a custom GPT as a control test: 18 out of 18 planted contradictions resolved correctly, with clean refusals on questions the data didn’t cover.

Published on the open web, the data was found and never used. Bots downloaded the pack files within three days. One retrieval agent fetched two files twice each inside a single minute. Across 20 weeks and eight platforms, none of the pack’s tracer values appeared in an AI answer. Every phone number reported was the website’s. Every set of hours was the website’s.

The safety information never got through, even with a sign pointing at it. Every eligibility rule and safety limit lived in one file. It was listed in the sitemap, downloaded by bots, and explicitly named in llms.txt as the dataset to check for qualification questions. We asked four platforms three qualification questions. Twelve responses; zero used it. That file was cited zero times in twenty weeks.

A live API endpoint changed nothing. Documented with an OpenAPI spec and advertised at the domain root with explicit instructions for AI agents: zero requests in twelve weeks.

AI visibility decayed on its own. With zero changes to the site, confirmed by commit history, one platform fell from 25 of 29 citations to 5 in a single week. Another forgot the company existed mid-answer and reverted to describing the fog machine.

The moment that captures the whole study

One platform told a hypothetical customer that regular pesticide spraying was “usually not a problem” while a public file on that company’s own domain specified a fourteen-day chemical suspension protocol, written for inhalation risk. That platform’s own crawler had downloaded the file ten weeks earlier.

The right answer was published, free, findable, signposted, and already fetched. The customer got told the opposite.

What this means if you run a business

It remains useful to publish accurate structured information. This process requires you to establish a definitive version of the facts. This information is effective when processed by artificial intelligence. If you do not provide these details, systems may generate incorrect information and attribute it to you.

Don’t assume your safety-critical or eligibility information reaches AI answers. In twenty weeks of controlled measurement, it never once demonstrably did.

Don’t judge AI visibility by citation dashboards. They count links, and links are a poor proxy for what AI actually says. Three of the seven citations our data received pointed at a file that didn’t contain the answer being asked about.

Don’t treat any of this as a one-time project. Recognition arrived on different timelines per platform, decayed without warning, and shifted while we changed nothing at all.

What we’re honest about

This is a rigorous experiment with a sample size of one fictional business. There was no untouched control company, our own weekly monitoring may have driven some of the crawling we observed, and the key consumption result rests on twelve hand-read answers plus twenty weeks of tracer monitoring. These results are strong enough to consider acting on, though not large enough to be final. The full limitations section is in the study.

The open question

Everything in this study was passive: files, discovery hints, and an endpoint published on a website, waiting to be used; the untested step is active. It involves formally registering the endpoint with AI platforms as a GPT Action or MCP tool, so the platform is told to call it rather than left to find it.

If that gets the constraint data into answers, it’s the path forward. If it doesn’t, AI skipping past business constraints is a deeper platform behavior that no amount of publishing will fix. That’s the next experiment, and this page is where its results will land.

What’s in the download

Download the full pdf here AI and Structured Data Nimbus Case Study

The full study runs 39 pages: complete methodology, the four-source tracer architecture, all 18 planted mismatches, week-by-week citation data, bot log analysis, verbatim AI responses from both manual test sessions, two figures, hypothesis outcomes, limitations, and references. There is no gate or email required. The link above goes straight to the PDF.

Shorter PDF versions if you want them:

Executive summary: the whole study in five pages
Plain-English edition: the full story with no jargon, written so anyone on your team can follow it

Moon & Owl Marketing built and ran this experiment between April and August 2026 to validate a product before building it. Questions about the methodology, or want to talk about what your business’s information looks like to AI right now? We are happy to talk.