LLM Product Data Is a Disaster

Brands Change Product Data Often
Brands provide retailers with scores of product details that change frequently. Common updates include new images, net weight modifications, discontinuation notices, and more.
Here's a widely-available food product with the brand’s facts clearly displayed on the front image, including the FDA-required attribute of net weight = 15 FL OZ.

The brand’s webpages show the same facts, in text and images: 15 FL OZ.
Retailers Often Get The Facts Wrong
The non-stop barrage of brand data is hard for retailers to manage. Most can’t keep up.
For example, Kroger displays a product version that the brand retired 4 years ago, with the wrong weight and an old UPC.

Here are links to 3 major retailers who are displaying incorrect and conflicting data for the same product.
ChatGPT Ingests Bad Retailer Data
Notice the 3 URLs above end with “utm_source=chatgpt .com“.
Those URLs and bad data came from ChatGPT’s response to a simple prompt:
What retailers offer “Outshine Creamy Coconut Fruit Bars” in 6 count,
and what is the product’s exact net weight?
ChatGPT could not provide the correct product weight.
ChatGPT also failed to resolve the conflicting data it presented.
Worse, ChatGPT Hallucinates The Product Facts
Making things worse, the LLM invented fake reasons for these discrepancies:

Hallucination #1: Retailer differences are “labeling conventions.”
Truth: The brand updated the product’s net weight years ago. Many retailers' updates have lagged.
Hallucination #2: “The exact measured package volume is 14.7 fl oz.”
Truth: The accurate net weight is 15 fl oz.
Hallucination #3: “Many retailers round” weight values.
Truth: Retailers do not round weight values. But they do display stale data.
We Use LLMs for Reasoning, Not Product Data
ChatGPT also hallucinates reasons for data discrepancies, presenting these reasons as facts.
Without industry-specific knowledge, LLMs can’t recognize stale data, can’t adjudicate conflicts, and invent facts.
How Foodgraph Sees When Product Data Is Bad
We currently curate 70 data sources to capture the breadth of U.S. CPG. Transforming multi-sourced data accurately — and keeping it fresh — is a special data challenge, different from selling or manufacturing goods.
Foodgraph's platform weaves together multi-sourced data + domain expertise + agentic systems into quality results, built on observable workflows that expect data chaos.

- Multi-Sourced Data must come from brands and retailers to capture the full U.S. assortment, including Private Label, major CPG products, and long-tail items.
- Domain Expertise brings the know-how to shape data accurately.
- Agentic Systems apply our knowledge to the data at scale.
Takeaway: You won’t find quality data in LLMs for matching, enrichment, analytics, or almost any grocery-related use case.
Send me a note. I’d love to hear about your tough challenges and discuss how we can support your work.
For more information on what we do, see these recent blog posts:
Now 70 Data Sources Strong & Growing (v54 release)
Our latest release curates 70 data sources, cataloging 1.8M+ products, up 55% YTD.
Private Label Isn't a Trend. It's a Takeover.
We ranked the top 5 grocers in our catalog for new PL products this spring. The numbers are astounding.
We asked LLMs what they “think” about us and others.
Warm Regards,

David Goodtree
Founder and CEO, Foodgraph
Discover additional articles, updates, and perspectives.
Book a demo today to learn how Foodgraph can help.


