Press release

The adherence gap: Why grocery AI fails between the data and the order

Retail fresh categories now produce some of the strongest documented evidence of AI's operational impact anywhere in the food system, and the stakes are not marginal: US retail generated roughly 4 million tons of surplus food in 2024, worth about $26.9 billion, with produce the largest contributor by tonnage.

It’s clear that new AI models enable impact, but it is adherence that limits and amplifies it. If a produce manager doesn’t see the store’s reality reflected in their daily order recommendations, then trust—and value—burn quickly.

Poor adherence starts upstream

Low adherence is easy to diagnose as a store execution problem. The usual response is change management: train the teams, mandate compliance, hold districts accountable. That approach obscures what overrides represent.

Say a produce manager open the order guide for Thursday's delivery. The system recommends 14 cases of strawberries. They order 6.

The manager knows four things the system does not:

  • There are nine cases in the backroom, not the four on the screen. Saturday's truck came in over, and nobody adjusted the receipt.

  • Those nine came in soft. That’s two days of life, not five.

  • Last week's 2-for-$5 ad moved 40 cases. The model read that as appetite for strawberries, not as a price event.

  • Thursday's truck has been late three of the last four weeks.

On the dashboard, this scenario shows up as one override. In the department, it is four data failures: a receiving error, a shelf life assumption, a contaminated demand signal, and a fixed lead time.

The manager is correcting the system’s blindspots. So the override rate is often not a measure of organizational discipline, but rather a measure of how often the system is wrong about the store the manager is standing in.

Fresh data doesn’t follow the rules

Most replenishment systems need accurate, single numbers to run through their calculations. They assume the on-hand number is right, the lead time is fixed, demand is stable, and the item master is accurate.

Fresh departments are defined by variability. A system built around single-point inputs cannot represent all of it:

  • Perpetual inventory doesn’t reflect reality. Culls, shrink, trim, sampling, mis-scans, quality issues, and fresh-cut production pull product out of the department with nothing recorded.

  • Shelf life is a variable, not a constant. It moves with vendor, season, lead time, and quality.

  • Receiving rarely matches the order. Trucks run late, DCs short the order, a different variety shows up.

  • Product attributes drift. The case size says twelve; the box on the dock holds six. Item IDs do not always line up cleanly across POS, ordering, inventory, and supply chain systems.

  • The demand signal is contaminated. A markdown sale reads like full-price demand and inflates the next order. An out-of-stock reads like weak demand and suppresses it. Promotions, seasonality, weather, and local buying patterns change constantly.

This is the ordinary texture of fresh operations. While data cleanup can reduce the noise, it cannot eliminate uncertainty in a department where physical product is constantly changing and not every change generates a reliable system event.

Retailers often go through lengthy cleanup and integration processes, but a system designed around clean, deterministic inputs ends up confidently wrong in exactly the moments the store team needs it, and overrides creep up as trust degrades.

The solution: AI models need to be built for uncertainty

If overrides are a symptom of the data foundation, the fix has to be architectural. AI models built for fresh must quantify uncertainty and carry it through the decision process. They should:

  • Treat inventory and demand as a range of probabilities. Instead of assuming there are exactly 6 cases on hand and demand will be exactly 10, model how likely the other possible outcomes are given what was ordered, received, sold, culled, and pulled for production.

  • Continuously estimate shelf life per item. A live estimate should move with vendor, lead time, season, and quality rather than relying on a static field set during implementation.

  • Separate price effects from preference. A markdown sale is a different signal than a full-price sale, and an out-of-stock is not evidence that demand disappeared.

  • Model supply and constraints as ranges. A four-day lead time is not always four days, and an order for ten cases does not guarantee ten cases will arrive. Model the likelihood of those outcomes while respecting hard constraints like case packs and delivery schedules.

The practical effect is a system that understands the store more often, including on the bad days. Being right on those days is what earns the tenth, twentieth, and hundredth order without an override.

That makes adherence one of the most honest performance metrics in fresh replenishment. It is also one almost nobody puts in an RFP.

What this means for supply chain solution RFPs

Grocers evaluating AI-powered replenishment should ask vendors how their systems behave when the data looks like actual grocery data.

  1. Put adherence in the RFP, and define it precisely. The percentage of recommendations accepted without modification, across all live departments rather than pilot stores, broken out by department, with four quarters of trend. A vendor who cannot produce that number at scale is telling you something.

  2. Ask how the data foundation works. How is on-hand estimated between counts? How is shelf life derived and updated per item? How are markdowns, promotions, and out-of-stocks handled in the demand signal? What happens when a receipt does not match the order? These answers predict adherence better than any accuracy metric a vendor volunteers.

  3. Read adherence as a leading indicator. Shrink and sales move on a lag and are contaminated by weather, promotions, competitive openings, and supply disruptions. Adherence moves immediately and sits almost entirely within the system's control. If it is climbing, the financial results are coming. If it is flat and low, patience will not produce them.

Adherence in Afresh’s AI-native store ordering solution

Across more than 12,700 fresh departments running Afresh, average adherence to store-level ordering recommendations is 94%, spanning produce, meat, deli, and bakery in national and regional chains.

  • A Midwest regional chain, produce: 92% of recommendations accepted, with 100% of produce orders placed through the system.

  • A Western regional chain, chainwide produce: 93.4% acceptance, 95% of items carrying a recommendation, and average user satisfaction of 4.12 out of 5.

  • A specialty fresh-format retailer, meat: adherence climbed to 89% over a year, with improvement in every department.