Sofpact field note

AI for Retail Decisions: When Simulation Is Evidence

Editorial retail floor plan with zones, flows and decision gates

A store layout change is one of the few retail decisions that resists clean testing. You cannot run the same week twice, the weather moves with the calendar, a competitor promotion lands mid-trial, and the staff who executed the change know they are being observed. By the time the numbers arrive, the honest answer to “did it work?” is usually “something changed, and the layout was one of several reasons”.

That is the real problem AI is being sold into. The pitch is that computer vision, simulation and generative design remove the uncertainty. They do not. Used carefully, they change where the uncertainty sits and make it easier to see. Used carelessly, they produce a confident number that survives challenge precisely because nobody can reconstruct how it was produced.

Three technologies, three different evidential weights

“AI for store layout” usually bundles three distinct things. They fail in different ways, and treating them as one category is the first mistake.

Measurement

Computer vision on existing camera infrastructure can count people, estimate dwell time at a fixture, detect queue length and identify blocked aisles. This is the most reliable of the three, because it observes something that actually happened. Its weaknesses are mundane and knowable: occlusion in dense periods, miscounting of staff and delivery traffic, camera angles that were installed for loss prevention rather than analysis, and calibration that drifts as fixtures move.

Simulation

Agent-based and flow models predict how shoppers would move through a layout that does not yet exist. The output is a projection, not an observation. Its quality depends entirely on the behavioural assumptions encoded in it — how agents choose routes, respond to congestion, abandon a queue or deviate towards a promotion. Those assumptions are frequently derived from a different store format, a different country or a synthetic population.

Generative design

Given constraints such as fixture dimensions, adjacency rules, fire egress and planogram requirements, generative tools can produce many candidate layouts quickly. This is genuinely useful for widening the option set beyond what a planner would draft by hand. It says nothing about which option is commercially better. A generated layout is a hypothesis with a rendering attached.

When simulation output counts as evidence

Simulation earns evidential weight under conditions that are easy to state and uncomfortable to meet.

  • Calibration against your own stores: the model reproduces observed behaviour in a layout you already operate, in the same format and trading pattern, before it is asked to predict a new one.
  • Backtesting on a past change: the model is shown a layout change you made previously, without the outcome, and its prediction is compared with what actually happened.
  • Directional rather than absolute claims: the model is used to rank options or identify a congestion risk, not to forecast a percentage uplift to two decimal places.
  • Sensitivity that is reported: you can see how the conclusion changes when footfall, basket mix or dwell assumptions move within plausible ranges.
  • A stated failure condition: the team has agreed in advance what result would count as the model being wrong.

Where those conditions hold, simulation is a reasonable input to a reversible decision. Where they do not, the output is a structured opinion. That can still be useful — structured opinions beat unstructured ones — but it should not be described to a board as analysis.

When it is not evidence

Three patterns recur. The first is a vendor model calibrated on a reference population that is never disclosed, producing an uplift figure for your estate. The second is a simulation validated only against the intuition of the person who commissioned it, which means it will be trusted when it agrees and discarded when it does not. The third is a pilot with no control: a layout changes, sales move, and the model is credited without anyone establishing what the comparable stores did over the same period.

The NIST AI Resource Center: AI RMF Core is explicit that measurement should use documented methods under conditions resembling deployment, and that generalising beyond tested conditions is a known limitation rather than an acceptable assumption. That framing transfers directly to retail: a model calibrated on a large-format store in one market is untested in a convenience format in another, and should be described that way.

The data-protection boundary most pitches skip

Counting people is not the same as recognising them, and the distinction carries legal weight. Anonymous footfall measurement, biometric recognition and any attempt to infer emotion or demographics from a face sit in materially different regulatory positions.

Under the European Data Protection Board: Guidelines 3/2019 on processing of personal data through video devices, video monitoring of a publicly accessible retail space engages transparency, lawful basis and data-subject-rights obligations, and the household exemption is construed narrowly. Repurposing cameras installed for loss prevention into an analytics estate is a change of purpose, not a technical upgrade, and needs to be assessed as one.

Where a system creates a biometric template capable of identifying an individual, the Information Commissioner’s Office: biometric data guidance on biometric recognition requires both a lawful basis and a separate special-category condition, with defined retention and review of the biometric reference database. The ICO has also set out its supervisory priorities in its Information Commissioner’s Office: AI and biometrics strategy, which is worth reading before commissioning anything camera-based.

In the EU, the European Commission: AI Act prohibits certain practices outright, including emotion inference in defined contexts and untargeted scraping to build facial-recognition databases. A retail analytics proposal that offers “shopper sentiment” from camera feeds needs legal review before it needs a business case.

Designing a test that can change your mind

Before commissioning any of this, write down the decision it is meant to inform. “Should we move the bakery counter?” is answerable. “Optimise the store” is not. Then specify the comparison: which stores act as controls, over what period, adjusted for what seasonal and promotional effects. Then state the threshold that would justify rolling the change out, and the result that would stop it.

The measurement layer usually deserves investment before the prediction layer. Reliable observation of what currently happens is reusable across many decisions and can be validated against till data, staff rotas and delivery schedules. A prediction engine sitting on unvalidated measurement inherits every error underneath it and adds its own.

Governance is proportionate here rather than heavy. The NIST: AI Risk Management Framework structure — Govern, Map, Measure, Manage — is enough: name who owns the decision, map what the system observes and infers, measure it against your own estate, and decide in advance what would make you stop.

What this cannot do

No amount of modelling will tell you whether a layout change is right when the underlying commercial question is unresolved. If the range, the pricing architecture or the service model is in dispute, a simulation will simply encode one side of that dispute and return it with a confidence interval. Nor will these tools resolve a decision that is genuinely about brand, colleague experience or long-term category position — they optimise what they can count, which is rarely the whole objective.

The right answer is sometimes to stop. If the measurement layer cannot be validated, if the model cannot be backtested against a change you have already made, if the data-protection position on camera use is unresolved, or if no one will own the outcome after the pilot, the responsible decision is to run a conventional trial and keep the money. A generated layout that nobody can defend is worse than a planner’s draft that somebody can.

Further reading: European Data Protection Board: Guidelines 3/2019 on video devices; Information Commissioner’s Office: biometric recognition guidance; NIST: AI Risk Management Framework; European Commission: AI Act.

See how a controlled AI Sprint tests one retail decision before it scales.