↖︎ Vishal Singh
D3M Part II · Quantifying Metrics Ch. 22 — Cross-Price Elasticity Teaching Case
Orange Juice · 116 weeks · One retail account

What your rival's price does to your sales

Coca-Cola wants to know who is really taking share from Minute Maid. Four regressions will answer that. They will also hand you one number that looks like a strategy and is actually a rounding error — and telling the two apart is the whole skill.

Fig. 1Retail juice demand does not drift — it detonates. Each brand spends most weeks near a flat shelf price and then spikes five- to twenty-fold during a promotion. Nearly all of the information we will use to estimate competition lives in those spikes.
01The assignment

Pepsi just bought your biggest rival

You have been hired by The Coca-Cola Company. Their concern is Minute Maid, and it is specific: Tropicana now belongs to PepsiCo, private label is taking shelf space, and nobody in the room can say — with a number attached — which of those threats is actually costing Minute Maid volume.

What they hand you is not a market research report. It is 116 weeks of scanner data from one of their largest retail accounts: for each week, the units sold and the shelf price of four orange juice brands. Tropicana, Florida's Natural, Minute Maid, and the store's own private label.

The question you are being paid to answer is a causal one dressed as a descriptive one. Not "whose sales move together," but "if Florida's Natural cuts price ten percent next week, how many cases does Minute Maid lose?" That quantity has a name.

What you will be able to do

Read a 4×4 elasticity matrix and say who competes with whom; separate a brand's clout from its vulnerability; and — the part most analyses skip — decide which cells of that matrix you are entitled to build a strategy on.

02The raw material

116 weeks, four brands, one aisle

Before any modelling, look at what actually happened on the shelf. Tropicana holds the premium slot — it averages $2.90 against Minute Maid's $2.27 and private label's $1.71. But averages hide the mechanism. Every brand runs deep, frequent promotions, cutting as much as half off its own regular price, and volume responds violently.

Fig. 2Price and volume are near mirror images, and the promotions are not synchronised. Because each brand discounts on its own schedule, the four price series are almost uncorrelated (pairwise |r| below 0.26) — which is exactly the independent variation a regression needs to separate one brand's effect from another's.

That last point matters more than it sounds. If all four brands promoted in the same weeks, their prices would move together and no regression on earth could tell you which one moved Minute Maid's volume. Here the retailer's staggered promotion calendar is doing the identification work for us, almost accidentally.

Volume share is remarkably even — Tropicana 27.9%, Minute Maid 27.3%, private label 22.7%, Florida's Natural 22.0%. Revenue share is not: Tropicana converts its price premium into 34.5% of category dollars while private label takes only 16.5%. Two brands, near-identical volume, very different businesses.

Data quality — read before trusting anything downstream

Every sales figure is an exact multiple of 64. These are not raw unit counts; the series has been scaled or aggregated from case-pack data. It does not bias elasticities (a constant multiplier vanishes into the intercept in a log model) but it does mean "units" here are not shopper units.

There is no feature or display variable. In real supermarket data, price cuts arrive bundled with an ad in the weekly circular and an end-of-aisle display. With those omitted, every coefficient below absorbs the promotion package, not the price tag alone. Hold that thought — it comes back in §08 and it changes what you are allowed to claim.

One account, no seasonality controls, no competitor stores. Volume that appears here as new demand may simply be a shopper who drove past a different supermarket.

03A tempting shortcut

Why the correlation matrix is a trap

The instinct is to compute market shares and correlate them. Brands that steal from each other should have negatively correlated shares. Run it and the result looks decisive:

Fig. 3All six correlations are negative — which proves nothing. Four shares that must sum to one are mechanically forced into negative correlation regardless of whether the brands compete at all.

Every pair is negative, from −0.215 to −0.384. You could write a confident memo about universal cannibalisation. You would be describing arithmetic, not competition.

Shares sum to 1.0 by construction. If one goes up, the others must fall. That constraint alone generates negative correlations among four roughly equal shares of about −0.33 — which is essentially the entire range we observe. The correlation matrix here contains almost no competitive information.

To measure competition you have to model the thing a manager actually controls — the price — and let the quantities respond.
04The estimator

Why elasticity is a slope in log space

Elasticity is a ratio of percentage changes, and a percentage change is a change in a logarithm. So if you regress log quantity on log price, the slope is the elasticity — no post-processing, no evaluation point, no units. That is the entire reason the log-log specification dominates pricing work.

log qMM = β0 + βMM,MM log pMM + βMM,Trop log pTrop + βMM,Fla log pFla + βMM,PL log pPL + ε

The diagonal coefficient βMM,MM is Minute Maid's own-price elasticity and should be negative. The three off-diagonal coefficients are cross-price elasticities: positive if the brands are substitutes, negative if complements, zero if unrelated.

Use the explorer below to see both halves of that claim. Switch the axes to log and the curved demand relationship straightens into a line whose slope you can read directly. Then switch the model to control for rivals and watch the slope move.

Fig. 4 — interactiveIgnoring rivals' prices makes every brand look less price-sensitive than it is. Minute Maid's own-price elasticity moves from −3.56 to −4.28 once competitors' prices enter the model, because weeks when MM discounted alone are pooled with weeks when a rival discounted at the same time, and the rival's damage is silently charged to MM's own price.

This is the first honest lesson of the case: an elasticity is not a property of a brand. It is a property of a brand and the model you estimated it in. Change the controls, change the number.

05The result

The elasticity matrix

Run that regression four times, once per brand, and stack the coefficients. Rows are the brand whose sales respond; columns are the brand whose price moved. Sixteen numbers describe the competitive structure of the category.

Each cell below carries its own 95% confidence interval, drawn to scale. Cells whose interval crosses zero are hatched — the data cannot distinguish them from no relationship at all.

Fig. 5 — interactiveA third of the competitive matrix is noise. Four of the twelve cross-price elasticities have confidence intervals spanning zero — and two of those four sit in Tropicana's row, meaning Tropicana's volume is statistically unmoved by what Minute Maid or private label charge. Hover any cell for its interpretation.

Read the diagonal first. Own-price elasticities run from −3.10 (Tropicana) to −5.56 (Florida's Natural). These are large — a 10% discount roughly doubles Florida's volume — and they are large for a reason we will return to.

Then read across Minute Maid's row. MM's sales rise when Tropicana raises price (1.12), when Florida's Natural raises price (1.20), and when private label raises price (0.75). All three are statistically solid. Minute Maid sits in the crossfire of the entire category.

Now read down Minute Maid's column — the effect of MM's own price on everyone else. Florida's Natural loses heavily (2.36), private label loses (1.04), and Tropicana… 0.06, with a confidence interval running from −0.40 to +0.53. Minute Maid's pricing does not reach Tropicana at all.

06Structure

Clout, vulnerability, and the asymmetry that matters

Two summaries turn sixteen coefficients into a competitive position. Sum a brand's column (excluding the diagonal) and you get its clout: total damage it can inflict on others by discounting. Sum its row and you get its vulnerability: total damage others can inflict on it.

Cloutj = Σijij| · Vulnerabilityi = Σjiij|
Fig. 6 — interactiveTropicana occupies the only genuinely defensible position in the category: it can hurt others roughly 2.7 times more than they can hurt it. Minute Maid has the single largest offensive weapon in the aisle — clout of 3.47 — but pays for it with the second-highest vulnerability. Select a brand to trace who it hits and who hits back.

The strategically interesting finding is not a level. It is an asymmetry.

The core finding

When Tropicana cuts price 10%, Minute Maid loses 11.2% of its volume. When Minute Maid cuts price 10%, Tropicana loses 0.6% — a figure statistically indistinguishable from zero (p = 0.79).

Minute Maid is fighting a rival that cannot feel the punches. Every dollar MM spends discounting against Tropicana buys volume from Florida's Natural and private label instead.

That is a defensible conclusion, because it rests on the contrast between a coefficient with a t-statistic of 4.0 and one with a t-statistic of 0.27. The gap is not subtle. Tropicana's buyers are not choosing between orange juices on price — they are choosing Tropicana.

07The discipline

The finding that isn't

Here is where most consulting decks go wrong, and where this one earns its fee.

Look again at Minute Maid's row. MM's response to Florida's Natural is 1.202. Its response to Tropicana is 1.119. Florida's is bigger. The obvious sentence writes itself: Florida's Natural, not Tropicana, is Minute Maid's closest competitor — target Florida.

Now test whether the data supports it.

Fig. 7The ranking is not real. The gap between the two coefficients is 0.083 with a standard error of 0.425 — the 95% interval for the difference runs from −0.76 to +0.92, so the data cannot even tell you which of the two is larger.
Statistical vs practical significance

Both coefficients are individually significant. That is not the question. The question is whether they differ, and the formal test gives p = 0.846.

To detect a gap that small with 80% power you would need roughly 207 times more data — about 24,000 weeks of it. That is 461 years of scanner records for a single supermarket account.

The two failures are opposite and equally common. Practical without statistical: a difference this size would matter commercially if it were real, but the evidence cannot establish it. Statistical without practical: with a large enough panel you could eventually prove the 0.083 gap exists, and it would still be too small to build a promotional calendar on.

So the honest version of the recommendation is not "target Florida's Natural rather than Tropicana." It is: Florida's Natural and private label are both reachable; Tropicana is not. That claim survives the standard errors. The ranking between the first two does not.

08Limits

What this model still can't tell you

The matrix is internally consistent and the fits are respectable (R² between 0.58 and 0.71). It is still not a demand system you should hand to a pricing committee. Three reasons, in ascending order of seriousness.

The elasticities are too big to be about price

Own-price elasticities of −3 to −5.6 are far above what shoppers' price sensitivity alone produces. The missing variables are the ones the callout in §02 warned about: in supermarket data, a price cut almost never travels alone. It arrives with a feature ad and a display. What the model calls the effect of price is really the effect of the whole promotion.

The arithmetic implies demand that does not exist

Fig. 8The model's own arithmetic is the best evidence against reading it literally: it attributes 69% of the volume Minute Maid gains from a 10% discount to new category demand rather than to any rival. Orange juice consumption does not work that way.

Feed a 10% Minute Maid price cut into the fitted equations and MM volume rises 57%, while rivals lose only about 2,000 units combined. Roughly two-thirds of MM's gain comes from nowhere. In reality that "nowhere" is stockpiling — shoppers buying four cartons instead of one and not returning for a month — and store-switching, shoppers driving to this account because it advertised the deal. Neither is incremental category demand, and neither persists.

Prices are not randomly assigned

The retailer sets prices, and it does so with information we do not observe: expected traffic, holidays, competing stores' circulars, inventory. If discounts are timed to weeks when demand was going to be strong anyway, the estimated elasticity absorbs that timing. This is the identification problem at the heart of Part II, and no amount of additional control variables fully solves it — it takes an experiment, or an instrument.

The right way to use these numbers

Treat the matrix as a map of competitive reach — who can affect whom, and in which direction — not as a calculator for forecasting the volume a specific discount will deliver. The pattern of zeros and non-zeros is robust. The magnitudes are not.

09Deliverable

The memo to Coca-Cola

Tropicana is out of reach, and that is the finding. Minute Maid's price has no measurable effect on Tropicana volume, while Tropicana's price has a large one on Minute Maid's. Discounting to fight Tropicana transfers margin to Coca-Cola's shoppers and volume from Florida's Natural. The response to the PepsiCo acquisition is not a price war MM cannot win; it is brand investment aimed at the loyalty Tropicana already has.

Minute Maid's clout is real but pointed at the wrong targets. MM has the largest clout score in the category (3.47) and it lands almost entirely on Florida's Natural and private label — two brands with roughly a fifth of category revenue each. Use it deliberately, and be clear that share taken from private label is share taken from the retailer's own margin, which has consequences at the next line review.

Minute Maid's exposure is the number to manage. Vulnerability of 3.07 against Tropicana's 0.80 means MM's weekly volume is largely determined by other people's promotion calendars. The operational fix is defensive: know rivals' promotion schedule in advance and decide, in advance, which weeks are worth defending.

Do not act on the Florida-versus-Tropicana ranking. The difference is 0.083 with a standard error of 0.425. Anyone presenting it as a targeting rule is presenting noise.

What to do next. Obtain feature and display indicators from the retailer and re-estimate — that single addition will do more for these numbers than any modelling refinement. Then run a genuine price test in a randomly selected set of stores. Two hundred and seven times more history will not arrive; a randomised experiment can be run next quarter.

AAppendix

Full regression output

Four OLS regressions, n = 116 weeks each. Dependent variable is log units; regressors are the four log prices. Standard errors in parentheses. Diagonal cells (own-price elasticities) are shaded.

Multicollinearity is not a concern: variance inflation factors range from 1.03 to 1.12. Durbin–Watson statistics range from 1.61 to 1.85, indicating mild positive residual autocorrelation; Newey–West standard errors with four lags leave every conclusion above unchanged.