↖︎ Vishal Singh
D3M · Quantifying Metrics · Case Study

How Many Questions Do You Need?

Factor analysis on 19,631 personality inventories, and what it costs a coffee company to ask ten questions instead of fifty.
§1 · The problem

Fifty questions is not a signup form

Bean & Basket wants to launch a Discovery Subscription — a rotating single-origin coffee delivered monthly, sold as a $45-a-year commitment. The marketing team has a theory about who will buy it: curious people. People who like trying things. The theory is plausible, and it is also useless until somebody can measure curiosity in a customer they have never met.

There is a well-validated instrument for exactly this. The International Personality Item Pool asks fifty questions and returns five numbers. It works. It has worked for decades. And no customer on earth is going to answer fifty questions to sign up for coffee.

So the real question is not does personality predict subscription uptake. The real question is the one every survey designer eventually faces: how much of the instrument can you throw away before it stops working? Ten questions might be tolerable at signup. Five might slip into a checkout flow. Fifty is a research protocol, not a product.

This case works that question end to end. We take fifty items answered by 19,631 people, recover the underlying structure with factor analysis, use it to target a coffee subscription offer, and then measure — in dollars — what gets lost when the instrument shrinks.

Read this before anything else

The personality data is real. The subscription offer is not. No one in this dataset was ever offered a coffee subscription. We generated the response outcome ourselves from a known statistical model, and we publish that model's exact parameters in §6 so you can check every number in this essay against the truth that produced it.

This is deliberate. Real targeting datasets do not come with an answer key, which makes it impossible to know whether a method worked or got lucky. Here we know. Every claim about what the short form loses is measured against a truth we control — and that is the only reason those claims can be trusted.

§2 · The sample

Nineteen thousand strangers on the internet

The data comes from an open-source personality test that ran online for years. Anyone could take it. Nobody was sampled, recruited, quota-matched, or paid. The raw file holds 19,632 responses; one row has a missing item coded as a zero, leaving 19,631 complete cases.

Twenty percent of respondents are under 18. Sixty-one percent are women. Fifty-six percent are outside the United States, and 37% describe themselves as non-English speakers despite answering an English-language instrument. This is not a customer base, and it is not a population.

Break out by
Who answered. Hover any bar for exact counts. The how they arrived view is the one to sit with: the test recorded its own referral source, which means selection into this sample is not merely suspected, it is measured.

That last variable is unusually valuable. Most convenience samples leave you guessing about who selected in. This one logged it. And the groups differ: respondents who arrived from a .edu address score about a third of a standard deviation higher on Agreeableness and roughly a third lower on Openness than people who clicked through from elsewhere on the test site.

What this does and does not license

Nothing here estimates a population parameter. The average Openness of this sample is a fact about people who found a personality quiz on the internet, and it transfers to no one. What does travel is the covariance structure — which items move together, and how strongly. That structure is what factor analysis extracts, and it is remarkably stable across samples that differ wildly in composition. We will exploit that, and we will not claim anything about levels.

Twenty-four respondents gave the identical answer to all fifty items. Another seven varied by almost nothing. Straightlining like this is the standard signature of someone clicking through to see their result. We leave them in — at 0.12% of the sample they change nothing — but you should know they are there, and a paid panel with a 3% straightlining rate would be a different conversation.

§3 · Structure

The correlation matrix already knows the answer

Before fitting anything, look at how the fifty items relate to one another. Below is the full 50×50 correlation matrix, with items in the order they appeared on the questionnaire.

Ordering Correlation
Fifty items, 1,225 unique correlations. Hover any cell for the two item texts. Five blocks of ten sit along the diagonal — that structure is visible before any model is fitted. Switch to sign-aligned to flip the reverse-worded items.

Two things jump out. The first is the block structure: five clumps of ten items along the diagonal, each internally correlated and largely independent of the others. That is the five-factor model appearing in raw data, unassisted.

The second is the checkerboard within each block. Half the correlations are negative. That is not a data error — it is deliberate. The instrument alternates between items keyed in opposite directions:

Positively keyedNegatively keyed (reversed)
I am the life of the party.I don't talk a lot.
I start conversations.I keep in the background.
I am always prepared.I make a mess of things.

Reverse keying exists to defeat acquiescence — the tendency to agree with whatever you are shown. If every Extraversion item were phrased positively, a habitual yes-sayer would score as an extravert regardless of personality. Alternating the direction means the yes-saying cancels out.

The practical consequence is that you cannot average these items as they stand. A respondent who is genuinely extraverted answers 5 to half the block and 1 to the other half; the mean is 3, indistinguishable from someone who answered 3 to everything. Items must be sign-aligned before they can be combined — and notice that factor analysis discovers which items to flip on its own, from the correlation signs, without being told the wording.

§4 · Retention

How many factors? The rules disagree

Factor analysis needs one decision from you that it cannot make itself: how many factors to extract. There are three standard ways to decide, and on this data they give three different answers.

Show
Eigenvalues from the 50-item correlation matrix. The Kaiser rule retains every factor with an eigenvalue above 1.0 — here, eight. Parallel analysis compares each eigenvalue against the 95th percentile from 100 matrices of pure random noise of the same size, a far more demanding bar; it retains seven. The scree elbow — where the curve visibly flattens — falls at five.

Kaiser says eight. Parallel analysis says seven. The elbow says five. Theory, and sixty years of personality research, says five.

The instinct is to trust the most sophisticated criterion. Parallel analysis is genuinely better than Kaiser's rule, which is known to over-retain on instruments this long. So should we take seven?

Before answering, look at what the extra factors actually contain. When we force a seven-factor solution, the Extraversion, Neuroticism, Agreeableness, and Conscientiousness blocks stay perfectly intact. The seventh factor picks up no items at all as its primary loading. And Openness splits cleanly in two:

Facet A — verbalFacet B — imagination
I have a rich vocabulary. (0.75)I am full of ideas. (0.71)
I use difficult words. (0.74)I have excellent ideas. (0.60)
I spend time reflecting on things. (0.24)I have a vivid imagination. (0.59)

So the sixth factor is not a sixth personality trait. It is Openness fracturing along a seam that personality researchers have argued about for decades — vocabulary and verbal display on one side, imagination and idea generation on the other.

Our first suspicion was that this was an artifact. Thirty-seven percent of this sample are non-English speakers answering an English instrument, and "I use difficult words" means something very different if the words are in your second language. That would produce exactly this split, for reasons having nothing to do with personality.

A clean hypothesis, tested and rejected

We split the sample and re-estimated the Openness block separately in each group. If the split were a language artifact, the vocabulary and imagination facets should be far less correlated among non-English speakers. They are not: r = 0.370 among English speakers and r = 0.351 among non-English speakers. The second eigenvalue of the Openness block is 1.32 and 1.25 respectively — essentially identical.

The split is real, and it is not about language. We report this because the obvious explanation being wrong is worth as much as it being right, and because a hypothesis you did not test is not a hypothesis you can dismiss.

We retain five. Not because the statistics demand it — they mildly favour more — but because the sixth factor is a facet of an existing trait rather than a new one, the seventh is empty, and five factors are interpretable, nameable, and connect to everything else known about the instrument. Retention is a modelling judgment informed by statistics, not a calculation performed by them. Defending that judgment out loud is the skill.

Five factors account for 45.5% of total item variance. Students trained on regression find this alarming. It should not be. The remaining 54.5% is item-specific: the particular wording of "I make a mess of things" captures something no other item captures, and none of that is trait variance we want. A factor model that explained 90% of item variance would be memorising the questionnaire, not measuring personality.

§5 · Rotation

Making the factors legible

Extraction gives you five factors that reproduce the correlations well but are, as they come out of the algebra, uninterpretable — every item loads a bit on everything. Rotation exploits a mathematical fact: the five-factor solution is not unique. Any rotation of the factor axes fits exactly as well. So you may as well choose the rotation that is easiest to read.

Varimax chooses the rotation that pushes loadings toward zero or toward one, so each item speaks mostly to a single factor. Here is the result.

Sort items Hide loadings below 0.00
Rotated loading matrix. Each row is one item, each column a factor. Blue is positive, red negative. Drag the threshold slider up to fade weak loadings — the simple structure emerges sharply around 0.30. All fifty items load highest on their intended factor, with no exceptions.

The recovery is clean: all fifty items land on the factor they were written for. That is unusually tidy, and it is worth naming why. These items were selected across decades precisely because they behave this way. Items that cross-loaded were culled long ago. You should not expect this from an instrument you wrote last Tuesday.

Not everything is equally strong. "I seldom feel blue" loads at just 0.32 on Neuroticism, and "I spend time reflecting on things" at 0.33 on Openness. These are candidates for removal — which becomes directly relevant in a moment, since we are about to remove forty of them.

One further note. Varimax constrains the factors to be uncorrelated, which is convenient but not true: in these data Extraversion and Agreeableness correlate at 0.34, and Neuroticism correlates −0.27 with Extraversion. An oblique rotation would let the factors correlate and fit the reality better, at the cost of a harder-to-explain solution. We use varimax here because the downstream model handles correlated predictors anyway — but that is a choice, and you should be able to say why you made it.

§6 · The outcome

Building an answer key

Everything to this point used real answers from real people. What follows does not. To ask what a shorter instrument costs, we need an outcome that personality actually drives — and no such outcome exists in this file. So we build one, and we build it in a way that makes the exercise checkable rather than merely plausible.

The construction has three steps.

First, a true trait. The factor model says each person has five latent trait values that the fifty items measure imperfectly. We draw each person's true traits from the model-implied posterior given their actual answers. These true values are not observable by any instrument — that is the point. The fifty-item score is a good estimate of them, recovering between 82% and 89% of their variance depending on the trait, which is roughly the reliability a fifty-item instrument achieves in practice.

Second, a response model. Probability of subscribing is a logistic function of the true traits plus age and country. Openness is the dominant driver, Conscientiousness matters because a subscription is a commitment, and Neuroticism works against it because variety is risk. Every coefficient was set by hand in numpy — no model was fitted to produce them, and nothing was tuned to make a result come out.

Third, a coin flip. Each person's response is a Bernoulli draw at their own probability. Nobody is deterministically a subscriber. The overall response rate lands at 8.37%.

truth.json — the full generating model
logit(p) = −2.60
+ 0.55 × Opennesstrue   (curiosity — the marketing team's hunch)
+ 0.32 × Conscientiousnesstrue   (follow-through on a commitment)
− 0.20 × Neuroticismtrue   (variety reads as risk)
+ 0.15 × Extraversiontrue    + 0.05 × Agreeablenesstrue
+ 0.15 × US   ·   age: <18 −0.55 · 18–25 −0.20 · 25–35 +0.15 · 35–50 +0.35 · 50+ +0.40
seed = 77  ·  n = 19,631  ·  realised base rate = 0.0837
The answer key, published up front. Effect sizes are deliberately moderate. A standard deviation of Openness multiplies the odds of subscribing by about 1.73 — real, but nowhere near deterministic. Personality that predicted behaviour more strongly than this would be fiction.
The one thing this design cannot show you

Because we generated responses from traits, personality is guaranteed to predict the outcome. Nothing here is evidence that personality predicts real subscription behaviour, and no result below should be quoted as if it were. What this design isolates is a narrower and more answerable question: given that a signal exists, how much of it survives a shorter instrument? That question is about measurement, and measurement is what transfers.

§7 · Targeting

Four instruments, one mailing list

Now the exercise. Bean & Basket will mail a Discovery Subscription offer. They can rank the list by predicted response using any of four measurement strategies:

InstrumentItemsWhat it costs the customer
OracleThe true traits. Unobtainable; included as the ceiling.
Full inventory50Fifteen minutes. Nobody will do this at signup.
Short form10Two items per trait. Plausible as an onboarding step.
Minimal form5One item per trait. Fits in a checkout flow.
Demographics0Age, gender, country. Already on file. Free.

The short forms take the highest-loading items in each block. Each instrument is scored, a logistic model is fitted on a random half of the list, and lift is measured on the held-out half. Everything below averages over 40 independent splits, because a single split moves the top-decile estimate by more than the effects we are trying to measure.

View
Held-out performance, averaged over 40 splits. In cumulative gain, the vertical axis is the share of all subscribers captured by mailing the top x% of the list; the diagonal is random mailing. In lift, the axis is how many times better than random each depth performs. Hover for exact values.

Mail the top 10% by the fifty-item instrument and you reach a group that responds at 2.34× the base rate. The oracle — with traits measured without error — reaches 2.44×. That gap is the cost of measurement error in a fifty-item instrument, and it is small: the long form captures about 96% of what perfect knowledge would buy.

The short forms do worse, and the ordering is what you would expect. But the size of the drop is the number this whole case was built to produce.

Top-decile lift by instrument, with the share of the fifty-item advantage retained. "Excess lift" is lift minus 1.0 — the part that beats random mailing, which is the only part worth paying for.
The headline

Cutting the instrument from fifty items to ten costs about 16% of the excess lift. Cutting to five items costs about 22%. Dropping personality entirely and targeting on the demographics already sitting in the CRM costs 45%.

Framed the other way: ten well-chosen items — one fifth of the instrument — deliver over four fifths of its targeting value.

Two details are worth pausing on.

The first is that the loss is far smaller than the reduction in reliability would suggest. The ten-item form recovers only 51–73% of true trait variance against the long form's 82–89%. That is a substantial measurement degradation, and it translates into a much gentler 16% loss in targeting. The reason is that targeting only needs to rank people correctly, and ranking survives noise better than estimation does. You do not need to know a customer's Openness score. You need to know whether they are in the top decile.

The second is a check we ran because the obvious strategy is not always the right one. We asked whether picking items by highest loading is actually the best way to build a short form, or whether it wastes coverage by choosing near-duplicates. So we compared it against 25 randomly chosen ten-item forms, and against a greedy rule that trades some loading strength for lower redundancy.

Two things we checked that did not pan out

Highest-loading selection reaches 2.117 top-decile lift. Random ten-item forms average 2.058 across 25 draws, ranging from 1.95 to 2.22. So loading-based selection does beat chance — but by roughly five points of retained lift, not the landslide you might expect. Any ten items from this instrument do most of the job.

The redundancy-aware rule scored 2.115 — indistinguishable from taking the top loadings. The clever idea bought nothing. We report it because a null result you ran is more useful than a clever idea you merely asserted.

There is one hazard in the short form worth flagging to any student who builds one. All ten items selected by loading strength happen to be positively keyed — "I start conversations," "I follow a schedule," "I am full of ideas." Not one reversed item survives. That form has no defence against acquiescence at all: a habitual yes-sayer scores high on everything. Our simulated respondents have no yea-saying bias, so this costs nothing here. Real ones do, and it would.

§8 · The money

Sixteen percent of lift, in dollars

Lift is a modelling metric. Nobody approves a budget in units of lift. So put costs on it.

A Discovery Subscription contributes $45 in annual margin. Reaching one person — printed mailer, sample sachet, postage, fulfilment — costs $4.20. Those two numbers set a break-even response rate of 9.33%.

The base response rate is 8.37%.

Notice what that means

Mailing the entire list of 19,631 people loses $8,515. The average customer is not worth reaching. This campaign is not a case where targeting improves a profitable programme — targeting is the only reason the programme exists at all.

The decision is therefore not whether to target but how deep to mail. Response rate falls monotonically as you work down a ranked list; profit rises while the marginal customer clears $4.20 and falls after. Move the sliders to see how the optimum shifts.

Margin per subscriber $45 Cost per contact $4.20
Campaign profit against mailing depth. Each curve is total profit from mailing the top x% of the list ranked by that instrument, scaled to the full 19,631. Dots mark each instrument's optimum. The dashed line is break-even. At the default economics the entire programme is underwater without a targeting model.

At the default economics, the fifty-item instrument earns $13,731 at its optimal depth of 28%. The ten-item form earns $11,25682% of the profit, giving up about $2,475. Demographics alone earn $8,293.

Note that the profit retention (82%) sits close to the lift retention (84%). We expected profit to degrade faster than lift, because profit is convex in response rate — errors near the top of the list are expensive twice over. The effect is there but small, well inside the noise across splits. It is a nice hypothesis that this data does not really support, and we would rather say so than dress up a two-point gap as a finding.

Now the question that actually decides the case. Is $2,475 a lot?

It depends entirely on what forty extra questions cost — and that cost is not zero. Longer forms mean lower completion rates, more abandoned signups, and a worse first experience of the brand. If asking forty additional questions costs Bean & Basket even a tenth of a percent of signups, the short form wins outright. The 16% is not a verdict. It is a price, and you cannot decide whether to pay it without knowing what is on the other side of the trade.

This is also where the sliders earn their place. Push the margin up toward $90 and mailing becomes profitable at almost any depth — the model stops mattering because everyone is worth reaching. Push cost toward $9 and only the very top of the list clears, so ranking quality becomes decisive and the gap between instruments widens sharply. How much your model is worth is a function of your economics, not your model.

§9 · Limits

What we would have to believe

Four assumptions carry this result, and each is worth stating as a claim that could be wrong.

Everyone answered. Every person on this list completed a fifty-item inventory. Real firms have this for a survey panel, not a customer base. Extending trait scores to unsurveyed customers requires imputing traits from observable behaviour, and that link is far weaker than anything measured here — correlations around 0.3 to 0.4 in the published literature. A signal that survives a shorter questionnaire may not survive being inferred from purchase history at all. That is a separate case study, and a more sobering one.

The long form is treated as near-truth. We defined true traits from the fifty-item posterior, which flatters the fifty-item instrument by construction. In the field even the long form carries error we have not modelled — mood, fatigue, socially desirable responding. Treat the retention percentages as an upper bound on how well the long form does, and therefore as a conservative estimate of the short form's relative standing.

The sample is not a population. 61% women, 20% minors, 56% non-US, self-selected from internet traffic. The covariance structure we extracted is stable enough to travel; the trait levels are not. Any claim of the form "our customers are more open than average" is unsupportable from this data.

Responses are synthetic. Stated three times now because it is the assumption most easily forgotten by the time a student reaches the profit chart.


Where this goes next

The natural sequel is the assumption that most needs breaking: drop the requirement that everyone answered. Give the class a bridge sample of a few thousand surveyed customers with transaction histories, ask them to predict traits from behaviour, score the unsurveyed base, and rerun this exact profit comparison. Our expectation — worth testing rather than assuming — is that imputed traits underperform the ten-item form badly enough to make the honest recommendation "just ask the ten questions."

The second sequel is uncomfortable and belongs in the same course. Everything here is the Cambridge Analytica architecture in miniature: measure traits on a consenting sample, extend to people who never answered, target accordingly. The mechanics are unremarkable and the failure was one of consent, not of method. A class that can build this pipeline should be made to argue about whether it ought to run — and should notice that the profit chart above is completely silent on the question.