The standard script in tasting rooms and vape shop sampling bars goes like this: offer a customer a flight of flavours, cap it at five to seven, and let them choose. That number comes from a well-worn reading of choice-overload research — jam studies, chocolate studies, the general sense that more options produce more paralysis than purchase. But flavour sampling is not jam. The options are chemically distinct, the sensory system fatigues in specific and predictable ways, and the customer is not comparing labels on a shelf. They are comparing internal states. When you actually track where decision quality collapses in a flavour flight, the cliff shows up later than seven and earlier than most people expect: somewhere around eleven samples.
Why the jam-study number does not transfer
Iyengar and Lepper's 2000 field study at a Draeger's supermarket in Menlo Park is the anchor here. Shoppers who saw a display of 24 jams sampled more than those who saw six, but only 3 percent of the large-display group bought, against 30 percent of the small-display group. The finding is real. The extrapolation is not.
Jam is a single-modality product. Strawberry and apricot differ along a handful of dimensions — sweetness, acidity, texture, seed presence. A taster can hold four or five of them in working memory simultaneously and rank them. Flavour concentrates for vaping differ along at least three independent axes: base profile (fruit, dessert, tobacco, menthol), intensity, and the specific aromatic top notes that shift with wattage and airflow. Two blue raspberry variants from different manufacturers can be more dissimilar in mouthfeel than a blue raspberry and a mango. The comparison set is not a line; it is a matrix.
That changes the arithmetic of cognitive load. George Miller's 1956 paper on the magical number seven described chunks, not items. A jam is roughly one chunk. A flavour sample evaluated across three axes is closer to three chunks, and a taster holding eleven of those is managing something like thirty-three active comparisons. The overload arrives later in item count because each item is doing more work — but when it arrives, it arrives harder.
The eleven-sample inflection
Where does the flip actually occur? The pattern that shows up in tasting-room observation, and that maps onto the dual-process literature, is a two-stage curve.
From samples one through roughly six, tasters are in what Kahneman would call System 2 territory. They are deliberate, they take notes, they ask about nicotine strength and PG/VG ratio, they re-taste. Decision quality rises with sample count because the comparison set is genuinely informative. This is the phase the jam study describes, and capping at seven is defensible here.
Between seven and eleven, something shifts. Tasters stop re-tasting earlier samples. They begin anchoring on the most recent or most intense sample — a recency and salience effect that Kahneman and Tversky documented in judgment under uncertainty. Verbal reports shorten. "That one's good" replaces "that one's sweeter than the third but less sharp than the fifth."
Past eleven, the reports collapse into binary. Tasters either default to a familiar profile or pick whatever they tasted last. The information content of the flight has stopped accumulating and started degrading. Eleven is not a magic number in the numerological sense; it is roughly where the working-memory budget for a three-axis comparison runs out for most people under typical tasting-room conditions — standing, time-pressured, with aromatic fatigue accumulating across the session.
Variable-ratio reinforcement and the tasting-room trap
There is a second force at work that the jam study does not capture, and it is the one that makes flavour sampling genuinely different from retail choice.
Flavour sampling is a variable-ratio schedule. Most samples are competent. A few are extraordinary — the one-in-fifteen profile that hits exactly the note the taster has been chasing. Skinner's work on intermittent reinforcement established that this schedule produces the most persistent responding, and the effect is not confined to animals pressing levers. It is why a taster who has had two mediocre samples will take a twelfth. The reward is uncertain, the cost is low, and the previous sample did not extinguish the expectation.
This is where the eleven-sample finding has practical teeth. The variable-ratio schedule does not stop at eleven. It keeps generating the impulse to continue. What stops is the taster's ability to evaluate what they are tasting. The result is a customer who takes fifteen samples, experiences peak reward on sample nine, and leaves with sample fourteen — the one that happened to be last, not the one that was best. The reinforcement structure and the cognitive limit are pulling in opposite directions, and the cognitive limit loses.
Loss aversion compounds it. Once a taster has invested eleven samples, abandoning the flight to buy the sample they already liked feels like writing off the sunk cost of the remaining four. Tasters routinely push past the point of useful information to avoid "wasting" the flight they have already paid for in time and attention.
What this means for how flights get built
The practical implication is not to cap flights at seven. It is to structure them so the comparison set stays legible as it grows.
Grouping by base profile — all fruit, then all dessert, then all menthol — reduces the active comparison load at any moment to a single axis, which pushes the useful ceiling higher. Interleaving a palate cleanser between groups resets aromatic fatigue, which is a sensory phenomenon distinct from cognitive fatigue and which the jam study has no analogue for. And placing the highest-conviction sample ninth or tenth, rather than last, exploits the recency effect instead of being victimised by it.
The forward-looking question for the category is whether the flight is the right format at all. A structured elimination — eight samples, forced ranking into two groups of four, then a head-to-head — extracts more decision-relevant information than fifteen samples tasted in sequence, because it converts a memory task into a comparison task. The tasting room that figures out how to run elimination brackets rather than open flights will sell the sample the customer actually preferred. That is a design problem, not a marketing one, and it is solvable.