The Emperor’s New Score
Why Napa’s Continued Reliance on Numerical Prestige is Failing the Dinner Table
We have looked at Napa Valley’s repetitive self-descriptions, the growing mismatch between tasting supply and visitor demand and the extreme squeeze in the luxury wine category. Now we come to the heart of the matter: the wine itself.
This is the first of a three-part series on how Napa Cabernet lost its way in the score-driven era, how luxury pricing compounded the problem, and what it will take to bring these wines back to the table.
Napa Valley has spent forty years training Cabernet to win a beauty pageant that rarely resembles dinner. In the professional “sniff-and-spit” arena, more color, more extraction, more oak, and more impact in the first thirty seconds often look like greatness. At the table, however, that same wine can feel like a monologue delivered at full volume over the food, over the conversation, and over the moment itself.
That is the trap hidden inside the 100-point era.
A system that once helped demystify fine wine for American consumers slowly hardened into a numerical monarchy, reshaping vineyard decisions, cellar choices, release strategies, and price. Now, in the middle of a demand correction, Napa is discovering that it optimized too many Cabernets for the tasting flight rather than the second glass.
Many will say this is old news.
It is not.
The overwhelming majority of Napa wineries still lean on scores in their marketing, their release language, their websites, and their allocation emails, according to our “Sea of Sameness” analysis.
The valley still clings to this time-worn ritual, as if the emperor’s score can keep ruling long after the crowd has stopped believing in it.
The Genesis of the 100-Point Rule
The history of modern wine evaluation is a tale of unintended consequences. Before the late 1970s, wine quality was a murky territory of flowery prose, merchant-led reputations, and historical classifications like the 1855 Bordeaux list, which was based on market pricing rather than sensory assessment.
In 1978, Robert Parker, a Baltimore lawyer turned critic, launched The Wine Advocate and introduced a 100-point rubric modeled after the American school grading system. It was a stroke of marketing genius. By starting every wine at a base of 50 points and awarding “grades” in the 80s and 90s, Parker provided a familiar frame of reference for consumers who found traditional European wine terminology elitist.
Parker’s rise to global authority was cemented by his assessment of the 1982 Bordeaux vintage; while established critics viewed the year as overly ripe, Parker recognized its profound concentration and aging potential. When the vintage eventually proved legendary, his 100-point system became the law of the land, allowing a single man’s judgment to move millions of dollars’ worth of wine.
This created the phenomenon of “Parkerization,” where wineries around the globe began tweaking their winemaking to match a seemingly specific palate: bold, ripe fruit, high alcohol, and aggressive oak treatment.
The dilemma we face today is that this evaluative framework is fundamentally disconnected from the “occasion of use.” A professional wine evaluation is an exercise in sensory isolation. Judges often taste flights of 60 to 100 wines in a single sitting, utilizing the “swirl, sniff, sip, and spit” protocol to avoid the intoxicating effects of alcohol. In this high-volume environment, the wines that stand out are the “gladiators”—outliers designed to cut through palate fatigue in a single, half-ounce sip.
The Ashenfelter Challenge: Moneyball for Wine
While Parker was consolidating his cultural power, another school of thought emerged from economics — and it landed in the wine world like a thrown brick. Princeton economist Orley Ashenfelter began showing that a wine’s long-term quality and market value could often be predicted with surprising accuracy from objective weather data: winter rainfall, temperatures during the growing season, and rainfall near harvest. In other words, some of the mystery surrounding fine wine could be stripped away by variables a farmer, statistician, or insurance actuary could understand.
This was heresy. The fine-wine world preferred to imagine itself as a domain of cultivated intuition, historical memory, and refined sensory gifts. Ashenfelter suggested that, at least in some cases, the emperor’s robes looked suspiciously like a spreadsheet.
His work was a kind of Moneyball for wine: a reminder that reputation and ritual are not the same as truth.
And it immediately created controversy, especially because it cut so close to Parker’s authority. Parker had built a global franchise on the force of palate, prose, and conviction, reinforced by his famous endorsement of the 1982 Bordeaux vintage. But that same vintage would later become part of Ashenfelter’s case.
Using only weather data—rainfall, temperature, and growing conditions—Ashenfelter showed that the quality and eventual market performance of vintages like 1982 could be predicted without ever tasting the wine. In effect, the model explained outcomes that had been attributed to critical insight.
That was the uncomfortable implication. It suggested that what looked like foresight might, in some cases, have been retrospectively validated by conditions that were measurable all along. The question was not whether Parker had great taste. It was whether the system had mistaken authority for explanation.
The Score as Style Governor
Once scores became currency, they began shaping taste itself. If a 96 or 98 could change allocation demand, mailing-list behavior, restaurant placements, and resale value, wineries had every incentive to craft wines that performed dramatically in the critic’s glass. That often meant deeper color, riper fruit, more new oak, more glycerol, more concentration, and more sheer force.
Over time, that stylistic drift often translated into very specific choices in the vineyard and cellar. Fruit was picked later and later in pursuit of concentration and plushness, bringing higher sugars and, with them, higher alcohol. New oak became attractive not just because it added complexity, but because it dominated the glass in a way critics could notice quickly. In some cases, even American oak — with its sweeter, more conspicuous aromatic profile — could become part of the effort to make a wine announce itself faster and louder.
The point is not that every one of these choices is inherently wrong. The point is that they were often rewarded for how they performed in a brief, isolated tasting rather than how they behaved over the course of a meal. Later picking did not just intensify flavor; it often pushed alcohol upward and altered the wine’s balance, making it softer and more fruit-forward in youth.
Wines built this way could feel opulent and flattering early on, but sometimes less suited to the long arc of development Napa producers were simultaneously claiming.
Sadly, many high-scoring wines have proved disappointing when opened after only ten years: overpromised and underdelivered, all in the name of getting a score early in the wine’s life. The score era did not just institutionalize a particular palate. It encouraged a chemistry set of decisions designed to win the tasting.
The result was not universal uniformity, but it did produce a recognizable archetype: the big, polished, heavily worked Napa Cabernet designed to register instantly in a professional tasting environment.
Dozens of Napa wineries were guided toward this outcome by a surprisingly small circle of influential consultants. The top five—Thomas Rivers Brown, Philippe Melka, Julien Fayard, Jean Hoefliger, and Aaron Pott—are associated with roughly 162 wineries.
Even after accounting for substantial overlap, just 15 consultants are responsible for the winemaking or blending at approximately 220 distinct Napa Valley wineries and brands. That amounts to roughly 40 percent of the active membership of Napa Valley Vintners.
Approximately 202 of these 220 entities (92%) have received a score of 94 or higher from a major critic (Wine Advocate, Vinous (Galloni), James Suckling, Jeb Dunnuck) in the last five years. Together, they account for nearly 90 percent of all 100-point scores awarded to Napa wines during that period.
It should not be surprising that wines from many different wineries came to share similar traits. They were being trained for the same exam.
That is where the score culture began quietly drifting away from wine’s primary purpose. Wine is not perfume. It is not a competitive diving score. It is not an IQ test for collectors.
It is an agricultural beverage meant to accompany food, company, and time.
When the Method Shapes the Product
Structured tasting has its place. Critics need methods. Buyers need comparability. But once the format begins dictating the product, the format becomes a distortion.
The “sniff-and-spit” world rewards instant impact. It does not reward restraint, grace, or the slow reveal over ninety minutes with a ribeye, lamb chops, or even a roast chicken. It does not reward the wine that grows more interesting with food and conversation. It rewards the wine that can dominate a crowded field in a very short audition.
Many have attributed award-winning scoring results to Robert Parker’s palate. It may be equally plausible that the results were the result of the tasting methodology. As experimental scientists will tell you, results are never independent of the measurement method. A flight of wine in a sterile sniff-and-spit format may have a predictable result.
That is why so many of the most celebrated modern Cabernets can feel curiously detached from the dinner table. They are masterpieces of auditioning. They are less reliable partners for drinking. They were not meant to be part of an ensemble.
The Illusion of Precision
Even if one grants that scores are a useful shorthand, scientific evaluation has long raised uncomfortable questions about the reliability and replicability of wine scores. Robert Hodgson’s papers in the Journal of Wine Economics, published by the American Association of Wine Economists, were especially damaging to the mythology of precision.
In one major U.S. wine competition study, Hodgson found that only about 10 percent of judges were able to replicate their score within a single medal group, while another 10 percent sometimes scored the same wine anywhere from Bronze to Gold.
Across 13 U.S. competitions, the results were even more unsettling: of the 2,440 wines entered in more than three competitions, 47 percent won Gold somewhere, but 84 percent of those same wines also received no award elsewhere. Other research has found that both within-judge reliability and between-judge consensus in wine judging are substantially lower than in many other expert domains, and that consistency varies widely across tasters. In short, the authority of the score often exceeds its reproducibility.
Even Robert Parker eventually conceded the point from the other direction. Asked how often he re-tasted a 100-point wine and repeated the same score, he answered: “Probably about 50% of the time.” If even the most influential critic of the 100-point era could not reliably reproduce perfection, the score begins to look less like a measurement than a moment.
This is one reason the whole culture of scoring deserves more skepticism than it usually receives. The scores are treated as if they were measurements, but they often behave more like opinions wearing lab coats.
The Cost of Numerical Monarchy
Scores also changed the psychology of the buyer. Once numerical prestige became central, consumers were trained to outsource judgment. A 97 from Parker, Galloni, Dunnuck, or Suckling became more than information; it became permission. It reassured the buyer that the bottle was worthy, collectible, and socially defensible.
But the arithmetic of abundance has made that system harder to navigate. How does a buyer choose among just the roughly 202 Napa wines that have scored 94 or higher from the small circle of highly influential winemaking consultants—let alone the dozens of other Napa wines that also scored highly? The signal that once simplified choice now overwhelms it.
A broader look at the data reinforces the same point. Based on an analysis of critical reviews from the past five years (2021–2026), approximately 318 unique Napa Valley wineries have received a score of 94 or higher for Cabernet Sauvignon from the four major critics—Wine Advocate, Vinous, James Suckling, and Jeb Dunnuck. Over that same period, these wineries collectively received more than 1,250 such citations.
What ties these wines together is not just their scores. It is the system they are built to satisfy. However different their vineyards, ownership, or backstories, they are being evaluated—and therefore shaped—within the same framework: a brief, comparative, score-driven tasting that rewards immediate impact, density, and presence in the glass.
The result is not simply abundance. It is convergence. The buyer is not choosing among 300 fundamentally different expressions of Napa Cabernet. He is choosing among hundreds of wines optimized, to varying degrees, for the same moment of evaluation—one that bears only partial resemblance to how wine is actually consumed at the table. In that process, distinctions of site, soil, and microclimate are often softened—sometimes lost altogether—in favor of a style that performs reliably in the critic’s glass.
Systems built on authority eventually weaken when authority fragments. Today the critic landscape is more diffuse, the consumer is more skeptical, and younger drinkers are less inclined to organize their taste around a numerical hierarchy.
The old monarchy still exists, but it no longer commands automatic obedience.
That leaves Napa in an awkward position. It has a large installed base of wines, brands, and price structures created under the old regime, even as the cultural machinery that supported that regime is losing force.
Over time, the score also began serving another constituency: the owner. For a certain class of Napa proprietor—especially those who came to wine after building fortunes and reputations in other arenas, like finance, technology, real estate, or private equity—the high score became a kind of trophy. It validated not just the wine, but the owner’s judgment, ambition, and successful arrival in yet another elite arena.
A 97 or 100 was no longer merely a signal to the consumer. It became a signal to peers. It no longer had to justify itself primarily at the table. It had to justify itself in the status economy of ownership.
The bottle became less an agricultural beverage meant to accompany dinner than an object built to win applause, earn rankings, and reinforce identity.
In that sense, the score did not merely commercialize Napa wines. It became a currency of status. But like any currency, its value depends on scarcity—and that scarcity has diminished.
The First Green Shoots of a Better Method
To be fair, the world of wine evaluation is not standing still. Newer frameworks are beginning to emerge that ask a more useful question than “How impressive is this wine in a lineup?” They ask how the wine behaves with food, in context, and over the life of an actual meal.
Some competitions have begun moving in this direction explicitly. The Sommeliers Choice Awards were designed around restaurant utility and “pairability” rather than abstract critic theater. The Sydney International Wine Competition, founded by Len Evans, has long judged wines not only on their own, but also with food. In Hong Kong, leading competitions have gone even further, awarding specific distinctions for how wines pair with different cuisines.
Even the Grand Central Oyster Bar in New York sponsors an annual competition built around a beautifully old-fashioned question: which wines actually work best with oysters? That is a more grounded test of usefulness than asking which Sauvignon Blanc, Chardonnay or sparkling wine shouts loudest in a sterile tasting flight.
These experiments matter because they begin to restore the missing variable: occasion of use. They recognize that wine is not just an object for pronouncement, but a companion to food, company, place, and time.
What does 94 points mean when you are having caviar? Is it different with fish? With a rib-eye?
Not everyone accepts this shift. Traditional critics have long argued that introducing food into evaluation compromises comparability—that it turns the judgment from the wine itself to the pairing, and introduces too many variables to produce a clean, transferable score. There is truth in that concern. But that objection also reveals the deeper divide. The traditional model optimizes for control and comparability—at least in theory. The gastronomic model optimizes for reality. And wine, ultimately, is consumed in the real world.
The evolving frameworks also point toward a standard that many consumers understand instinctively but the score era often ignored: a great wine should not merely impress in the first sip. It should make sense at the table.
What the Dinner Table Knows
The dinner table is a more honest critic than the tasting sheet. It reveals whether a wine invites another glass, flatters the food, and can live inside a real evening rather than a professional performance. Many Napa Cabernets still do this beautifully. At their best, they do more than accompany a meal. They elevate it. But too many have been trained for spectacle rather than companionship.
A high score is not a mistake. A 95-point wine is almost always well-made, expressive, and excellent. The problem is not the presence of the score. It is the weight we have asked it to carry. A number can capture technical quality and momentary impact. It cannot, on its own, tell you how a wine will live at the table.
This is not just a matter of opinion. A growing body of commentary from within the wine world has drawn the same distinction between what might be called clinical evaluation and gastronomic reality. Benjamin Lewin, Master of Wine, has noted that even experienced tasters can be misled by wines tasted in isolation, observing that “the sheer deliciousness of a wine taken in isolation will give a misleading impression of how it will taste with food.” A wine that dazzles in a sip can prove overbearing across a meal.
Eric Asimov, the longtime wine critic of The New York Times, has made a similar point more bluntly. He has argued that many modern, high-scoring Cabernets—“soft, fruity and oaky”—lack the structure, freshness, and restraint that allow a wine to sustain interest at the table. They may impress at first, but often fail to carry a meal.
Jon Bonné has described the same phenomenon from another angle: a generation of wines shaped by what critics reward rather than by how they are meant to be consumed. In that world, ripeness and impact can crowd out the acidity and balance that make a wine useful with food.
None of this is an argument against ambition, craftsmanship, or greatness. It is an argument against confusing extraction with depth, immediate impact with longevity of pleasure, and a professional tasting score with a wine’s real usefulness in life.
Napa does not need less excellence. It needs a broader definition of excellence—one that includes drinkability, context, and joy.
The 100-point system helped build the valley’s modern prestige. But prestige built on a ritual that no longer matches how people actually drink is a fragile thing. That is the problem confronting Napa now: we have spent decades creating wines for the judges’ table while the real market increasingly lives somewhere else.
The emperor may still have a score. The harder question is whether he still has a dinner invitation.
Upcoming
Part II: Life Beyond the Score
Bringing Napa Cabernet Back to the Table
* * *
Ted Hall is a former Senior Partner at McKinsey & Company and a founder of the McKinsey Global Institute. He writes about economics, incentives, and how complex systems shape real-world outcomes, drawing on decades of work across agriculture, food, wine, and consumer markets. A winemaker for more than 50 years, he is co-founder of Long Meadow Ranch and a former chairman of Robert Mondavi Corp.









excellent post. I would simply add that the Parkerization of the wine world had resulted in better wines overall. would that have happened anyway over time? well, nothing like an economic incentive to goose investment in equipment, vineyard, and expertise!
Ted
Solid review of one of the oldest problems in the American wine culture. As I have been involved in the wine trade for over a half century, like you, as a wine judge, writer, retailer etc ad nauseum, I have never subscribed to 100 pt system, for all of the reasons you note; and one more, effectively. It actually obscures the bigger picture that wines compared microscopically by grading out of 100 like 91 vs 93 pts fails to note that both are excellent but different. And that is more important to enjoyment of the wine. One note of correction (IMO): You wrote: It is an argument against confusing extract with depth, immediate impact with longevity..." The issue is not Extract, but Extraction. "Dry Extract" is a measurable thing and has real consequences for identifying a finer wine than one less so, whether it is a 9% alcohol German Riesling or a 14% Barolo. Trying to extract too much from the grapes via higher sugar levels and alcohol, over-long macerations etc etc is generally the bigger problem as it invariably disrupts the balance and harmony the fruit originally provided. But paying attention to those aspects of farming and winemaking that lead to mature fruit and phenolics commensurate with the climate of a place and the appropriate variety generally leads to good Dry Extract and fine quality plus balance. That is something that too many Napa Cab producers have lost site of in their 'quest for points'. So much more one can say, but you definitely make the case for taking points out of the equation. Personally, for years and now with WIneKnowLog Reports and other writing, I continue to use a simple three star system: Good/very good; Very good/excellent; Outstanding. Simple understandable and judgmental only insofar as the notes for each wine's description. But I will say, as one of your commenters did, that the Australian show system and the old 20pt system which sets minimum benchmark scores such as 16 pt for a Bronze medal (good quality, some personality, no faults etc) provided the basis for my 3 star system. Which seems eminently sensible in the wider scheme of things! Joel