Opinion

Why We Don't Use Star Ratings — and Why You Shouldn't Trust Them

A 4.3 out of 5 tells you almost nothing worth knowing. Here's what the average is hiding, and what we do instead.

  • star ratings
  • reviews
  • consumer advice
  • buying guides
  • trust

A star rating is the most confident-looking number in shopping, and one of the least honest. It arrives pre-digested, universally understood, comparable across every category — and it collapses everything you’d actually want to know into a single decimal that mostly measures the mood of a self-selected crowd. We don’t publish star ratings on RBE. Not because scoring is hard, but because the number lies in ways that are predictable, systematic, and almost impossible to see once it’s been averaged.

Here’s the core problem: an average assumes the thing it’s measuring has a middle. Most products don’t.

The average describes a person who doesn’t exist

When a product lands at 4.3 stars, the instinct is to picture a typical buyer who was a bit better than satisfied. But pull the distribution apart and you rarely find a gentle hump around four. You find two crowds. A large group who gave it five stars because it worked fine, and a smaller but angry group who gave it one star because it broke, arrived damaged, or died in six months. Average those together and you get a 4.3 that describes nobody — not the happy majority, not the burned minority.

That shape has a name: bimodal. And it matters enormously, because the one-star cluster is usually where the truth lives. “It stopped charging after four months” is a far more useful sentence than “4.3 stars.” The average doesn’t just hide that failure mode — it actively dilutes it, folding a serious defect into a reassuring number. Two products can share the same score while one is reliably decent and the other is a coin flip between delight and disaster. The star rating cannot tell those two apart. A verdict can.

Every thumb is on the scale

Even if averages described reality, the inputs are compromised before they’re counted. This isn’t conspiracy; it’s just how the incentives run.

Reviews are gamed. Sellers seed listings with reviews from free or discounted units, bundle in “leave us five stars” inserts, and quietly redirect unhappy buyers to customer service instead of the review box. Independent investigations into review manipulation keep finding the same playbook, and platforms keep purging batches of fake reviews — which tells you both that it’s happening and that a lot of it never gets caught.

Reviews are incentivized. A buyer who got a product free, or heavily discounted in exchange for feedback, is not a neutral witness. They rate higher, and they rate sooner, which skews the number upward exactly when a product is new and has the fewest honest data points to balance it.

Reviews are front-loaded in time. People rate on the high of unboxing, not after the hinge cracks or the battery sags a year in. The failures that matter most for a durable purchase — longevity, support, how it ages — are the ones least likely to make it into the average, because by the time they show up, nobody’s still writing reviews.

And ratings inflate by category. Star scales are relative to their neighborhood. In some categories almost everything clusters above four stars, so a four-star product is actually below average; in others a four is genuinely strong. A single decimal ripped out of that context is worse than no information, because it feels like it means something universal when it doesn’t.

What we do instead

Our answer isn’t a smarter score. A weighted, adjusted, better-normalized number would still be a number — still pretending a bimodal, gamed, time-skewed, category-relative mess can be honestly compressed into one figure. The fix isn’t a better average. It’s refusing to average at all.

So we give a plain verdict: Buy, Skip, or It depends — and then we show our work. Who is this genuinely right for, and who should walk away. The specific failure modes owners keep reporting, stated as failures and not smoothed into a mean. The tradeoffs a category forces on you, so a solid budget pick isn’t punished for not being premium, and a premium one isn’t excused for costing a fortune. We synthesize independent testing and real owner reports rather than running our own lab, and we tell you plainly when the honest answer is “it depends” instead of manufacturing false confidence. If you want the full method, it’s laid out in how we review.

The point of a verdict is that you can argue with it. A 4.3 offers nothing to push against — you either trust the number or you don’t. A verdict says “buy this unless you need X, in which case skip it, because owners keep hitting Y.” You can read that reasoning, weigh it against your own situation, and reach a different conclusion. That’s not a weakness of the format. That’s the whole job.

So what should you do

When you see a star rating, treat it as a smoke alarm, not a verdict. A product buried at two stars with thousands of ratings is worth avoiding. But in the crowded four-to-four-and-a-half range where most shopping actually happens, the number has stopped discriminating — so ignore the decimal and go read the one- and two-star reviews first. That’s where the real product is described. Look for repeated, specific complaints; a pattern of the same failure is a signal, while a scatter of one-off gripes is just noise.

Or skip the excavation and read a verdict that already did it. That’s the entire reason RBE exists: we tell you what to skip, not just what’s popular. Start with our verdicts, and next time a glowing average tempts you, remember that it was built to reassure you — not to inform you.

Real Buyer Experiences is reader-supported and independent. We synthesize independent testing and owner reports into one honest verdict — we never trade a recommendation for payment. How we review →

More from The Journal