Ask any AI assistant to find you the best roofer within twenty miles. You’ll have an answer in about four seconds. Three names, ranked, with reasons attached. It will sound authoritative, because it is designed to sound authoritative.

Now ask yourself whether you’d spend $50,000 on it.

Most people wouldn’t, and the skepticism runs well beyond anything Yelp has measured about itself. Adoption is accelerating far faster than confidence. Pew Research Center reported this summer that 49% of American adults now use AI chatbots, up from 33% in 2024, and that 60% read AI-generated summaries in their search results. Separate Pew research found that among Americans who had encountered those summaries, 53% had at least some trust in them. But only 6% trusted them a lot. Forty-six percent had little or no trust at all.

Yelp’s own research with Morning Consult this February found the same tension from another angle. Among 2,202 American adults, 65% had used an AI-powered search tool in the previous six months, but only 15% trusted what it told them a great deal. Sixty-three percent said they double-check AI results against other sources. Half described AI answers as a kind of walled garden — difficult to verify, impossible to inspect.

AI adoption is widespread. Deep trust isn’t. Which means people are now doing something faintly absurd at enormous scale: getting an instant answer, and then going to do the work by hand anyway.

That gap may turn out to be one of the defining economic realities of the AI era, and almost nobody is pricing it correctly.

AI made answers cheap. Verification became the scarce resource.

Earlier this year, I sat with the board of a Fortune 100 consumer products company working through its AI strategy. What struck me wasn’t the answer they reached. It was that nobody around the table believed competitive advantage would come from the AI itself. Almost no time went to which model to license. The question that consumed the afternoon was harder: once every competitor has access to the same frontier models, what do we own that a model can’t reproduce? The answers that survived scrutiny weren’t technologies. They were trusted customer relationships, proprietary evidence, and governance the company had built to protect its credibility at the expense of short-term efficiency. That’s not primarily a technology problem. It’s a leadership problem. And as I started looking closely at Yelp, I realized we had spent that afternoon wrestling with essentially the same strategic question.

Yelp Was Supposed To Lose

If you’d asked a room of strategists three years ago which companies conversational AI would flatten first, review platforms would have made every list. The logic was airtight. Yelp existed because finding a good plumber required reading forty strangers’ opinions and forming a judgment. If a machine could read the forty opinions and hand you the judgment, what exactly was left?

This was not a fringe view, and the numbers still give it oxygen. Yelp’s second-quarter results released August 6, showed net revenue of $376 million, up just 1% year over year, while advertising revenue fell 3% and adjusted EBITDA fell 9% to $91 million. For the full year, the company now expects adjusted EBITDA of $315 million to $325 million, against $369 million in 2025.

That is a company choosing to absorb significant near-term financial pressure to finance a strategic transition. Yelp acquired the AI lead-management firm Hatch for roughly $264 million in cash, drew $100 million on its credit facility, watched its cash balance fall from $216 million to $94 million over six months, and has now paused its share repurchase program to pay down the revolver, with buybacks expected to resume in 2027.

You can read that as a declining business financing a pivot it didn’t choose. You can also read it as a management team converting its balance sheet into a position the market hasn’t priced. Both readings sit on the same set of facts. Which is precisely why the underlying asset is worth examining on its own terms, separate from any given quarter.

Yelp also arrives at this moment carrying twenty years of scar tissue. It is currently suing Google , alleging the company used its dominance in general search to favor its own local results — the latest chapter in a fight over local discovery that has lasted more than a decade. This is not a company that has floated serenely above the platform wars, and any argument about its strategic position has to survive that history.

In July, OpenAI licensed Yelp’s content — ratings, reviews, photos, business details — into ChatGPT. Those ratings and reviews now power ChatGPT’s local experience in relevant categories. Yelp recently announced the next step: ChatGPT users can book a table or join a waitlist through Yelp without leaving the conversation, with quote requests for home services to follow. That builds on integrations with Apple Maps, Amazon’s Alexa+, Microsoft Bing, DuckDuckGo, Yahoo and a set of automotive brands.

Watch what just changed. The AI ecosystem faced a practical choice: spend years building trusted local experience datasets from scratch, or partner with organizations that already had them. Partnership won, repeatedly, across nearly every major platform in the category. But licensing content is one thing. Routing a reservation through someone else’s infrastructure is another.

The organic citation data points the same direction, with an asterisk worth stating plainly: Yelp commissioned the study. An analysis it funded, measuring citations across ChatGPT, Gemini, Google AI Mode and Perplexity in the fourth quarter of 2025, counted roughly 512,700 citations to Yelp against 149,700 for the next-closest home services platform — a 3.4x lead over the Better Business Bureau, with Angi, Thumbtack, HomeAdvisor and Nextdoor further behind. Company-funded research deserves scrutiny. But the counts are specific, the methodology is disclosed, and the direction matches what the partnerships already show.

The systems that were supposed to make Yelp unnecessary started depending on it. Now they’re beginning to transact through it.

But there’s another way to read exactly the same development, and it’s the reading a board should sit with. If ChatGPT owns the interface while Yelp supplies the evidence and the transaction rails underneath it, Yelp becomes more useful and less visible at once — a trusted input inside someone else’s customer relationship. Licensing and routing can deepen a moat. They can also convert a company into a commodity supplier, paid per query, with no one left who knows its name. The question isn’t whether AI will replace intermediaries. It’s which intermediaries AI decides it still needs, and on what terms. Whether this deepens Yelp’s position or accelerates its commoditization is genuinely unresolved.

Why AI Needs Ground Truth

Here’s what’s easy to miss about large language models: they don’t know things. They predict. Given everything written before it, a model computes the most probable next token, then the next, then the next. This is a remarkable capability and it is genuinely useful — but it is not the same as knowing whether a particular roofer actually showed up on the day he promised.

Machine learning has a term for this. Ground truth refers to the verified reference data a model’s outputs are evaluated against — the answer key. In business, the analogous asset is harder to acquire: authentic, accumulated evidence that people trust enough to act on. And it has quietly become the scarcest input in the AI economy, because everything else can now be generated on demand.

Yelp’s reviews aren’t ground truth in the strict machine-learning sense. Human experience is too subjective for that. But they illustrate the business equivalent — evidence whose provenance and governance make it more reliable than generated assertion.

This isn’t a framing the industry invented for marketing purposes. It’s where the research points. The standard technique for reducing hallucination is retrieval-augmented generation — pairing a model with an outside corpus so its answers are anchored to documents rather than statistical patterns. But retrieval alone doesn’t solve it. What the corpus contains turns out to matter enormously.

The clearest demonstration came in Nature this February. Researchers built OpenScholar, a comparatively small model paired with a curated store of 45 million open-access scientific papers and a pipeline that verifies its own citations, then tested it against GPT-4o. The smaller specialized system outperformed the frontier model on a challenging multi-paper synthesis task. More striking: asked to cite recent literature, GPT-4o fabricated its citations 78 to 90 percent of the time, while the grounded system achieved citation accuracy on par with human experts.

Read that as a business proposition rather than an engineering one. A smaller, specialized system built around better evidence beat a larger general-purpose model without that grounding. If grounding is only as good as the ground underneath it, the leverage in the AI stack shifts to whoever owns a corpus worth grounding in. That is a sentence about competitive advantage, not machine learning.

Akhil Kuduvalli Ramesh, Yelp’s Chief Product Officer, frames the company’s role in the AI stack this way: “We’re not here to replace human judgment. We are here to aid it. And one of the best ways to aid human judgment is to share perspectives.”

That sounds soft until you watch it operate.

Here is a sentence from a five-star Yelp review of Ace Roofing, an Austin contractor. The customer, Lewis G., loved his new standing seam metal roof. Then he added this: when the crew removed his skylights, rather than cleaning them before reinstalling, “someone decided to draw faces and hearts in the grime.”

Sit with what that sentence is. No marketing department would ever publish it. No starrating can contain it — the review is five stars, so the number tells you nothing. And no language model could know it happened, because it exists in the record only because one person climbed onto one roof and looked. That is what evidence is, as distinct from information. It is also, in miniature, the entire asset Yelp has spent twenty years accumulating: a few hundred million small facts that nobody had a commercial reason to record.

Kuduvalli describes the shift as finding the needle in the haystack. A business might carry 400 reviews. Before, you read as many as your patience allowed and hoped the one relevant to your situation surfaced.

He offered an example from his own weekend. Particular about South Indian filter coffee, he’d found a place forty minutes away, but he was bringing a friend with a gluten allergy and a dog. So before driving, he asked. The assistant surfaced photographs of the dishes and cross-referenced the restaurant’s own website on ingredients. On the dog question it found no published policy. Rather than inventing an answer, it surfaced customer photos showing outdoor seating, said it couldn’t confirm whether dogs were allowed, and let him decide.

The system’s most valuable move was declining to answer — and showing him the evidence anyway. Every part of what it returned was traceable to a person: a reviewer who ate the dish, a customer who photographed the patio, a business that published its own menu. The recommendation wasn’t the product. The evidence behind it was. And Yelp’s real asset isn’t simply the evidence. It’s twenty years of judgment about which evidence deserves to count.

Trust Isn’t A Value. It’s An Architecture.

Every company says it values trust. Almost none of them will show you the plumbing.

Yelp will. And what’s in there is stranger than the marketing suggests.

The software that decides which reviews count toward a business’s star rating evaluates every review against hundreds of signals. Here’s the part that matters: no employee at Yelp can override it. Not a salesperson, not an executive, not the CPO. The company states plainly that this is deliberate, designed to eliminate conflicts of interest. Advertisers and non-advertisers are subject to identical treatment.

Read that again as a governance decision rather than a product feature. A public company built the system that generates its revenue, then made itself structurally incapable of putting a thumb on it. That constraint is precisely what helps make the underlying data credible enough to serve as a grounding source for other systems.

Kuduvalli describes the same instinct running through product design. “We could have gone to market faster, but we wanted to design Yelp Assistant in a way where every answer is evidence led, clear and transparent.” They could have simply delivered the answer. Instead they built the harder thing: the answer, plus the receipts.

Every institution claims integrity. Far fewer design systems that remove the temptation to compromise it. Which raises a question I’d put to every leadership team reading this.

What have you made yourself unable to do?

The architecture has a price. Of reviews contributed to Yelp in 2025, 70% were recommended; 17% were not recommended, 11% were removed by the user operations team, and 2% were pulled by reviewers themselves. Roughly three in ten submissions didn’t make the set that counts — an extraordinary amount of self-inflicted subtraction for a company whose product is reviews.

And note the nuance, because it’s the most interesting design choice in the whole system: not recommended doesn’t mean fraudulent. Sometimes the software simply doesn’t yet know enough about a reviewer to vouch for them. Those reviews aren’t deleted. They stay accessible through a link at the bottom of every business page — visible to anyone who wants to look, just not counted.

An evidence system that publishes its own rejects. I can’t think of another one.

There’s an obvious objection here, and it deserves airing. A system that decides which experiences count isn’t neutral simply because the deciding is automated. Yelp’s software makes judgments about credibility, and those judgments determine whose voices carry weight. Raval’s data gives the concern teeth: in his sample, Yelp’s filter hid roughly 46% of reviews of high-quality businesses — essentially the same share it hid for low-quality ones. Either fake reviews are everywhere, he observes, or a great many legitimate reviews are getting caught in the net. Yelp’s answer is that the alternative, letting every submission move the score equally, produces a system that’s trivially easy to game. Fair enough. But notice what that concedes. Yelp hasn’t removed judgment from the process. It has institutionalized it. Every ground-truth system has a governor somewhere, and the honest question is never whether one exists but whether you can see it.

The comparison with competitors is more interesting than a morality tale. Google and Tripadvisor both invest heavily in fraud detection; nobody in this industry is indifferent to fake reviews. Tripadvisor reported rejecting or removing 2.7 million fraudulent reviews in 2024, up from two million the year before, using an anti-fraud system it models on the banking sector.

The difference is architectural, and it sits upstream. Google permits businesses to solicitreviews within its policies. Tripadvisor concentrates its effort on catching fraudulent submissions after they arrive. Yelp goes further up the pipe — prohibiting solicitation outright and withholding even genuine reviews from the rating until it has enough confidence in them. The distinction isn’t who fights fraud harder. It’s where each company draws the line around acceptable influence, and how much revenue it’s willing to leave on the far side of that line.

The enforcement operation behind it runs on similar logic. According to its 2025 Trust & Safety Report, Yelp filtered nearly half a million reviews bearing the characteristics of AI generation, closed 1,991 accounts tied to review-exchange rings, and filed more than 1,020 reports to outside platforms flagging paid-review operations. Sixty percent resulted in content being removed. When review-extortion schemes hit businesses on Google in several cities that year, Yelp says its systems blocked similar suspicious activity before it reached nearly all business pages.

This is the unglamorous work of trust infrastructure: spending money to defend the integrity of an asset even when the abuse occurs beyond the platform you own.

Regulation is raising the stakes here too. The Federal Trade Commission’s Consumer Reviews and Testimonials Rule , in force since October 2024, prohibits a range of fake-review, conditioned-incentive and review- suppression practices, and permits civil penalties for knowing violations. After a quiet first year, the Commission began putting companies publicly on notice, sending warning letters to 10 businesses in December 2025 and reminding them that violations can carry civil penalties of up to $53,088 each. Nothing in that rule requires Yelp’s particular architecture. But parts of what once looked like voluntary institutional discipline are quietly becoming compliance infrastructure.

Why The Architecture Can’t Be Copied

The scale is real: more than 330 million cumulative reviews, 8.4 million claimed business locations, roughly half a billion photos, 74 million monthly visitors.

But scale isn’t the moat. Adjudication is.

The clearest evidence comes from outside the company. Independent research by economist Devesh Raval, who built quality tiers for local businesses using consumer-protection complaint data, found that a low-quality business on Google carries roughly the same average rating as a medium-quality business on Yelp. Only 2% of Yelp reviews run 100 characters or less, against 50% on Google — where 32% of “reviews” contained no text at all. Raval conducted the analysis independently and states the views are his own, though Yelp supplied the underlying data on its reviews.

And Yelp’s ratings distribution looks nothing like a vanity metric. Just 16% of businesses average five stars. Thirty-eight percent land at four, 26% at three.

Yelp is a harder grader. That’s not a marketing claim; it’s a measurable property of the dataset, independently documented.

Which is the whole point. Anyone can generate 330 million pieces of review-like text this afternoon. No one can generate twenty years of deciding which human experiences deserve to count.

The Question Every CEO Is Skipping

This isn’t really a story about Yelp, and it certainly isn’t a story about review platforms. Hospitals hold clinical outcomes, banks hold transaction histories, manufacturers hold operational telemetry, universities hold learning results — proprietary institutional evidence that no model can synthesize and no competitor can fabricate, made valuable not merely by possessing it but by knowing where it came from, how it was governed, and why anyone should trust it.

Most of those institutions are currently asking how fast they can deploy AI. Very few are asking the harder question: what evidence makes our AI believable?

Here’s the reframe I’d push. Trust infrastructure belongs on the same shelf as intellectual property and brand equity — an asset that took years to accumulate, can be destroyed far faster than it was built, and now determines whether your AI outputs are worth anything to the people receiving them. We have decades of practice valuing patents and brands. We have almost none valuing the evidence base underneath them. That gap is going to get expensive.

Models will improve. Compute will get cheaper. Answers will become ubiquitous and, eventually, free. Verification will only become more valuable — because verification is what converts an answer into a decision someone is actually willing to act on. And infrastructure is what converts that decision into action. Increasingly, Yelp owns pieces of both.

One last thing worth saying plainly, because it’s where the real lesson lives.

Yelp wasn’t building for this. Nobody in 2009 was constructing a trust architecture in anticipation of large language models. The company enforced those standards for its own reasons, at genuine cost, for two decades — refusing solicited reviews, un-recommending its own inventory, walling revenue off from ranking — and absorbed a great deal of criticism for it along the way.

Then the technology arrived that made those standards scarce.

Durable advantage rarely looks like genius while it’s being built. It usually looks like an organization stubbornly investing in something the market undervalues — until the market changes, and everyone else calls it a moat.

Which leaves the rest of us with the uncomfortable part. You cannot begin accumulating ground truth at the moment you discover you need it. By then, the only companies that have any are the ones who were willing to look unreasonable for twenty years.