Brave’s chatbot placed first in a June 2026 rerun of the company’s own AI answer benchmark, beating Grok, Google AI Mode, ChatGPT, and Perplexity on a scale the company calls Bradley-Terry Elo. The company that designed the query set, picked the judging models, and built the scoring method is also the company selling the product the benchmark promotes: an API priced at $5 per 1,000 requests. The numbers deserve more scrutiny than Brave’s LinkedIn post gave them, because the second scoring column on the results page does not mean what most readers will assume it means.

The problem sits in that second column. Brave’s table header describes it as a five-point Likert absolute category rating, the standard survey-style measure of how good a single answer is judged to be. The bar chart directly beneath the table, plotting the identical six values, labels the same column a Copeland score, defined on its own axis as opponents beaten out of five. Those are not interchangeable statistics, and the chart’s own numbers settle which label is correct: integers spaced perfectly at 5.0, 4.0, 3.0, 2.0, 1.0, and 0.0 across six entrants are a Copeland signature, not an averaged rating. Perplexity’s 0.0 records that it lost every one of its five head-to-head matchups. It does not mean judges rated its answers worthless, and any reader who repeats “Perplexity scored zero” as a quality verdict is repeating a mislabeled chart.

Three more choices in the presentation point the same direction. The Elo chart is captioned with a 95 percent confidence interval, but no error bars appear on the bars and no interval values are published anywhere in the post, so that confidence figure cannot be checked against anything on the page. The chart’s horizontal axis starts at 650 rather than zero, a common convention for Elo scores that also stretches the 12-point gap between Ask Brave’s 1169 and Grok’s 1157 across more of the visible chart than the raw numbers justify. Labeling is inconsistent even before the Likert-Copeland mixup: the winning entrant appears in both bar charts as Ask Brave v2, yet the table sitting directly above calls that same entrant plain Ask Brave, and sets it alongside a distinct Ask Brave v1. Nowhere does the post say what divides the two versions, or when the older one shipped.

Some of the design holds up. Every pairwise comparison in both rounds ran twice, once in each order, specifically to guard against position bias in the two judging models, Claude Opus 4.5 and Claude Sonnet 4.5. That is a real control, not a cosmetic one. Brave also did not always win this exercise. In the first round, run November 30, 2025 on a genuine five-point scale, Grok placed first at 4.71 and Ask Brave followed at 4.66, a spread of 0.70 points across all five entrants, with Google AI Mode, Google’s conversational search experience, at 4.39, ChatGPT at 4.32, and Perplexity at 4.01. Brave’s own commentary on that round conceded the result, crediting its grounding data rather than its model for keeping the gap close.

The benchmark exists to sell the LLM Context API, which ranks context fragments rather than full web pages for AI systems to consume. Brave now bills that product under a Search plan at $5 per 1,000 requests and an Answers plan at $4 per every 1,000 web searches, plus $5 per million combined tokens. Every plan includes $5 of free credit each month, but claiming it requires publicly crediting the Brave Search API on the developer’s own site or about page. Brave says the underlying endpoint already handles more than 22 million answers a day inside Brave Search, the scale figure the benchmark is meant to support.

None of this proves Brave’s underlying claim, that better grounding data can beat a larger model, is wrong. It does mean the specific figures circulating on LinkedIn this week were designed, scored, and labeled by the company selling the product those figures promote. Teams evaluating grounding or context APIs should ask Brave directly for the missing confidence intervals and the Ask Brave version definitions before citing any number from this post in a procurement comparison.

PPC Land, written by Luis Rijo, published the original report on July 26, 2026.