Viral BridgeBench Post Claims Claude Opus 4.6 Was ‘Nerfed,’ Critics Call It Bad Science

BridgeMind AI claimed Anthropic’s Claude Opus 4.6 was secretly degraded after a hallucination benchmark retest. The viral post has since drawn sharp criticism for flawed methodology.

The claim triggered widespread debate over whether AI companies are quietly downgrading paid models to reduce costs.

BridgeMind Claims a 98% Surge in Hallucinations

BridgeMind, the team behind the BridgeBench coding benchmark, posted that Claude Opus 4.6 had fallen from second to tenth place on its hallucination leaderboard. Accuracy reportedly dropped from 83.3% to 68.3%.

The post framed this as proof of “reduced reasoning levels.” However, a closer look at the underlying data tells a different story.

Critics Say the Comparison Is Fundamentally Flawed

According to computer scientist Paul Calcraft, the claim is “incredibly bad science,” highlighting a critical problem with the methodology.

The original high score came from just six benchmark tasks. The new retest expanded the benchmark to 30 tasks.

On the six overlapping tasks, performance was nearly identical, dropping only from 87.6% to 85.4%.

That small swing came mostly from a single extra fabrication in one task. With no repeated runs, this falls well within normal statistical variance for AI models.

Large language models are not deterministic, and one bad output on a small sample can shift results significantly.

Broader Frustrations Fuel the Narrative

Still, the post struck a nerve. Since its February 2026 launch, Claude Opus 4.6 has faced persistent complaints about perceived quality decline.

Developers report shorter responses, weaker instruction-following, and reduced reasoning depth during peak hours.

Some of this traces to deliberate product changes. Anthropic introduced adaptive thinking controls that let the model self-adjust its reasoning budget. The default effort level was later set to medium, prioritizing efficiency over maximum depth.

An independent analysis of over 6,800 Claude Code sessions found reasoning depth dropped roughly 67% by late February.

The model’s file-read ratio before editing code fell from 6.6 to 2.0. That suggests it attempted fixes on code it had barely reviewed.

What This Means for AI Users

This reflects a growing tension in the AI industry. Companies optimize models for cost and scale after launch, while heavy users expect consistent peak performance. The gap between those priorities erodes trust.

Based on the available evidence, the BridgeBench data does not prove a deliberate downgrade. The benchmark comparison was apples-to-oranges, and the overlapping results were nearly identical.

However, the underlying frustration is not entirely baseless. Adaptive compute controls and service-level optimizations have changed how Claude Opus 4.6 behaves in practice. For developers relying on consistent output, those changes matter.

Anthropic has not issued a public statement on the specific BridgeBench claims as of April 13.

The post Viral BridgeBench Post Claims Claude Opus 4.6 Was ‘Nerfed,’ Critics Call It Bad Science appeared first on BeInCrypto.

Source: https://beincrypto.com/claude-opus-nerfed-bridgebench-claim-backlash/

Viral BridgeBench Post Claims Claude Opus 4.6 Was ‘Nerfed,’ Critics Call It Bad Science

BridgeMind Claims a 98% Surge in Hallucinations

Critics Say the Comparison Is Fundamentally Flawed

Broader Frustrations Fuel the Narrative

What This Means for AI Users

You May Also Like

US Blockades Iranian Ports in Strait of Hormuz: Oil Prices Spike Higher – Bitcoin News

Uphold’s Massive 1.59 Billion XRP Holdings Shocks Community, CEO Reveals The Real Owners

Trump 'crashing out' over Orban loss: 'He can't handle the truth'

Trending News

Ethereum Price News: ETH Holds Above $2,000 as Network Activity Stays Near All Time High

Gold continues to hit new highs. How to invest in gold in the crypto market?

WLFI Threatens Lawsuit Against Justin Sun as Token Blacklist Dispute Turns Public

Covéa Chooses Shift Technology as Strategic Partner for Fraud and Risk Management

USD: Blockade supports cautious rebound – Scotiabank

24/7 Live News

Quick Reads

AAVE Flashes Reversal Signal: Is the $100 Reclaim Finally Near?

BEEG Token Analysis 2026: Opportunity or Hype After a 98% Crash?

Why North Korea Openly Steals Crypto: Inside the World's Most Brazen State-Sponsored Heist Operation

BEEG Price in 2026: Short-Term Flip or Long-Term Hold?

BEEG Arbitrage 2026: The Meme Coin Price Gap Most Traders Are Missing

Crypto Prices