Skip to content
IdeaScout.
← Back to archive

LLM quality degradation tracker ("nerf tracker")

AI-discovered

Problem

AI model providers quietly degrade models after launch — quantizing them more, or routing requests to older models with a changed system prompt — while users have no way to prove it. A developer reports tasks that were previously one-shotted now getting stuck in reasoning loops, and suspects silent model swaps are common across the industry.

Opportunity

A continuously-running benchmark service that executes identical tasks daily against each model/API endpoint and publishes longitudinal quality regressions — an independent 'nerf detector' for LLMs that devs and companies would pay to trust.

Market analysis

Real, recurring demand: every 'they nerfed the model' cycle on HN shows developers want longitudinal proof of quality regressions, and providers structurally cannot self-report it credibly. The catch is that independent measurement is already anchored by Artificial Analysis, and the closest match to this exact idea (livenerf) is a free open-source repo.

Market · AI engineers, agent builders and CTOs depending on frontier model endpoints; demand resurfaces every time a model quietly changes behavior.

Pricing · No clear price anchor surfaced; independent leaderboards are free to read, so the sellable layer would be alerts, historical data or API access.

score 6/10 by glm-5.1

Pros

  • + Strong, recurring demand signal every time a model 'feels worse'.
  • + Cheap to run at small scale: frozen task panel plus a daily cron.
  • + Trust niche that model providers cannot credibly occupy themselves.

Cons

  • − Artificial Analysis already owns independent LLM measurement mindshare.
  • − Methodologically hard to separate real nerfs from infra bugs and A/B rollouts.
  • − Closest analog (livenerf) is free and open source, capping willingness to pay.

Source

Hacker News (comments)

Open original thread ↗

The hard part is not running tasks daily — it is being right in public. livenerf’s own methodology notes that the 2025 “quality incidents” turned out to be infrastructure bugs rather than deliberate downgrades, and that a model as served through a subscription client is a different object from the raw API endpoint. That distinction points at the credible solo wedge: longitudinal tracking of coding agents as actually served on consumer plans, which is where most nerf reports originate and where cross-model leaderboards do not go. Trust is the entire product here — a single wrong “nerf confirmed” headline burns it permanently.