LLM quality degradation tracker ("nerf tracker")
Problem
AI model providers quietly degrade models after launch — quantizing them more, or routing requests to older models with a changed system prompt — while users have no way to prove it. A developer reports tasks that were previously one-shotted now getting stuck in reasoning loops, and suspects silent model swaps are common across the industry.
Opportunity
A continuously-running benchmark service that executes identical tasks daily against each model/API endpoint and publishes longitudinal quality regressions — an independent 'nerf detector' for LLMs that devs and companies would pay to trust.
Market analysis
Real, recurring demand: every 'they nerfed the model' cycle on HN shows developers want longitudinal proof of quality regressions, and providers structurally cannot self-report it credibly. The catch is that independent measurement is already anchored by Artificial Analysis, and the closest match to this exact idea (livenerf) is a free open-source repo.
Market · AI engineers, agent builders and CTOs depending on frontier model endpoints; demand resurfaces every time a model quietly changes behavior.
Pricing · No clear price anchor surfaced; independent leaderboards are free to read, so the sellable layer would be alerts, historical data or API access.
Pros
- + Strong, recurring demand signal every time a model 'feels worse'.
- + Cheap to run at small scale: frozen task panel plus a daily cron.
- + Trust niche that model providers cannot credibly occupy themselves.
Cons
- − Artificial Analysis already owns independent LLM measurement mindshare.
- − Methodologically hard to separate real nerfs from infra bugs and A/B rollouts.
- − Closest analog (livenerf) is free and open source, capping willingness to pay.
Existing / similar tools
Source
Hacker News (comments)
The hard part is not running tasks daily — it is being right in public. livenerf’s own methodology notes that the 2025 “quality incidents” turned out to be infrastructure bugs rather than deliberate downgrades, and that a model as served through a subscription client is a different object from the raw API endpoint. That distinction points at the credible solo wedge: longitudinal tracking of coding agents as actually served on consumer plans, which is where most nerf reports originate and where cross-model leaderboards do not go. Trust is the entire product here — a single wrong “nerf confirmed” headline burns it permanently.