LLM Quality Nerf Tracker
Problem
AI model quality silently degrades after launch hype: quantization, request routing to older models, or prompt changes make outputs worse, but users only have anecdotes. In the SWE-2 benchmark thread an HN user observes this is quite common once PR hype wears off and suggests someone should build a 'nerf tracker'.
Opportunity
A service that continuously runs standardized, versioned eval suites against production AI model endpoints over time and alerts subscribers when a specific model version's quality drops — an independent model-quality watchdog for teams building on LLM APIs.
Market analysis
Eval platforms (Confident AI, Braintrust, Galileo) monitor your own pipeline and Artificial Analysis benchmarks models independently, but none alerts when a specific production endpoint quietly degrades; the nerf-tracker niche is real, narrow, and cheap to run.
Market · Engineering teams building on third-party LLM APIs plus AI newsletters and procurement; demand spikes every time a silent nerf allegation goes viral on HN or X.
Pricing · Dev-tool SaaS norms: observability and eval platforms run tens to hundreds of dollars per month per seat or usage; a focused alert subscription could plausibly charge $29 to $99 per month per team.
Pros
- + Operationally cheap: scheduled evals against public endpoints, no customer data plumbing.
- + Inherently viral: every caught nerf is a story that markets the product.
- + Independent watchdog positioning is hard to out-market once trusted.
Cons
- − Providers rarely expose model versions, so attribution is fuzzy without canary fingerprinting.
- − Artificial Analysis or any eval platform could add time-series alerts as a feature.
- − Willingness to pay for alerts alone, without a full observability suite, is unproven.
Existing / similar tools
Source
Hacker News (comment thread)
The hard part is not running evals, it is identity: without provider-side version identifiers, the claim “the same model got worse” requires fingerprinting endpoints with canary prompts and statistical drift detection, which is exactly the defensible trick a copycat cannot fake. That is also the credibility moat: a watchdog that cries wolf gets ignored, so the product needs confidence intervals and public methodology from day one. A solo builder can genuinely ship this, because the core loop is a cron, a prompt suite, and an alert feed, before any team dashboard.