Skip to content
IdeaScout.
← Back to archive

Anti-LLM-Training License & Opt-Out Compliance Layer

AI-discovered

Problem

Open-source maintainers want to publish code while preventing AI labs from ingesting it into training data, but no license or mechanism reliably does this today. Comments confirm the consensus: crawlers don't even look at licenses, enforcement is impossible for individuals, and existing licenses are 'a gentleman's agreement in a world of no gentlemen.'

Opportunity

A compliance/licensing toolkit for the AI era: license templates with registered, timestamped terms; a robots/ai-txt opt-out manifest standard; a public registry of opted-out artifacts; and monitoring that detects when opt-out code surfaces in model outputs — giving maintainers evidence and standing they currently lack.

Market analysis

The demand is emotional and real, but the evidence says the whole anti-AI-licensing stack is failing on enforcement: ~30% of AI scrapes ignore robots.txt outright, the OSI holds that no-AI-training clauses are incompatible with open source, and the practical conclusion in the field is that once code is public it is effectively available for training regardless of license terms. A registry plus output monitoring gives evidence, not prevention — and the buyer who needs that evidence (a law firm or platform) is not the maintainer feeling the pain.

Market · OSS maintainers and small rightsholders wanting training-data control; high sentiment, but the entity that would actually pay for evidence (enterprises, lawyers, platforms) is one step removed from the pain.

Pricing · Maintainers expect this to be free or near-free (opt-out lists like ai.robots.txt are community-maintained); comparable monitoring services exist at the API level (e.g. Patronus AI CopyrightCatcher) but sell to LLM vendors, not maintainers.

score 3/10 by glm-5.1

Pros

  • + Strong, recurring community sentiment every time AI scraping hits the news.
  • + Evidence-gathering (registry + output monitoring) is technically tractable for a solo builder.
  • + EU AI Act opt-out mandates may create compliance-driven demand post-2026.

Cons

  • − No enforcement teeth: scrapers ignore opt-outs, individuals cannot afford litigation.
  • − OSI position means restricted licenses forfeit open-source status, splintering the audience.
  • − Free community alternatives already exist (ai.robots.txt list, GitHub opt-out settings).
  • − Willingness to pay sits with maintainers who expect it free, not with those needing evidence.

Source

Hacker News (Ask HN)

Open original thread ↗

The uncomfortable truth from the field: anti-AI licensing is failing not for lack of tools but because it tries to solve an architectural problem with a legal one. The only mechanism with teeth is provenance tracking inside training pipelines — which is why the realistic buyer is the AI lab needing to demonstrate opt-out compliance (EU AI Act), not the maintainer. A solo builder could pivot this into “opt-out compliance evidence as a service” selling to model providers: registry, timestamping, and audit logs that let a lab prove it honored reservations. That flips the customer from someone who cannot pay to someone legally required to.