Skip to content
IdeaScout.
← Back to archive

Automatic identification/matching for raw media rips

AI-discovered

Problem

People ripping their own Blu-rays must manually name and tag every episode before media servers (Plex, Jellyfin, Kodi) can identify them, even though the ripped data is bit-identical and a simple hash could uniquely identify it. In the Jellyfin 12.0 HN thread, a lifetime Plex user says they will switch 'as soon as someone has a good solution for identifying raw Bluray rips automatically' and asks why there is no CDDB-style hash database for raw video rips, 25+ years after CDDB solved this for audio CDs.

Opportunity

A CDDB-for-video service: a crowd-sourced hash database of disc rips (BDMV structures, MKV hashes) with an API that media servers and ripping tools can call to auto-name and tag releases. Also unlocks the broader 'identify this file by content hash' pattern for other media formats.

Market analysis

The gap is real — Automatic Ripping Machine identifies discs via title lookups against OMDb/TMDB, not content hashing, and episode naming on multi-episode discs is still manual — but the naive 'hash the file' premise is technically flawed, which is exactly why CDDB-for-video never happened.

Market · Self-hosted media enthusiasts (r/jellyfin, r/PleX, ARM users); a dedicated niche that already pays for media-server convenience (Plex Pass).

Pricing · Free OSS dominates the ripping stack (ARM, MakeMKV beta); comparable monetization is Plex-Pass-style: free API tier with paid hosted sync, or a one-time supporter license.

score 6/10 by glm-5.1

Pros

  • + Clear, unsolved gap confirmed by ARM's own maintainers and users.
  • + Natural integration point: ripping tools (ARM, MakeMKV wrappers) and media servers both want this API.
  • + Crowd-sourced database has classic network effects: more rippers, better coverage.

Cons

  • − File hashes break with encoding settings — identification needs stable disc-structure hashes or perceptual video fingerprinting, a much harder problem.
  • − Chicken-and-egg: the database is useless until it reaches critical mass, and ripping volume is declining.
  • − Copyright-adjacent service (hash DB of commercial disc content) may attract legal friction.

Source

Hacker News (thread: Jellyfin 12.0)

Open original thread ↗

CDDB worked because an audio CD’s table of contents is a stable, disc-level fingerprint independent of how you rip it. Video rips have no equivalent anchor: two MakeMKV runs with different title selections or a HandBrake re-encode produce completely different files, so a hash of the MKV is meaningless. The viable design hashes the source structure — BDMV playlist and clip files read straight off the disc, which are bit-identical across copies of the same pressing — and pairs it with a segment-level episode mapping contributed by the first person who names that disc. That is a genuine engineering moat, but also the reason 25 years of enthusiasts haven’t already built it.