Automatic identification/matching for raw media rips
Problem
People ripping their own Blu-rays must manually name and tag every episode before media servers (Plex, Jellyfin, Kodi) can identify them, even though the ripped data is bit-identical and a simple hash could uniquely identify it. In the Jellyfin 12.0 HN thread, a lifetime Plex user says they will switch 'as soon as someone has a good solution for identifying raw Bluray rips automatically' and asks why there is no CDDB-style hash database for raw video rips, 25+ years after CDDB solved this for audio CDs.
Opportunity
A CDDB-for-video service: a crowd-sourced hash database of disc rips (BDMV structures, MKV hashes) with an API that media servers and ripping tools can call to auto-name and tag releases. Also unlocks the broader 'identify this file by content hash' pattern for other media formats.
Market analysis
The gap is real — Automatic Ripping Machine identifies discs via title lookups against OMDb/TMDB, not content hashing, and episode naming on multi-episode discs is still manual — but the naive 'hash the file' premise is technically flawed, which is exactly why CDDB-for-video never happened.
Market · Self-hosted media enthusiasts (r/jellyfin, r/PleX, ARM users); a dedicated niche that already pays for media-server convenience (Plex Pass).
Pricing · Free OSS dominates the ripping stack (ARM, MakeMKV beta); comparable monetization is Plex-Pass-style: free API tier with paid hosted sync, or a one-time supporter license.
Pros
- + Clear, unsolved gap confirmed by ARM's own maintainers and users.
- + Natural integration point: ripping tools (ARM, MakeMKV wrappers) and media servers both want this API.
- + Crowd-sourced database has classic network effects: more rippers, better coverage.
Cons
- − File hashes break with encoding settings — identification needs stable disc-structure hashes or perceptual video fingerprinting, a much harder problem.
- − Chicken-and-egg: the database is useless until it reaches critical mass, and ripping volume is declining.
- − Copyright-adjacent service (hash DB of commercial disc content) may attract legal friction.
Existing / similar tools
Source
Hacker News (thread: Jellyfin 12.0)
CDDB worked because an audio CD’s table of contents is a stable, disc-level fingerprint independent of how you rip it. Video rips have no equivalent anchor: two MakeMKV runs with different title selections or a HandBrake re-encode produce completely different files, so a hash of the MKV is meaningless. The viable design hashes the source structure — BDMV playlist and clip files read straight off the disc, which are bit-identical across copies of the same pressing — and pairs it with a segment-level episode mapping contributed by the first person who names that disc. That is a genuine engineering moat, but also the reason 25 years of enthusiasts haven’t already built it.