Plan-vs-code review gate for AI-generated PRs
Problem
PR volume and PR size have exploded with AI-generated code, and human reviewers can no longer meaningfully review the diff itself. One team moved peer review to the implementation plan and validates PRs against it in CI, but had to build this workflow with internal custom tooling — when asked, the market answer was that no publicly available product does this.
Opportunity
A review gate product: peer-review the plan (not the diff), then automatically validate every PR against the approved plan in CI, classifying deviations (missing work, changed approach, extra scope) so only deviations need human attention.
Market analysis
Verified whitespace: the team that pioneered plan-review-over-diff-review confirmed they built internal tooling because no public product exists, and their measured economics (45 minutes of plan feedback vs ~16 hours of PR rework) make the ROI argument for you. Adjacent AI review tools (Tembo, Augment Cosmos) review diffs, not plan-faithfulness — this is a different, unoccupied gate.
Market · Teams with heavy AI-generated PR volume; demand is explicit (the HN thread literally asked for this product) and growing with agent adoption.
Pricing · No direct comparable priced yet; adjacent AI code-review SaaS implies per-developer team pricing (roughly $15-30/dev/month) with room for a plan-seat model.
Pros
- + Confirmed gap — early adopters built internal tooling and publicly asked for a product.
- + Sharp, measurable value: review effort moves upstream where feedback is 20x cheaper.
- + Solo-buildable as a GitHub Action + plan-file convention; no infrastructure-heavy MVP.
Cons
- − Deviation classification is LLM judgment — fuzzy output must earn trust before gating merges.
- − Garbage-in risk: the gate only validates against the plan, so plan quality becomes the bottleneck.
- − GitHub Copilot review / Tembo / Augment could add plan-faithfulness checks as a feature.
Existing / similar tools
Source
Hacker News (Ask HN)
The insight early adopters keep confirming: AI is more reliable as a reviewer than as a generator, so the scarce human attention should fund the plan, not the diff. But the product’s hard part is not the CI hook — it is the deviation taxonomy. ‘Changed approach’ can be the agent wisely adapting or silently ignoring the agreed design, and distinguishing those two is exactly the judgment reviewers were doing before. Ship the classifier as advisory-first (flag, never block) until its precision earns gating rights; a review tool that produces false alarms on merge day gets uninstalled the same week.