Skip to content
IdeaScout.
← Back to archive

Model-to-workflow router for LLM-backed apps

AI-discovered

Problem

Teams shipping production LLM apps end up with dozens of AI workflows (the OP has 54 in a single Django app), each silently coupled to one provider/model choice. As new models ship weekly with shifting speed/intelligence/cost tradeoffs, there is no systematic way to track which workflow uses which model, what each workflow optimizes for, and whether the currently assigned model is still the right one. The OP hand-maintains a registry of workflow purpose, optimization target, eval and quality bar — pure manual spreadsheet work.

Opportunity

A model-routing control plane: a registry that maps every LLM call site/workflow to its optimization goal (cost, latency, intelligence), runs evals per workflow, and flags or auto-switches assignments when a cheaper/faster model passes the bar. Sells directly to eng teams already paying multiple AI vendors.

Market analysis

Per-request model routers (NotDiamond, Martian, OpenRouter Auto) are crowded and well funded, but the OP's problem is different and mostly unaddressed: an asset registry of workflow intent with per-workflow eval gates. The wedge is governance of existing call sites, not routing new traffic.

Market · Engineering teams running many production LLM workflows across multiple providers; demand signal visible in fast-growing gateway and LLM observability spending.

Pricing · Routers charge platform fees on top of model costs (around $20 per 5,000 requests for auto-router services); eval and observability platforms use per-seat plus usage pricing, which sets buyer expectations.

score 6/10 by glm-5.1

Pros

  • + Real, felt pain: model sprawl across dozens of call sites is now normal for production teams.
  • + Sits between routers (per-request) and observability (post-hoc), a genuine gap neither category fills.
  • + Registry is a natural beachhead into evals, cost reports and compliance.

Cons

  • − Router vendors and gateways are converging on this space from below with big budgets.
  • − Value hinges on per-workflow evals, the hardest part to build and to sell.
  • − Enterprises may expect this to arrive inside their existing LLM observability stack.

Existing / similar tools

Source

Hacker News (Ask HN)

Open original thread ↗

The hard part is not the registry, it is the eval: a control plane that cannot prove a cheaper model still passes the quality bar is just a spreadsheet with extra steps. A solo builder should ship the registry plus drift alerts first, since cataloguing call sites and flagging stale assignments is cheap and instantly useful, and treat auto-switching as a later opt-in step. Silent model changes are exactly what production teams fear most, so the safest product posture is “recommend and prove”, never “switch on its own”. If the eval story works, this becomes the system of record for every AI workflow in the company, which is a far stickier position than being a router.