Skip to content
IdeaScout.
← Back to archive

LLM API mock/replay server for load testing

AI-discovered

Problem

Teams stress-testing AI features face a dilemma: hitting real OpenAI/Claude APIs turns a load test into a token-burning exercise, so they end up hand-building and maintaining mock services of the provider APIs (streaming, tool calls, errors) just to test their own gateways, queues and retries.

Opportunity

A drop-in LLM API simulator: point your SDK base URL at it and get realistic streaming responses, tool-call flows, latency profiles and error rates at any concurrency — for a fraction of inference cost, with record/replay of real traffic.

Market analysis

Functional LLM mocks are already abundant — MockServer mocks OpenAI/Anthropic/Gemini/Bedrock with streaming and tool calls, and llm-mock-server, MockAI and OpenAI Mock API cover the OSS space — but they are correctness-testing tools. The load-testing angle (token-rate-accurate SSE pacing, latency distributions, injected error rates at thousands of concurrent streams, record/replay of real traffic) is a genuinely underbuilt layer.

Market · Platform/infra engineers at companies shipping AI features; demand is validated by the Ask HN thread itself admitting teams hand-build these mocks.

Pricing · Dev-tool norms: free OSS core plus paid team/cloud tier; comparable infra tooling monetizes at $50-500/month per org.

score 6/10 by glm-5.1

Pros

  • + Pain is admitted first-hand by the target audience on HN.
  • + Clear wedge above existing mocks: concurrency realism, latency/error profiles, record/replay.
  • + Solo-buildable — it's a server, no infrastructure moat needed for v1.

Cons

  • − Free OSS mocks are good enough for many teams, capping conversion.
  • − Providers could ship first-party sandbox/load-test modes (OpenAI already has a partial one).
  • − Dev-tool monetization is notoriously hard below team size.

Source

Hacker News (Ask HN)

Open original thread ↗

The engineering difference from existing mocks is the product: a functional mock can return instantly, but a load-testing simulator must reproduce token-rate-paced SSE with configurable TTFT and inter-chunk delays, backpressure behavior, and provider-accurate failure modes (429s with retry-after, mid-stream disconnects) at high concurrency — that’s a small amount of hard code nobody has written well yet. Record/replay of real traffic is the killer feature that turns it from a demo tool into something a platform team can’t rebuild in a weekend, and it’s also the natural open-core boundary: cassettes and OSS free, distributed load generation and CI integration paid.