Skip to content
IdeaScout.
← Back to archive

Lightweight CI-native eval harness for multi-step AI agents

AI-discovered

Problem

Teams shipping production LLM agents find existing eval frameworks too heavy or too brittle once they move past single-turn prompts: testing multi-step trajectories, retries and non-deterministic state produces flaky LLM judges, slow test runs and unmaintainable snapshot diffs. There is no settled, CI-friendly harness for non-deterministic agent behavior.

Opportunity

An eval harness designed for CI from day one: deterministic replay of tool-calling trajectories, tolerance-based assertions for non-determinism, and fast incremental test runs that don't rot as prompts change.

Market analysis

Acute and widespread pain: agent teams openly complain about flaky LLM judges and rotting snapshot diffs in CI. The space is crowded (promptfoo, DeepEval, Langfuse), but the specific wedge of deterministic trajectory record/replay with tolerance-based assertions is sharper than what the generalist platforms offer, and small projects like agent-evals prove appetite without yet owning it.

Market · Engineering teams shipping production LLM agents; eval tooling is one of the hottest dev-infra categories with strong budget signals.

Pricing · Eval platforms are open core with paid hosted tiers (promptfoo, DeepEval, Langfuse all free OSS + SaaS); CI speed and caching are features teams demonstrably pay for.

score 6/10 by glm-5.1

Pros

  • + Pain is specific, recurring and voiced by teams with budgets already spending on observability.
  • + Deterministic replay of tool-call trajectories is a credible technical wedge against judge-based platforms.
  • + Fast incremental CI runs directly address the slow-test complaint that eval platforms handle poorly.

Cons

  • − promptfoo and DeepEval have large communities and are one roadmap item away from first-class replay.
  • − Record/replay of non-deterministic, stateful trajectories is a deep engineering problem to keep correct.
  • − Framework fatigue is real: teams resist adopting a fifth harness alongside pytest and their eval platform.

Source

Hacker News (Ask HN)

Open original thread ↗

The wedge here is real but time-boxed, so the build order matters more than the architecture diagram. The winning pitch is not “another eval framework” (teams are drowning in those); it is “your existing agent tests stop flaking and stop taking twenty minutes”. That means the unglamorous engineering is the product: cached deterministic replays keyed by trajectory hash, tolerance-based diffs instead of exact snapshots, incremental runs that skip what didn’t change. Sell it as a pytest/CI layer that wraps whatever harness the team already runs, because displacing promptfoo head-on is a losing fight, while making it reliable inside CI is something none of the incumbents do well yet. Do that before the platforms absorb replay, which they will within a couple of release cycles.