Lightweight CI-native eval harness for multi-step AI agents
Problem
Teams shipping production LLM agents find existing eval frameworks too heavy or too brittle once they move past single-turn prompts: testing multi-step trajectories, retries and non-deterministic state produces flaky LLM judges, slow test runs and unmaintainable snapshot diffs. There is no settled, CI-friendly harness for non-deterministic agent behavior.
Opportunity
An eval harness designed for CI from day one: deterministic replay of tool-calling trajectories, tolerance-based assertions for non-determinism, and fast incremental test runs that don't rot as prompts change.
Market analysis
Acute and widespread pain: agent teams openly complain about flaky LLM judges and rotting snapshot diffs in CI. The space is crowded (promptfoo, DeepEval, Langfuse), but the specific wedge of deterministic trajectory record/replay with tolerance-based assertions is sharper than what the generalist platforms offer, and small projects like agent-evals prove appetite without yet owning it.
Market · Engineering teams shipping production LLM agents; eval tooling is one of the hottest dev-infra categories with strong budget signals.
Pricing · Eval platforms are open core with paid hosted tiers (promptfoo, DeepEval, Langfuse all free OSS + SaaS); CI speed and caching are features teams demonstrably pay for.
Pros
- + Pain is specific, recurring and voiced by teams with budgets already spending on observability.
- + Deterministic replay of tool-call trajectories is a credible technical wedge against judge-based platforms.
- + Fast incremental CI runs directly address the slow-test complaint that eval platforms handle poorly.
Cons
- − promptfoo and DeepEval have large communities and are one roadmap item away from first-class replay.
- − Record/replay of non-deterministic, stateful trajectories is a deep engineering problem to keep correct.
- − Framework fatigue is real: teams resist adopting a fifth harness alongside pytest and their eval platform.
Existing / similar tools
- → promptfoo ↗
- → DeepEval ↗
- → agent-evals ↗
- → Langfuse ↗
Source
Hacker News (Ask HN)
The wedge here is real but time-boxed, so the build order matters more than the architecture diagram. The winning pitch is not “another eval framework” (teams are drowning in those); it is “your existing agent tests stop flaking and stop taking twenty minutes”. That means the unglamorous engineering is the product: cached deterministic replays keyed by trajectory hash, tolerance-based diffs instead of exact snapshots, incremental runs that skip what didn’t change. Sell it as a pytest/CI layer that wraps whatever harness the team already runs, because displacing promptfoo head-on is a losing fight, while making it reliable inside CI is something none of the incumbents do well yet. Do that before the platforms absorb replay, which they will within a couple of release cycles.