AI EngineeringGuides & Tutorials

Test Your AI Before Users Do: Evaluation Tools for Hallucinations, RAG, and Agents (2026)

OpenAI just shut down its Evals platform and bought Promptfoo. Here is the current eval tool landscape, from Braintrust and LangSmith to DeepEval and RAGAS, plus a fixed-benchmark starter suite with factual, adversarial, structured-output, and tool-use cases you can build this week.

Toolbit AI - Team
13 min read
Test Your AI Before Users Do: Evaluation Tools for Hallucinations, RAG, and Agents (2026)

Your LLM application is almost certainly failing right now, and nobody has noticed. Not crashing, not erroring, not showing up in your uptime dashboard. Just quietly answering worse than it did last month, in ways no metric you currently track can see.

That is the default state of production AI. A prompt gets "improved" on Friday. A model gets swapped because a new one shipped with better benchmarks. A retrieval index gets re-chunked to save tokens. Each change looks fine in the playground demo. Each one can regress the specific cases your users actually care about, and nothing pages you, because the answer still comes back fluent and confident. We have covered before why AI hallucinates in the first place; the harder engineering question is how you catch it when it starts happening in your product.

The answer is evaluations: systematic tests you run against your AI system, the way you run unit tests against code. The tooling around this matured fast, and September 2026 is a genuinely good moment to set up an eval stack, partly because one of the biggest vendors just forced everyone's hand. Here is what the landscape actually looks like, which tools are worth your time for hallucinations, RAG, and agent workflows, and a concrete starter suite you can build this week.

Why evals, and why your dashboard does not count

Traditional software testing works because behavior is deterministic: same input, same output, assert and move on. LLMs break that contract. The same prompt can produce a great answer and a fabricated one. Fluency is not correctness, so a response that reads beautifully can be completely wrong, and latency graphs, error rates, and token costs, the things most teams monitor, are all silent on quality.

Evaluations close that gap with three ingredients: a fixed dataset of test cases, a way to score outputs, and a place to run the comparison. When you have those, questions that were previously vibes become numbers. Did the new prompt make answers more grounded or less? Is the cheap model actually worse on your support tickets, or just worse on public benchmarks? Did the re-index hurt anything? Without evals you are guessing; with them, you have a regression suite for a system that has no deterministic behavior to assert against.

The other reason this matters right now: OpenAI is shutting down its Evals platform. The official docs now carry a deprecation notice: the platform goes read-only for existing users on October 31, 2026, and fully shuts down on November 30, 2026. If your team's eval workflow lives in the OpenAI dashboard, you have weeks to export your datasets and graders and re-anchor them somewhere else. And in a telling move, OpenAI's recommended path for agentic security testing now runs through Promptfoo, which OpenAI acquired in March 2026. The center of gravity in this space just moved, which makes it a good time to choose your harness deliberately instead of defaulting into one.


The four layers of an eval stack

Before picking tools, it helps to know what you are actually building. Mature AI teams layer four kinds of checking, and the tools below each emphasize different layers.

Layer 1: Unit assertions. Deterministic checks that run in CI: does the output parse as valid JSON, does it contain the required fields, does it refuse when it should refuse, does it stay under the length limit. Cheap, instant, no judge model needed. Most teams under-invest here; a surprising share of "model quality" bugs are actually schema and contract bugs that plain asserts catch.

Layer 2: LLM-as-judge scoring. A second model scores outputs against a rubric: faithfulness to retrieved context, answer relevance, tone. This is the workhorse for hallucination detection, because a judge can check whether every factual claim in an answer is supported by the retrieved documents. One warning that current research makes sharply: a large 2026 audit of 21 judge models found that highly consistent judges (test-retest reliability above 0.95) can still carry severe position bias, favoring whichever answer appears first, and judge rankings shift by up to 14 positions across benchmarks. A stable judge is not necessarily a valid one. Keep the judge model fixed across comparisons, swap positions in pairwise tests, and spot-check scores against human labels before trusting a threshold.

Layer 3: Regression suites. A fixed set of cases, run on every prompt change, model swap, or retrieval tweak, gated in CI. This is where evals stop being an experiment and start being tests. It is also the layer most teams skip, which is why quality erodes silently over months. If you have ever wondered why AI agents fail in production in ways staging never showed, the missing regression suite is usually the root cause.

Layer 4: Production tracing and sampling. Once you ship, you need visibility into real traffic: traces of every chain, agent step, and tool call, plus evals run on a sample of production outputs to catch drift. This is what Braintrust, LangSmith, and Phoenix are fundamentally built around, and it pairs naturally with the context engineering discipline of treating prompts and retrieved context as versioned, observable artifacts.

Diagram of the four layers of an AI eval stack, from unit assertions and LLM-as-judge scoring to regression suites and production tracing

The tool landscape: what each one is actually best at

ToolTypeBest atWeakest atOpen source
BraintrustEval + tracing platformExperiments, regression comparison, scoring at scaleNo self-host below enterpriseNo
LangSmithTracing + eval platformLangChain/LangGraph agent lifecycleCost at high trace volumeNo (Enterprise only)
Arize PhoenixOSS observability + evalsTracing-first debugging, self-hosting, agent evalsLess turnkey than hosted rivalsYes
PromptfooOSS CLIRed-teaming, adversarial testing, YAML-driven CI evalsNot a production monitoring platformYes (MIT)
DeepEvalOSS pytest frameworkCI-gated metric tests, RAG + agent metricsNo tracing/UI story of its ownYes (Apache 2.0)
RAGASOSS metric libraryRAG component diagnosis: retrieval vs generationNarrow: RAG metrics onlyYes

Braintrust: the evals-first platform

Braintrust meters you on scores, not just traces, which tells you what it is optimized for: running lots of evaluation. The free Starter tier includes $10 of model credits, 1 GB of processed data, and 10,000 scores per month with unlimited users and projects; the Pro tier is $249/month with 5 GB and 50,000 scores included, per its pricing page. The core loop is the experiment: point two prompt versions or two models at the same dataset, get side-by-side scores, promote the winner, and keep the history as your regression baseline. If your team's bottleneck is "we can't compare options rigorously," this is the strongest hosted answer. Where it is weaker: if you want self-hosting without a sales conversation, it is not your tool.

LangSmith: deepest if you live in LangChain

LangSmith is the commercial platform behind LangChain, and its strength is the agent lifecycle: tracing for LangGraph workflows is best-in-class, and it now spans deployment and evaluation in one product. Pricing is per-seat friendly at the start: free Developer tier with 5,000 base traces monthly, then $39 per seat for Plus with 10,000 base traces included, per LangChain's pricing page. The catch is scale: trace overage and extended retention get expensive fast, a recurring complaint in independent comparisons, and self-hosting only appears at Enterprise. Choose it when you are committed to the LangChain ecosystem and want everything integrated; think twice if you are evaluating across many providers at volume.

Arize Phoenix: the open-source observability bet

Phoenix is Arize's open-source AI observability and evaluation platform, and it takes the opposite stance: self-host the whole tracing + eval stack, free, with a managed option (Arize AX) if you want it hosted. Its eval approach is unusual in a good way: evaluators are attached to traces, judges are swappable via adapters (OpenAI, LiteLLM, LangChain, and others), and every judge run is itself traced, so you can audit why the judge scored something the way it did, per its evaluation docs. It ships pre-built evaluators for RAG (faithfulness, retrieval relevance) and, notably, for tool-calling agents: tool selection, invocation, and response-handling checks. If you want an open-source LangSmith alternative you can run in your own VPC, this is the current default answer.

Promptfoo: red-teaming in a YAML file

Promptfoo describes itself as a CLI for test-driven LLM development, and its two modes map to the two halves of quality: evaluation (compare prompts and models with assertions) and red-teaming (attack your own app before other people do). The declarative YAML config is genuinely pleasant: prompts, providers, and pass/fail criteria live in one reviewable file, and a non-zero exit fails the build, so CI integration is trivial. The red-teaming side is its signature strength, probing for prompt injection, jailbreaks, PII leakage, and excessive agency, which pairs naturally with reading up on prompt injection as a hidden security risk. The governance footnote you already know: OpenAI acquired Promptfoo in March 2026 and is folding the enterprise capabilities into its Frontier platform, while the MIT-licensed CLI continues to ship. If you standardize on it, pin your version and keep configs in your repo.

DeepEval: pytest, but for LLMs

DeepEval is the tool for teams whose quality gate is a CI pipeline. It wraps evaluation in pytest: you write test functions, assert metric thresholds, and run deepeval test run, which adds caching and parallelism and exits non-zero on failure. The metric library is the broadest of the open-source frameworks: G-Eval for custom rubrics, a full RAG suite (answer relevancy, faithfulness, contextual precision and recall), agentic metrics (task completion, tool correctness, argument correctness, plan adherence), and wrappers for benchmarks like TruthfulQA and GSM8K. It comes from an independent startup (Confident AI), which matters if vendor neutrality is part of your calculus this year. What it does not give you is production tracing; pair it with Phoenix or a hosted platform for that layer.

RAGAS: the RAG specialist

RAGAS is narrower and unashamed of it: it is the metric library that gave RAG evaluation its shared vocabulary. Faithfulness (is the answer grounded in retrieved context?), answer relevancy, and contextual precision and recall (did retrieval fetch the right chunks, ranked well?) are its core, and its docs pitch the move from "vibe checks" to systematic eval loops. Its diagnostic split is the useful part: low faithfulness points at a hallucinating generator, low contextual recall points at a retriever that missed the needed chunks. The common production pattern you will see in the wild is RAGAS for retrieval tuning during development and DeepEval or another harness for the CI gate. One known quirk: judge JSON parse failures can produce NaN scores; newer tooling, including DeepEval's own RAG metrics, works around this.


A starter eval suite you can build this week

The angle that separates teams that benefit from evals from teams that collect dashboards: a fixed benchmark. Hold the test cases constant, change the system, and compare numbers. If the benchmark changes every run, you are measuring nothing. Here is a four-case-type starter suite, sized to be genuinely runnable in a few days.

Step 1: Build the golden dataset. Collect 30 to 50 real inputs from your actual domain: support tickets, user queries, documents you really process. Not synthetic, not generic. Split them into four case types, roughly equal:

  1. Factual cases. Questions with a verifiable right answer, drawn from your own docs or data. These ground your accuracy baseline.
  2. Adversarial cases. Tricky inputs: unanswerable questions the system should decline or hedge, attempts to make the model fabricate, edge phrasings, prompts with embedded instructions. These are your hallucination stress test.
  3. Structured-output cases. Inputs that must produce valid JSON, specific fields, or constrained formats. These are scored by plain code: parse, validate schema, pass or fail. No judge needed.
  4. Tool-use cases. Agent scenarios where the right tool must be called with the right arguments. Scored against an expected trajectory: did it call lookup_order before issue_refund? Did it pass the right ID?
The four case types of a starter eval suite: factual, adversarial, structured-output, and tool-use questions with their scoring methods

Step 2: Score with two judges, not one. Run LLM-as-judge scorers for the factual and adversarial cases, using the strongest model you can justify as judge, never the model under test. Two judges, or one judge run with positions swapped, given the position-bias findings above. For structured and tool-use cases, keep it deterministic: schema validation and trajectory matching are free and unambiguous.

Step 3: Gate it in CI. Wire the suite to fail builds: DeepEval's pytest integration or Promptfoo's YAML and exit codes both do this with an afternoon of work. Define thresholds from your first baseline run, not from ambition. If faithfulness lands at 0.82 on day one, your threshold is 0.80, and you improve it from there. A suite that blocks merges on day one is a suite that gets deleted by day three.

Step 4: Add production sampling. Once shipped, sample 1 to 5 percent of real traffic into the same scoring pipeline, with tracing on. This catches drift, the failure mode where nothing changed in your code but the world moved under you: users ask different questions, the docs corpus shifts, a provider quietly updates a model.

The exact tool choice for each step matters less than the shape. A reasonable all-open-source path: Phoenix for tracing and judge auditing, DeepEval for the CI gate, RAGAS when you need to debug retrieval specifically, Promptfoo when security testing joins the requirements. A reasonable hosted path: Braintrust or LangSmith for the platform, Promptfoo or DeepEval retained for CI-local runs so your gate never depends on a vendor's uptime.


What tends to go wrong

The judge disagrees with reality. LLM-as-judge scores are relative signals, not ground truth. The 2026 judge audit found rankings that swing by double-digit positions across benchmarks; the same applies to your rubric. Before trusting any threshold, hand-label 20 outputs yourself and compare. If judge and human agree under 80 percent of the time, fix the rubric, not the model.

The suite grows stale. A benchmark that never updates stops representing your users. Budget a monthly review: pull fresh failure cases from production traces into the golden dataset. The suite is a living artifact, not a compliance checkbox.

You evaluate the demo, not the system. Single-turn Q&A evals miss most agent failures, which live in multi-turn context, tool selection, and error recovery. If your product is agentic, your suite needs trajectory-level cases, and your tracing needs to capture intermediate steps, not just final answers.

Vendor lock quietly returns. Two of the three most-adopted open-source eval tools changed owners in early 2026. The insurance is the same in every case: keep your datasets, rubrics, and configs in your own repo, version-controlled, provider-neutral, so switching a harness is a weekend, not a migration program.

Two quick answers

Can I fully automate hallucination testing? Mostly. Faithfulness metrics (DeepEval, RAGAS, Phoenix all ship them) automatically check whether every claim in an answer is supported by retrieved context, and that catches the majority of grounding failures. But adversarial cases, subtle context conflicts, and tone-of-voice judgments still need periodic human review. Automate the floor, human-review the ceiling.

Do I need a hosted platform, or is open source enough? Open source covers evaluation itself completely: DeepEval or Promptfoo in CI, RAGAS for RAG diagnosis, Phoenix for tracing, all self-hostable at zero license cost. Hosted platforms earn their fee when multiple people need to review outputs, annotate, and share experiments, or when you want managed production sampling with retention and compliance handled. Solo project: stay open source. Team shipping to users: the platform tax is usually worth it.

Pricing and plan details are as published by the vendor around September 2026 and can change: confirm on the official site.

Share this article

Related articles

Continue exploring similar guides and insights