Here is the short answer: evaluating an agent means scoring four dimensions - task success, tool-use correctness, cost and latency, and safety - on a fixed set of real traces, before and after every change. Checking only the final answer is not evaluation. A right answer can hide a wrong tool call, ten times the tokens you budgeted, or two agents politely passing the same task back and forth in a loop. OpenAI's evaluation best practices name circular agent handoffs among the classic failure modes to evaluate for, and they give the opposite approach its proper name: "vibe-based evals," using "it seems like it's working" as an evaluation strategy. It is a named anti-pattern.
If you have ever shipped an agent on vibes and hoped for the best, the most common automation mistakes are worth a read - skipping the failure check is the oldest one in the book.
By the end of this post you will know how to evaluate AI agents in practice: what to measure, how to build an eval set from real traffic, how to run a loop that catches regressions before users do, and you will have a small harness you can copy today.
In short:
- Evaluate four dimensions: task success, tool-use correctness, cost and latency, safety. Any single score will miss agent failures.
- Start with 5-10 hand-curated examples per critical component, then mine production traces to grow the set.
- Run the loop on every change: baseline, change, rerun, compare. Ship only if nothing regressed.
- Treat cost and latency budgets like unit-test assertions.
- OpenAI's Evals platform becomes read-only for existing users on October 31, 2026 and is scheduled to shut down November 30, 2026 - do not build new eval infrastructure on it.
Measure these four dimensions

Agent evaluation means scoring four dimensions together - task success, tool-use correctness, cost and latency, and safety - because any single metric misses agent failures. An agent is a trajectory, not an answer. It plans, calls tools, retrieves documents, hands work to other agents, and then - maybe - produces an answer. DeepEval, an open-source evaluation framework, is built around exactly this idea: it grades complete agent trajectories across every decision and action, and it grades individual steps too, such as LLM calls, tool use, retrieval, and sub-agent handoffs. LangSmith makes the same point from the other direction: break your system into its critical components - LLM calls, retrieval steps, tool invocations, output formatting - and define quality criteria for each one.
The four, with their headline agent metrics:
- Task success - did the agent actually achieve the goal? Headline metric: task completion rate.
- Tool-use correctness - right tool, right arguments, no wasted steps? Headline metric: tool-call accuracy.
- Cost and latency - what did the whole trajectory cost and how long did it take? Headline metric: cost and p95 latency per task.
- Safety - did it stay inside the lines? Headline metric: incident count per 1,000 tasks.
If your agent reaches its tools over MCP, a plain-English tour of the MCP protocol is useful background here - every tool call over that wire is exactly what your tool-use evaluation needs to grade.
Score task success without a research team
Task success rate asks one question: did the agent achieve the goal? Score it with deterministic checks where you can and a judge where you must. DeepEval ships two metrics for this. Task Completion evaluates whether the agent accomplished its goal, and Goal Accuracy measures how precisely it hit the intended target. If you want per-criterion scoring, OpenAI's eval runs hand you result counts - total, passed, failed, errored - so every example lands in a clear bucket.
The hard part is open-ended output. Exact-match grading dies the moment two correct answers use different words, so the standard move in LLM evaluation is LLM-as-a-judge: a strong model reads the agent's output against your criteria and scores it. The good news comes from the researchers who studied this carefully: strong LLM judges like GPT-4 match human preferences at over 80% agreement - the same level of agreement humans reach with each other (Zheng et al., arXiv).
The same paper documents the honest caveats. Position bias: judges favor the first answer they read. Verbosity bias: longer answers look better. Self-enhancement bias: a model grades its own family's output more generously. So practice judge hygiene, straight from the paper: swap the answer order and check that the verdict holds, control for answer length, and never use the same model family as both judge and contestant.
One more knob worth knowing: pass@k. It is the fraction of problems solved at least once within k samples from the same model. Repeated sampling trades cost for success - on the HumanEval coding benchmark, one sample per problem from Codex solved 28.8% of them, while 100 samples per problem solved 70.2% (Chen et al., arXiv). Treat it as a deliberate cost dial, not a free lunch.
Grade the tool calls, not just the outcome
A right answer reached through the wrong tool is a production bug waiting for its moment. The agent queried the search API when it should have used the database, passed a malformed argument that happened to work anyway, or burned four unnecessary steps to look thorough. The output looks fine. The trajectory is quietly wrong - and the day the environment changes, it breaks.
So grade the calls, not just the outcome. The metric set, assembled from DeepEval's agentic metric families and LangSmith's guidance:
- Tool correctness - did it call the right tool?
- Argument correctness - were the arguments well-formed and sensible?
- Step efficiency - did it take unnecessary steps?
- Plan adherence - did it follow the plan it laid out?
LangSmith describes good agent examples as ones with "correct tool selection and proper argument formatting" - the trajectory matters as much as the destination.
The cheap way to grade all this: record trajectories from your own traces, then split the check in two. A deterministic check verifies the expected tool sequence - pure code, no model needed. A judge handles the fuzzy part, like whether an argument's value was sensible given the input. Deterministic where you can be, judge where you must be. That is the whole trick, and it keeps this dimension from becoming a research project.
Track cost and latency per task, not per prompt
Per-prompt numbers lie for agents, because retries, loops, and handoffs multiply cost invisibly. A prompt that costs 500 tokens sounds cheap until the agent retries it three times, loops once, and hands it to a sub-agent for good measure. The honest unit of accounting is the task: add up every token across every step, and measure every wait across every step.
The metrics:
- Cost per completed task - total tokens across all steps. OpenAI's eval runs expose per-model usage directly: prompt tokens, completion tokens, total tokens, and cached tokens, plus the invocation count - so token accounting is built into the run itself.
- p50 and p95 latency per task - the median tells you the typical experience; the 95th percentile tells you the worst one users actually get.
- Steps per task - a simple count that catches loops and runaway retries before they show up on an invoice.
LangSmith run traces carry latency metrics alongside the rest, so this data exists the moment you trace your agent.
Then set budgets and treat them like unit-test assertions. The eval fails the change if cost per task or p95 latency regresses past the baseline - not "looks a bit high," but fails, loudly, in CI. Budgets that live in a document get ignored; budgets that live in a test get enforced.
If your costs keep creeping, the context engineering habits replacing prompt engineering are the practical fix.
Test the failure modes before users find them
Users and attackers will find your edges eventually. Safety evaluation - the adversarial half of agent testing - means getting there first, on purpose. OpenAI's best practices tell you to define evals for the use cases you specifically support and the ones you specifically block - and the block list deserves its own test set:
- Jailbreak pass rate - how often do attempts to redirect the model actually work?
- Refusal correctness - does the agent decline out-of-scope requests cleanly, without over-refusing legitimate ones?
- Prompt-conflict behavior - when the user prompt and the system prompt disagree, which one wins?
Build the safety eval set once: a curated batch of adversarial and out-of-scope prompts. Rerun it on every change, the same as the rest of the eval set. Count incidents per 1,000 tasks so the number stays comparable across runs of different sizes.
One honest note from the tooling itself: DeepEval ships red-teaming support, but guardrails still sit on its roadmap rather than in the shipped feature list. The blocking layer - actually refusing the bad request at runtime - is one you may need to build yourself. Evaluation tells you the door is unlocked; you still have to install the lock.
Build your eval set from real traces
Your eval set comes from two sources: hand-curated ground truth first, real production traces second.
First, curate by hand. LangSmith's guidance is specific: start with 5-10 examples of what "good" looks like for each critical component of your system - each LLM call, each tool invocation, each retrieval step. These are your ground truth. Ten great examples beat a hundred mediocre ones, because every example you write encodes a judgment about what quality means, and that judgment is the whole point.
Second, mine production. OpenAI's best practices say it plainly: log everything as you develop so you can mine your logs for good eval cases. The named anti-pattern is creating eval datasets that do not faithfully reproduce production traffic patterns - a dataset of toy prompts tells you nothing about your real agent.
What does a dataset look like? LangSmith defines it cleanly: a dataset is a collection of examples, and an example is a test input plus a reference output, with optional metadata for filtered views. Phoenix describes the same loop from the tracing side: you group traces into datasets and rerun them through different versions of your application.
The workflow that makes it hum, per LangSmith:
- Flag bad runs for review as they happen in production.
- Annotate them - annotation queues collect human judgments on the flagged runs.
- Promote them - transfer the annotated runs into datasets for future evaluations.
And one rule that separates teams that improve from teams that thrash: every production failure becomes a new eval case before you fix it. Fix the bug, sure - but first capture the failure, so it can never silently return. LangSmith calls this the online-to-offline loop: online evaluations surface issues that become offline test cases, and offline evaluations validate the fixes. (LangSmith evaluation docs)
Run the evaluation loop before every change

Here is the whole discipline in one sentence: baseline, change, rerun, compare - and ship only if nothing regressed on any dimension.
OpenAI's eval workflow is three steps: describe the task as an eval, run it with test inputs, then analyze and iterate - which the guide itself compares to behavior-driven development. Their best-practices guide stretches it to a five-part process: define the objective, collect the dataset, define the metrics, run and compare - then set up continuous evaluation that runs on every change and grows the eval set over time.
LangSmith turns the loop into code. Evaluation metrics can be converted into tests: regression tests assert that new versions must outperform baseline versions on the relevant metrics, and they run inside pytest like any other test. Phoenix does the same version-vs-version comparison through experiments - run two versions on the same inputs and see the delta side by side.
The steps, concretely:
- Baseline - run the full eval set against the current version and store the results.
- Change - tweak the prompt, swap the model, adjust a tool.
- Rerun - the full eval set, same inputs, same criteria.
- Compare - every metric on every dimension against the stored baseline.
- Ship or revert - ship only if no dimension regressed. Otherwise, revert or fix.
The regression check is the whole point. LangSmith names regression testing as a core purpose of offline evaluation: ensure new versions do not degrade quality. Without the stored baseline, "it feels better now" is just vibe-based evals wearing a lab coat.
Know what the tools actually do - and what is being deprecated
Before picking tools, one time-sensitive fact: OpenAI is deprecating its Evals platform. It becomes read-only for existing users on October 31, 2026, and the platform is scheduled to shut down on November 30, 2026, with Datasets named as the successor direction (OpenAI's evals guide states this directly). If you were planning to build new eval infrastructure on OpenAI Evals, plan on something else.
What do the alternatives actually do? Here is the map, from each tool's own docs:
| Tool | What it is | What it gives you |
|---|---|---|
| LangSmith | Evaluation platform in the LangChain ecosystem | Offline evals on datasets, online evaluators on live runs, experiments, pytest integration, annotation queues |
| Arize Phoenix | Open-source, OpenTelemetry-based tracing and eval | Tracing of model calls, tool use, and retrieval; datasets and experiments; runs locally at localhost:6006; brings Ragas/DeepEval evaluators |
| DeepEval | Open-source Python framework (Apache-2.0) | Pytest-style CLI, agentic metric families including Tool Correctness and Task Completion, G-Eval and custom metrics; syncs to Confident AI |
How to choose? Skip rankings and look at three criteria:
- Open-source vs SaaS - Phoenix and DeepEval are open-source and self-hostable; LangSmith is primarily a hosted platform. If your traces cannot leave your infrastructure, that decides it.
- Code-first vs UI - DeepEval is pytest in spirit; LangSmith offers a UI for annotation and experiment comparison. Teams that live in CI tend one way, teams that review in a browser the other.
- Agent-trajectory support - DeepEval grades full trajectories and steps; Phoenix traces them via OpenTelemetry; LangSmith breaks systems into components. All three fit agent work, but their defaults differ.
Ship a minimal harness you can copy today
You do not need a platform to start. Four confirmed primitives make a complete harness:
- A JSONL dataset of input and expected-output pairs, seeded from your traces the way the previous section describes.
- A pytest loop over the examples. For each one, record pass or fail per criterion, latency, token usage, and the tool calls made.
- Stored results per run, diffed against a stored baseline run before you ship any change.
- Auto-append every production failure to the dataset - the online-to-offline loop, running on its own.
Here is the shape of it, illustrative and vendor-free:
# illustrative - the shape of a minimal eval harness
import json, time, pytest
CASES = [json.loads(l) for l in open("eval_set.jsonl")]
BASELINE = json.load(open("baseline_results.json"))
def run_case(case):
start = time.time()
result = my_agent(case["input"]) # your agent, however you call it
return {
"passed": judge_or_check(case, result), # criteria from this post
"latency": time.time() - start,
"tokens": result.usage.total_tokens, # or your framework's counter
"tool_calls": [c.name for c in result.steps],
}
@pytest.mark.parametrize("case", CASES)
def test_no_regression(case):
now, before = run_case(case), BASELINE[case["id"]]
assert now["passed"], "task success regressed"
assert now["latency"] <= before["latency"] * 1.1, "latency budget blown"
assert now["tokens"] <= before["tokens"] * 1.1, "cost budget blown"
Grow it from there: add a tool-sequence check once you record trajectories, add the safety prompt set once you curate it, add pass@k measurement if sampling fits your workload. The harness is a seed, not a ceiling.
You now have the four dimensions, the eval set, the loop, and the harness. Evaluation is a habit, not a project - run the set before every change, feed every failure back in, and the question stops being "it seems like it's working" and starts being "nothing regressed, ship it."
FAQ
Is OpenAI's Evals platform shutting down?
Yes. It becomes read-only for existing users on October 31, 2026, and is scheduled to shut down on November 30, 2026, with Datasets named as the successor direction. Do not build new eval infrastructure on it - pick a tool that will still be there next year.
What is the difference between evaluation and testing for LLM apps?
Evaluation measures performance according to metrics; testing asserts correctness. The practical bridge: convert your metrics into tests, so a regression test asserts that a new version outperforms the baseline and fails loudly in pytest when it does not.
How many eval examples do I need to start?
Start with 5-10 hand-curated examples per critical component. They are your ground truth. Then grow the set by mining production traces, and keep growing it with every failure the loop catches.
Is LLM-as-a-judge reliable?
Yes, with care. Strong judges like GPT-4 reach over 80% agreement with human preferences - the same level humans agree with each other - but they carry position, verbosity, and self-enhancement biases. Swap answer order, control for length, and never judge with the same model family you are testing.
What is pass@k and why does sampling more help?
Pass@k (pass at k) is the fraction of problems solved at least once within k samples. On HumanEval, one sample from Codex solved 28.8% of problems while 100 samples solved 70.2% - but 100 samples also cost 100 times more. It is a cost-versus-success dial, not a free win.




