what i learned about llm evals and why they matter
september 28, 2026
one thing that clicked for me while learning about llm applications is that getting a good response from a model is not the same as building a system that works reliably. you can change a prompt, switch models, modify a tool, change retrieval, or update the agent harness, and something that worked yesterday starts failing and you might not notice immediately. this is where evals started making sense to me.

so, what is an llm eval?
an eval is basically a controlled test for an ai system. the simplest way to think about one is: given this input, the system should do x, not y. suppose an agent handles billing questions you might want to test that it looks up the user's account before giving billing-specific information, or that it must never invent a refund policy.
an eval usually has a few basic pieces:
- task (the input and what success means)
- trial (one attempt at that task)
- grader (code, an llm, or a human that decides how it performed)
- transcript (what happened during the run)
- outcome (the resulting state)
this matters because an application eval isn't asking whether a model is generally capable it's asking whether this particular system behaves correctly for the things we care about.
why bother with them?
the most useful thing about evals, at least to me, is regression prevention. imagine an agent correctly follows the rule "always look up the account before answering a billing question." you later change the system prompt it works better for twenty other cases, but accidentally causes the agent to skip account lookup in some billing conversations. without an eval, that old failure can quietly come back.
this is exactly what happened to me the first time it actually clicked. i tweaked a system prompt to fix a tone issue the agent was being too apologetic and the change worked great. except three days later i noticed, buried in a batch of transcripts, that the same agent had stopped looking up the account before answering billing questions. nothing had flagged it. no error, no crash, just a slightly-too-confident answer built on stale context. i only caught it because i happened to be reading transcripts that week. that's the moment i stopped thinking of evals as "nice to have documentation" and started thinking of them as the thing that would have caught this automatically, the day it happened.
a useful eval isn't just something that's easy to score it should have a clear target behavior, reproducible input, understandable grading, and a connection to something the product actually needs.
but an eval set is still a model of reality
an eval dataset is controlled; real users aren't. your test might contain "how do i get a refund?" while a real user says something like "hey i was charged again and i thought i cancelled this last week, can you reverse it?" or makes a typo, changes their goal mid-conversation, or asks for something the team never thought to test.
your eval can tell you whether the system behaved correctly on a known case. it can't tell you what unexpected problems users are having this week. evals catch the known knowns, and production behavior is where unknown failures start showing up. that doesn't make evals less useful it just gives them a boundary.
evals, traces, and conversations answer different questions
it's useful to separate three things that are often discussed together.
evals ask did the known behavior work. they're controlled and repeatable, useful for regression checks and release gates.
traces show what actually happened inside the system model calls, retrieval, tool calls, retries, routing decisions, latency, errors. if an eval tells you the billing test failed, the trace is what tells you the agent skipped account lookup and answered from stale context.
conversations show what happened from the user's perspective. a user repeatedly correcting the agent, rephrasing the same request, or abandoning a task can reveal problems that aren't obvious from a test case or an individual trace.
these three layers fit together: the conversation shows what went wrong for the user, the trace shows how the system got there, and the eval turns that failure into something repeatable so it can be caught again later.
a good response can still be the wrong answer
another distinction i found useful was response quality vs. resolution. an llm can produce a response that's grammatical, confident, and well formatted while still failing to accomplish what the user actually wanted. take "can i use this with salesforce?" a reply like "we support many crm platforms" isn't false, but it doesn't answer the actual question. the same gap shows up with questions about billing, security, exports, troubleshooting, and setup. that changed how i think about evaluating conversational systems instead of only asking "was the response good," the more useful question is often "did the user get what they came here to accomplish."
where intent resolution fits
intent resolution asks whether the user's stated or implied goal was actually addressed. response ≠ resolution, and there isn't one universal definition of "resolved": for a support agent it might mean the user got the answer without needing human help, for a coding assistant it might be whether the suggested code was actually used, and for a tutor it could mean the user moved on to the next concept instead of repeating the same foundational question.
you can't directly observe intent, so you infer it from behavior rephrasing a request, escalating to a human, or abandoning a conversation can all signal non-resolution, while a follow-up that builds on the previous answer can signal progress. none of these signals is perfect alone, which is why the useful part is defining what "resolved" means for the product and measuring toward that.
the interesting part happens after production
this is probably the biggest thing i took away from the whole topic: the eval isn't the end of the process. say users keep correcting the agent on billing questions you inspect the conversations, find the pattern, then check the traces and discover the agent isn't performing the account lookup. now there's an actual engineering problem to fix, whether that's a prompt change, retrieval, the harness, or fallback behavior. the important part is not stopping at "the response quality is bad" you change the system actually causing it, then measure again. if the pattern matters, that production case turns into an eval, so the next regression gets caught automatically.

where evals fit
evals aren't proof that an ai application works they're a way of making known important behavior measurable and repeatable. production conversations help you discover failures you didn't know to test for, traces help you understand why those failures happened, and the useful failures among them become evals. that's also why generic model benchmarks aren't enough for an application: a benchmark says something about general model capability, but your product depends on the whole pipeline model, prompts, retrieval, tools, memory, routing, and harness behavior.
the goal isn't an eval suite that says everything is perfect. it's making the behaviors that matter visible, catching regressions before they quietly return, and turning real production failures into better tests.