All insights

Build an Evaluation Harness for LLM Features

October 7, 2026 · 7 min read

LLMQAAI EngineeringTesting
Build an Evaluation Harness for LLM Features

Traditional software testing asks whether a system produced the expected result. LLM testing asks a harder question: whether a variable result is acceptable for a specific user, task, and risk level.

That distinction breaks many familiar QA habits. Exact string assertions reject valid answers. A few polished demos hide long-tail failures. Average quality scores conceal severe mistakes. Even a stable prompt can behave differently after a model update, retrieval change, or provider-side modification.

The solution is not to abandon deterministic testing. Keep it for deterministic components. Add an evaluation harness for the probabilistic layer: a versioned system that runs representative cases, applies explicit grading rules, records evidence, and compares changes.

Define quality as observable behavior

“Good response” is not a testable requirement. Start by decomposing quality into behaviors that can be observed independently.

For a customer-support assistant, those behaviors might include:

  • Answers the user’s actual question.
  • Uses only supported policy claims.
  • Cites the relevant source when required.
  • Refuses requests outside its permitted scope.
  • Avoids exposing personal or confidential information.
  • Escalates when evidence is missing or conflicting.
  • Produces the required structure for downstream systems.

Not every dimension deserves equal weight. A slightly verbose answer is inconvenient; an invented refund policy creates liability. Classify dimensions as critical, important, or advisory before choosing metrics.

Also define the unit under evaluation. It may be a single response, a multi-turn conversation, a tool call, a retrieved context set, or an end-to-end task. Testing only final prose can miss the actual defect. A convincing answer may rest on poor retrieval, an invalid tool argument, or unsupported reasoning.

Build the evaluation set from real work

A useful evaluation set is not a random collection of prompts written by the engineering team. It should approximate production demand and deliberately expose known risks.

Start with anonymized production examples when available. Cluster them by intent, complexity, language, customer segment, and required capability. Then add synthetic cases to cover rare but consequential conditions.

Maintain at least four categories:

  • Typical cases: Common requests representing normal traffic.
  • Boundary cases: Ambiguous wording, long context, conflicting documents, and incomplete information.
  • Adversarial cases: Prompt injection, data-extraction attempts, unsafe instructions, and attempts to bypass policy.
  • Regression cases: Every important defect discovered in testing or production.

Attach metadata to each case. Record its category, risk level, expected behavior, permitted sources, grading rubric, and owner. For tasks with one correct result, include a reference answer. For open-ended tasks, specify required facts and prohibited claims instead of forcing one canonical phrasing.

Do not let the set become static. Production traffic changes, policies change, and users learn new ways to interact with the feature. Review the distribution regularly and retire obsolete cases without deleting their history.

Use layered grading instead of one score

No single evaluator is reliable enough for every failure mode. Combine deterministic checks, model-based grading, and human review.

Deterministic checks should handle anything that can be expressed as code. Examples include schema validity, required fields, citation existence, allowed tool names, latency limits, token budgets, and prohibited patterns. These checks are fast, inexpensive, and explainable.

Model-based graders are useful for semantic qualities such as relevance, completeness, groundedness, and tone. Give the grader a narrow rubric and require structured output with a score, rationale, and quoted evidence. Avoid vague prompts such as “Rate this response from 1 to 10.”

Use pairwise comparison when choosing between two variants. Evaluators are often more consistent at deciding which response better satisfies a rubric than assigning absolute scores. Randomize response order and periodically test for position bias.

Human review remains necessary for high-risk or disputed cases. Reviewers should see the same rubric, not rely on intuition. Measure agreement between reviewers and between human and model grades. Low agreement usually indicates an ambiguous criterion, insufficient context, or a genuinely subjective requirement.

Calibrate model graders against a human-labeled sample before trusting them. Recheck calibration after changing the grader model, rubric, or task distribution. The evaluator is part of the test system and must itself be versioned.

Test the pipeline, not just the final answer

Most LLM features are pipelines. A typical retrieval-based assistant may include query rewriting, access filtering, retrieval, reranking, prompt assembly, generation, citation mapping, and output validation. An end-to-end score tells you whether the pipeline failed but not where.

Add component-level evaluations alongside end-to-end cases:

  • Retrieval recall for documents required to answer the question.
  • Ranking quality for the first few context items.
  • Context precision to detect irrelevant or conflicting material.
  • Tool-selection accuracy and argument validity.
  • Groundedness between claims and supplied evidence.
  • Citation entailment, not merely citation presence.
  • Conversation-state retention across multiple turns.

Preserve traces for failed runs: prompts, retrieved document identifiers, tool inputs and outputs, model parameters, response, grader result, latency, and cost. Redact sensitive data, but retain enough evidence to reproduce the defect.

This separation improves triage. If the correct policy never reached the model, prompt tuning is unlikely to help. If the correct source was present but the response contradicted it, generation or instruction-following is the more likely problem.

Control variance during testing

LLM outputs vary, so one run per case can produce false confidence. Run critical cases multiple times and report pass rates rather than only pass or fail. The number of repetitions should reflect impact and expected variance; routine cases may need a few runs, while safety-critical behavior deserves more.

Pin every controllable dependency: model version, system prompt, temperature, retrieval index, tool definitions, evaluator version, and dataset revision. Store these values with results.

Separate two questions:

  1. Did the proposed change improve behavior under controlled conditions?
  2. Is behavior stable enough under realistic sampling conditions?

For the first, reduce randomness to make variants comparable. For the second, test the production configuration and examine the distribution of outcomes. Report confidence intervals for aggregate metrics when sample sizes permit.

Turn failures into an engineering workflow

An evaluation harness creates value only if failures lead to action. Each failed case should be classifiable by likely cause: retrieval, instruction conflict, missing knowledge, tool misuse, output formatting, policy gap, evaluator error, or model limitation.

Assign owners by component rather than sending every failure to the “AI team.” Search engineers should own retrieval defects. Product and legal stakeholders should resolve ambiguous policy. Application engineers should own broken validation and tool contracts.

When a production incident occurs, minimize it into a reusable regression case. Keep the original trace separately, then create the smallest test that reproduces the behavior without customer data. This turns operational learning into durable coverage.

Track trends by slice, not just overall averages. A 2% aggregate improvement can hide a serious decline for Spanish queries, long documents, or a regulated workflow. Publish the worst-performing important slices and their sample counts.

Keep offline evaluation connected to users

Offline scores are proxies. They cannot prove that the feature helps users complete work.

Validate promising variants through staged production experiments. Depending on the feature, useful outcomes may include successful task completion, correction rate, escalation rate, time to resolution, accepted suggestions, or user abandonment. Capture explicit negative feedback, but do not treat thumbs-up rates as a complete quality measure; feedback is sparse and selection-biased.

Compare offline metrics with production outcomes. If a metric improves while users perform worse, the rubric or dataset is misaligned. Revise the evaluation system rather than optimizing harder against the wrong target.

Takeaway

LLM QA becomes manageable when quality is decomposed into observable behaviors, tested on representative cases, and graded with multiple methods. Version the dataset and evaluators, test pipeline components, measure variance, and convert every meaningful defect into regression coverage. The goal is not perfectly repeatable language. It is repeatable evidence that the feature behaves acceptably for its users and risks.