Video summary

Research Insights Made Simple #26: как строить работающие evals для AI-агентов

Main summary

Key takeaways

Educational

Main ideas, concepts, and lessons

  • Evals shouldn’t just test “did it pass?”—they should evaluate the whole AI system and its business/production behavior.

    • Instead of evaluating only a model (e.g., on benchmark-style success/failure), evaluate:
      • the task outcome
      • the result
      • the execution trace (how the agent/system arrived there)
      • operations questions like stability and cost when integrated into a process
  • Work with “reproducible scenarios” and a “gate” for rollout.

    • Similar to engineering testing:
      • define a starting state
      • define a contract (what the agent can do)
      • run the agent on reproducible scenarios
      • use judgement/grading to decide whether to roll out or not
  • Use an evaluation loop that mirrors production reality (offline ↔ online loops).

    • Offline loop (pre-rollout):
      • collect user-like traces/corner cases
      • improve prompts, graders, and evaluation datasets
    • Online loop (post-rollout):
      • in production, collect traces from real users
      • annotate corner cases
      • expand the evaluation dataset
  • LM-as-judge and deterministic checks are complementary.

    • Some evaluations can be deterministic and fast (regex/blacklists, code-level checks, mathematical checks).
    • Others require LM judgment (quality/rubric-based evaluation).
    • LM-as-judge may require calibration (e.g., using reasoning, comparing to reference preferences, or asking other models).
  • Eval design should support non-determinism and agent behavior.

    • Agents are not just single outputs—they involve:
      • planning
      • tool usage
      • multiple steps/turns
      • context management
      • trajectories that can go off-track
    • Therefore, evals should measure trajectory features, not only final answers (e.g., number of steps/turns, tool calls, whether the agent stayed within guardrails).
  • Different “types of evals” for business/product settings.

    • The speaker frames evals into multiple categories (explained in the talk), emphasizing how they map to business needs.

Methodology / framework presented (detailed bullet points)

1) Build evals as part of product logic (not just a benchmark)

  • Treat evals/prompts/graders as business-logic components that must be tested like any production system.
  • Don’t assume third-party accuracy will remain stable—models and toolchains change frequently.

2) Use a “two-circuit” evaluation approach

  • Circuit A: offline (development + rollout preparation)

    • prepare/improve:
      • prompts
      • context assembly strategy
      • evaluation datasets
      • graders/judges
    • use reproducible scenarios and run them repeatedly before deployment
  • Circuit B: online (production feedback loop)

    • real users generate traces
    • you splice/merge/manual-markup corner cases
    • expand the dataset and refine evals (and prompts/graders) over time

3) Choose the appropriate grader/eval type (four categories described)

  • Type 1: Deterministic fast checks

    • binary pass/fail where possible
    • implemented via business rules:
      • rules for prohibited content/PII
      • forbidden words via regex/blacklists
      • code-level or math-based checks
  • Type 2: “LM as a judge”

    • use an LLM to score outcomes
    • can be binary or graded/ranged
    • should include/encourage reasoning or justification (with care, since reflection/explanations can be gamed)
    • may require hacks like:
      • asking other models for consistency checks
      • improving prompts to get better judge behavior
  • Type 3: Subject-matter expert evaluation (“subject meta-experts”)

    • human experts prepare/label datasets:
      • build representative scenarios
      • do markup on examples (confirm what’s wrong, what’s missing)
    • can start small internally (even ~100 cases as a starter dataset)
    • scaling considerations:
      • internal trusted annotators first
      • external expert platforms may be needed for larger rollouts or higher risk
  • Type 4: User-facing/feedback-loop signals (“online metrics” and product analytics)

    • incorporate user feedback:
      • thumbs up/down, ratings
      • behavioral traces from the product (click paths, user actions)
    • use analytics baselines and compare variants (especially in B2C/internal products with high feedback volume)

4) Build evals using “task → graders → metrics → outcome/trajectory”

A system-level view was described (inspired by an external “anatomy” of eval systems):

  • Task definition

    • includes expected input/output
    • works for both prompt-based and agent-based flows
  • Graders

    • combine deterministic checks + LM judging
  • Metrics

    • not always obvious; derive them from the task and operational constraints
    • examples discussed:
      • latency/cost impact (long checks affect performance)
      • PI/secret leakage detection via in-request checks
      • success path efficiency (how many steps were needed)
  • Outcome / trajectory measurement

    • for agents:
      • measure how the agent moved toward the goal
      • count steps/turns/tool calls
      • detect early stops vs excessive wandering
      • correlate mistakes with trace segments

5) Make evaluation instrumentation part of the system (tracing/telemetry)

  • Collect execution traces and product analytics together:
    • technical traces (e.g., via OpenTelemetry/DataBricks-like tooling)
    • user journey/product interactions (what the user clicked, which buttons, etc.)
  • Use these traces to:
    • build evaluation sets
    • replay realistic behavior in offline evals
    • treat real interactions as positive/negative examples
    • detect failure modes that pure final-answer scoring misses

6) Use evals to counter model/tool “cheating” and regressions

  • Models can exploit eval loopholes (e.g., creating dummy files or placeholders that satisfy surface checks).
  • Evals must evolve with:
    • new models (behavior changes)
    • new tool providers / pricing / availability
    • prompt changes becoming insufficient for newer models
  • Without evals, teams fall into:
    • “trusting the demo” or “it’s the vendor—so it’s fine” behavior

7) Roll out progressively using the maturity pyramid idea (L0 → higher levels)

  • Start with prototypes / one-shot prompts + small datasets:
    • “demo-level” correctness
    • hackathons are mentioned as an easy starting point
  • Then integrate into loops and production traffic:
    • use telemetry + feedback to improve prompts/graders
  • Over time:
    • prompts become business logic
    • you aim for closed-loop improvements (offline + online)

8) “Skill” / tool evaluation as a control mechanism (measurability)

  • Many agent “skills”/tools should provide their own evaluation signals.
  • Measure whether turning a skill on/off affects task quality.
  • Potentially deactivate skills if they are harmful or unnecessary.
  • Guidance summary: testing/measurability pushes closer to standard engineering/testing maturity.

Key examples mentioned (how evals apply)

  • Localization automation

    • correct behavior across languages and models
    • evaluation can include expert validation and feedback loops
  • Virtual assistant / regulated domains

    • some sections should not be edited (risk of breaking rules/constraints)
  • Agent trajectories

    • measure step counts and tool-call behavior (too many turns implies issues)
  • PII leakage / policy checks

    • run checks during request/response cycles to detect unwanted data
    • tie metrics to operational cost/latency
  • Code-related examples

    • discussion of “cheating” in benchmarks and the need for robust checks

Speakers / sources featured (identified)

Speakers

  • Sasha (host; name not fully captured in subtitles)
  • Zhenya (Evgeny/engineering director) — Engineering Director at Flo (as stated), discussing:
    • eval development
    • product analytics/tracing approach

Organizations / tools / sources mentioned

  • Flo
  • MLflow
  • DataBricks (tracing integration concept)
  • OpenTelemetry (telemetry/tracing standards)
  • Toloka / Mechanical Turk (examples of crowdsourcing/external labeling platforms)
  • TSSR (tool mentioned for generating datasets)
  • .gitignore (mentioned in the context of cheating/benchmarks; “Git ignore” implied)
  • Harness (mentioned as a benchmark/evaluation platform)
  • AI providers/models mentioned (in passing):
    • Anthropic (Claude, Opus)
    • Gemini
    • OpenAI
    • Hugging Face
    • MCP (mentioned in a tooling compatibility context)
  • Meta (mentioned in context of a paper/benchmark discussion; exact paper title not given)
  • A third-party referenced “anatomy of systems / eval building” framework (author not specified in subtitles)

Original video