ACAgentic Craft

FoundationHazelJS

Evaluation (evals)

Offline and online methods to score whether agents meet task, safety, and cost goals.

Authors
editorial-team
Published
Last reviewed

Direct answer

Offline and online methods to score whether agents meet task, safety, and cost goals.

Plain language

Tests and scorecards for agents—so you don’t rely on a few lucky demos.

Technical explanation

Golden tasks, rubrics, trajectory assertions, LLM-as-judge (calibrated), production sampling, and regression gates in CI—plus AgentRuntime loop critique/validate stages for online quality loops.

Example

CI fails if forbidden tools are called on the incident-summarizer suite; online loop successScore gates drafts.

Related

← All glossary terms