FoundationHazelJS
Evaluation (evals)
Offline and online methods to score whether agents meet task, safety, and cost goals.
- Authors
- editorial-team
- Published
- Last reviewed
Direct answer
Offline and online methods to score whether agents meet task, safety, and cost goals.
Plain language
Tests and scorecards for agents—so you don’t rely on a few lucky demos.
Technical explanation
Golden tasks, rubrics, trajectory assertions, LLM-as-judge (calibrated), production sampling, and regression gates in CI—plus AgentRuntime loop critique/validate stages for online quality loops.
Example
CI fails if forbidden tools are called on the incident-summarizer suite; online loop successScore gates drafts.