Evals that keep your AI analyst accountable.
Verify that your AI analyst stays accurate
as your data models, definitions, and analytical context evolve.

Test answers against the
source of truth
For metric questions, evaluate answers against deterministic goldens generated from your semantic layer and warehouse. Every check stays tied to the definitions your business already trusts.

Turn trusted conversations
into coverage
Promote a real, multi-turn analyst conversation into an Eval. Sundial replays the original user turns as context and verifies the final answer, so a trusted workflow becomes lasting regression coverage.

Ship model changes
with evidence
The Data Modeling Agent finds gaps in coverage, generates and verifies targeted Evals, then runs them in a branch. Your team can see the impact of a model change before it reaches more users.

Evaluate deeper analysis
with the right rubric
Not every business question has one numeric answer. For narrative and complex analysis, evaluate the substance of the response against a clear rubric, with repeated runs that reduce judge variability.

A tighter loop from real usage to better answers
Keep improving the system behind every answer without losing the behaviors your business already trusts.
- 1
See real usage
Find the questions and answers worth improving.
- 2
Improve the system
Refine the model, definitions, or analytical context.
- 3
Verify the change
Run Evals before it reaches more users.
Questions about Evals
What can Evals test?
Evals cover everything from metric lookups and breakdowns to multi-step analysis. Use them for the questions, definitions, and decision workflows your business relies on most.
How are metric answers evaluated?
For reproducible questions, Sundial can evaluate an answer against a deterministic golden generated from the semantic layer and warehouse. That keeps the test grounded in the same governed definitions the analyst uses.
Can we evaluate a real conversation?
Yes. A trusted multi-turn conversation can become an Eval that replays the original user turns and checks the final answer, preserving valuable real-world coverage as the system changes.
When should we run Evals?
Run Evals on a schedule and before changes to models, semantic definitions, Playbooks, or other agent context. Compare runs to spot improvements and regressions before they reach more users.
Keep every answer ready for the next change.
Give your data team the proof that Sundial's AI analyst is improving without losing the answers your business depends on.