Datasets
0 created in last 7 days
from persisted registry
Control plane
Monitor dataset ingestion, rubric health, model regressions, and human review throughput from one operational surface.
Datasets
0 created in last 7 days
from persisted registry
Pending runs
Pending or running
from evaluation statuses
Pass rate
No completed runs
score threshold: 70
Adjudication
Manual decisions required
from adjudication queue
Latest persisted evaluation activity across datasets.
Dispatch evaluation runs to populate live operational activity.
Page 1 of 1
Completed scored evaluations will appear here after execution runs finish.
Run evaluations with model identifiers to compare pass rates.
No evaluations currently require manual adjudication.
Live pass rates and review load grouped by model identifier.
Run evaluations with model identifiers to compare operational quality.
Page 1 of 1