Collect structured judgments on recorded traces and preserve the reasoning behind them.
Use Human Scores to define criteria, then create a Review containing the traces and criteria to evaluate. A criterion can be numeric, categorical, multi-select, or text-based.
Use the whitespace failure from your first evaluation, or two traces from your own application with known outcomes.
{"text":" hello "}, an output of {"text":" HELLO "} should receive 0 with a comment explaining the extra spaces.A useful criterion describes a decision a reviewer can make from the captured evidence. Split independent concerns into separate criteria. For model judgments, include known passing and failing examples in the reviewer instructions.
Open the review, inspect a trace's input, output, and execution evidence, and answer the attached criteria. Add a comment when the rating needs an explanation. The review player saves changes and supports moving between items.
An item is complete when the required criteria have answers. A session completes when every item has a complete review. Optional notes can preserve context without completing an unanswered rating.
Numeric Human Scores have normalized numeric values. Text and categorical responses remain typed judgments; they are not automatically averaged as numeric scores.
Browser feedback is human-attributed. API-key and OAuth submissions are AI-labelled, with the authenticated principal and optional agent/model metadata. AI completion and human completion have separate counts. A completed AI review is not human-verified ground truth. Notes-only updates preserve scores and completion; review operations never update dataset expectations.
Organization API keys with reviews:write can submit notes, scores and annotations. Reading reviews requires reviews:read; inspecting evidence and creating a session also requires traces:read. Legacy ingestion keys cannot review.
Use datool agent tools record_review to inspect the deployed schema, then read a session and item before submitting with its current revision:
datool reviews get 1
datool reviews item 1 --item-id item-id
datool reviews record 1 --item-id item-id --input @finding.json
datool reviews export 1 --out review.ndjsonMCP exposes the same operations. API clients cannot choose their reviewer identity or claim a human source. Exported items preserve explicit provenance. See CLI reference and authentication.
A review preserves a judgment; it does not silently rewrite a dataset. After checking a human review, explicitly curate the reference answer in a dataset, preserve the original trace as evidence, and use a snapshot for the next evaluation.
Keep human and AI-labelled counts separate when assessing review coverage. Exported review items include criterion definitions and provenance so another reader can distinguish who made the judgment and what they were asked to decide.