# Datasets

Turn representative inputs and known failures into repeatable evaluation cases.



Open **Datasets** to create and organize cases. A useful case includes an **input**, an **expected output** when you have one, and **metadata** that explains its purpose or category. Cases can come from recorded traces or be written directly.

## Choose meaningful cases [#choose-meaningful-cases]

Start with a few common requests and known failures. Keep expected outputs aligned with your scorer: an exact-match scorer needs an exact target, while a rubric-based scorer may use requirements or reference facts.

Review input/output formatting in the editor before starting a run. Dataset values may be structured JSON, text, or messages, depending on the connected app's contract.

## Import and export [#import-and-export]

The CLI supports portable dataset files:

```json
{
  "format": 1,
  "kind": "dataset",
  "key": "uppercase-examples",
  "description": "Basic uppercase behavior",
  "items": [
    {
      "key": "greeting",
      "input": { "text": "hello" },
      "expectedOutput": { "text": "HELLO" },
      "metadata": { "category": "greeting" }
    }
  ]
}
```

Save this as `cases.json`, then use an authenticated [CLI](/docs/reference/cli):

```sh
npx datool datasets push cases.json --dry-run
npx datool datasets push cases.json
npx datool datasets pull uppercase-examples --out exported-cases.json
```

Keep stable case keys when editing a file. Dataset push retains omitted cases. A sync sidecar tracks revisions; stale changes conflict instead of silently overwriting a newer revision. Current resource imports require both `datasets:write` and `scorers:write`.

## Preserve evaluation evidence [#preserve-evaluation-evidence]

Evaluation runs save case snapshots. Editing a dataset later does not rewrite the input or expected output used by a completed run. The CLI can also create an explicit immutable dataset snapshot:

```sh
npx datool datasets snapshot dataset-id --label release-1
```

Next, [define a scorer](/docs/evaluation/scorers) and [run the dataset](/docs/evaluation/runs).

## Promote a production span [#promote-a-production-span]

In the span inspector, choose **Create dataset case**, select a dataset, preview the case, and save. The case retains source trace/span IDs, timestamps, model/prompt attributes and the selected invocation subtree. Sibling invocations and scorer executions are excluded. Observed output stays separate from expected output; the reference defaults to null. Copy it only after deliberate review.

If captured messages differ from the app's variable schema, enable **Map input to app variables** and supply explicit JSON. The captured input is preserved. Evidence scoring sees the original invocation; a connected run uses the mapped input. Datool does not infer variables from prompt text.

Agents use `promote_spans` or `datool datasets promote dataset-id --input @promotion.json`. Preview defaults to true. Save with `preview: false` and the returned `expectedEvidenceHash`. Batches contain 1–100 spans and are atomic. A selected capture is limited to 1 MiB and 1,000 spans; the batch limit is 8 MiB.

For large items, `get_dataset` accepts `includeItems: false` to read only the dataset header and count. `list_dataset_items` accepts `preview: true`: fields over 16 KiB are replaced by placeholders with explicit `omittedFields` entries containing their byte size and a 128-character preview. The dataset editor keeps these fields behind a blurred cover until you click **Load field**. Each click fetches only that field; small fields remain editable without loading the others. Saves send only changed fields and preserve unloaded values.

Use `get_dataset_item` with `fields: ["input"]` to load input while keeping other large fields as previews. An empty `fields` array requests previews only; omit `fields` to retrieve the complete item for export. The REST item GET and PATCH endpoints accept the equivalent comma-separated `fields` query parameter to control response fields. Never write an omitted field's placeholder back as its value.

Complete single-item reads are bounded to 32 MiB. Dataset item creation, updates and bulk requests accept up to 4 MiB through REST and MCP; other request limits and the 8 MiB evaluation evidence budget remain unchanged. A reverse proxy may enforce a lower upload limit. Preview pages never replace or truncate stored evidence.

Span-backed snapshots preserve captured evidence. Re-score an evaluation with `sourceRunId` after judge changes. Execute the connected app again after changing its extractor, prompt or model. Keep calibration cases separate from untouched evaluation cases.

