Evaluations Commands
See Evaluations module docs for permissions and data model.
soat list-datasets
List datasets
- Method:
GET - Path:
/api/v1/datasets
Usage
soat list-datasets
Options
--project-id
Project ID (required if not using project key auth)
- Source:
query - Required: no
- Type:
string - Example:
proj_V1StGXR8Z5jdHi6B
--limit
Maximum number of results to return
- Source:
query - Required: no
- Type:
integer - Default:
50
--offset
Number of results to skip
- Source:
query - Required: no
- Type:
integer - Default:
0
soat create-dataset
Create a dataset
- Method:
POST - Path:
/api/v1/datasets
Usage
soat create-dataset --name <string>
Options
--project-id
Project ID (required if not using project key auth)
- Source:
body - Required: no
- Type:
string - Example:
proj_V1StGXR8Z5jdHi6B
--name
Unique name within the project
- Source:
body - Required: yes
- Type:
string - Example:
billing-regressions
--description
What this suite covers
- Source:
body - Required: no
- Type:
string \| null - Example:
Questions the billing agent regressed on in Q2
soat get-dataset
Get a dataset
- Method:
GET - Path:
/api/v1/datasets/{dataset_id}
Usage
soat get-dataset --dataset-id <string>
Options
--dataset-id
Dataset ID
- Source:
path - Required: yes
- Type:
string - Example:
dset_V1StGXR8Z5jdHi6B
soat update-dataset
Update a dataset
- Method:
PUT - Path:
/api/v1/datasets/{dataset_id}
Usage
soat update-dataset --dataset-id <string>
Options
--dataset-id
Dataset ID
- Source:
path - Required: yes
- Type:
string - Example:
dset_V1StGXR8Z5jdHi6B
--name
—
- Source:
body - Required: no
- Type:
string - Example:
billing-regressions
--description
—
- Source:
body - Required: no
- Type:
string \| null - Example:
Questions the billing agent regressed on in Q2
soat delete-dataset
Delete a dataset
- Method:
DELETE - Path:
/api/v1/datasets/{dataset_id}
Usage
soat delete-dataset --dataset-id <string>
Options
--dataset-id
Dataset ID
- Source:
path - Required: yes
- Type:
string - Example:
dset_V1StGXR8Z5jdHi6B
soat list-dataset-items
List dataset items
- Method:
GET - Path:
/api/v1/datasets/{dataset_id}/items
Usage
soat list-dataset-items --dataset-id <string>
Options
--dataset-id
Dataset ID
- Source:
path - Required: yes
- Type:
string - Example:
dset_V1StGXR8Z5jdHi6B
--limit
Maximum number of results to return
- Source:
query - Required: no
- Type:
integer - Default:
50
--offset
Number of results to skip
- Source:
query - Required: no
- Type:
integer - Default:
0
soat create-dataset-item
Add a dataset item
- Method:
POST - Path:
/api/v1/datasets/{dataset_id}/items
Usage
soat create-dataset-item --dataset-id <string> --input <array<object>>
Options
--dataset-id
Dataset ID
- Source:
path - Required: yes
- Type:
string - Example:
dset_V1StGXR8Z5jdHi6B
--input
Messages replayed verbatim as the generation's input
- Source:
body - Required: yes
- Type:
array<object>
--expected-output
Reference answer for exact_match / embedding_similarity / llm_judge scorers
- Source:
body - Required: no
- Type:
string \| null - Example:
Your invoice is issued on the first of each month.
--metadata
Free-form tags, opaque to the platform
- Source:
body - Required: no
- Type:
object<string, unknown> - Example:
\{"topic":"billing"\}
soat create-dataset-item-from-generation
Curate a dataset item from a generation
- Method:
POST - Path:
/api/v1/datasets/{dataset_id}/items/from-generation
Usage
soat create-dataset-item-from-generation --dataset-id <string> --generation-id <string>
Options
--dataset-id
Dataset ID
- Source:
path - Required: yes
- Type:
string - Example:
dset_V1StGXR8Z5jdHi6B
--generation-id
The completed generation to promote. Must belong to the same project as the dataset.
- Source:
body - Required: yes
- Type:
string - Example:
gen_V1StGXR8Z5jdHi6B
--expected-output
Reference answer. Omit to use the generation's own answer; pass null to store the item with no reference answer.
- Source:
body - Required: no
- Type:
string \| null - Example:
Your invoice is issued on the first of each month.
--metadata
Free-form tags, opaque to the platform
- Source:
body - Required: no
- Type:
object<string, unknown> - Example:
\{"topic":"billing"\}
soat update-dataset-item
Update a dataset item
- Method:
PUT - Path:
/api/v1/datasets/{dataset_id}/items/{item_id}
Usage
soat update-dataset-item --dataset-id <string> --item-id <string>
Options
--dataset-id
Dataset ID
- Source:
path - Required: yes
- Type:
string - Example:
dset_V1StGXR8Z5jdHi6B
--item-id
Dataset item ID
- Source:
path - Required: yes
- Type:
string - Example:
dsit_V1StGXR8Z5jdHi6B
--input
Messages replayed verbatim as the generation's input
- Source:
body - Required: no
- Type:
array<object>
--expected-output
—
- Source:
body - Required: no
- Type:
string \| null - Example:
Your invoice is issued on the first of each month.
--metadata
—
- Source:
body - Required: no
- Type:
object<string, unknown> - Example:
\{"topic":"billing"\}
soat delete-dataset-item
Delete a dataset item
- Method:
DELETE - Path:
/api/v1/datasets/{dataset_id}/items/{item_id}
Usage
soat delete-dataset-item --dataset-id <string> --item-id <string>
Options
--dataset-id
Dataset ID
- Source:
path - Required: yes
- Type:
string - Example:
dset_V1StGXR8Z5jdHi6B
--item-id
Dataset item ID
- Source:
path - Required: yes
- Type:
string - Example:
dsit_V1StGXR8Z5jdHi6B
soat list-evals
List evals
- Method:
GET - Path:
/api/v1/evals
Usage
soat list-evals
Options
--project-id
Project ID (required if not using project key auth)
- Source:
query - Required: no
- Type:
string - Example:
proj_V1StGXR8Z5jdHi6B
--limit
Maximum number of results to return
- Source:
query - Required: no
- Type:
integer - Default:
50
--offset
Number of results to skip
- Source:
query - Required: no
- Type:
integer - Default:
0
soat create-eval
Create an eval
- Method:
POST - Path:
/api/v1/evals
Usage
soat create-eval --name <string> --agent-id <string> --dataset-id <string> --scorers <array<object | object | object | object | object | object | object>>
Options
--project-id
Project ID (required if not using project key auth)
- Source:
body - Required: no
- Type:
string - Example:
proj_V1StGXR8Z5jdHi6B
--name
Unique name within the project
- Source:
body - Required: yes
- Type:
string - Example:
billing-regression-suite
--agent-id
The agent under test
- Source:
body - Required: yes
- Type:
string - Example:
agent_V1StGXR8Z5jdHi6B
--dataset-id
The dataset to run it against
- Source:
body - Required: yes
- Type:
string - Example:
dset_V1StGXR8Z5jdHi6B
--scorers
Scorer configs, a discriminated union on type. Each type may appear at most once. Every scorer produces \{ score: 0–1, passed: boolean \}; binary scorers emit 0 or 1. exact_match compares the trimmed output text to expected_output. contains looks for value in the output text. json_logic evaluates expression over \{ input, output, object, expected, item.metadata \}, where object is the structured output (absent when the agent has no output_schema). output_schema validates the structured output against the scorer's own schema, falling back to the agent's; it requires the agent to carry an output_schema, because without one the platform emits no structured output and every item would score 0. llm_judge grades the output with a model completion, returning a continuous score plus its reasoning. Its pass_threshold is required: a continuous score says nothing about where "good enough" is, and a defaulted cutoff would silently decide the gate. embedding_similarity embeds the output text and expected_output with the platform's configured embedding model (EMBEDDING_PROVIDER / EMBEDDING_MODEL — the same stack document ingestion uses) and scores their cosine similarity, clamped to 0-1. Its pass_threshold is required for the same reason as the judge's. An item without an expected_output scores 0; an embedding backend failure marks the item errored, never a score of 0. tool runs a custom scoring algorithm: a server-callable project tool the engine invokes once per item with the item's context. Unlike the built-in types it may appear several times, each under a distinct name — outcomes and aggregates key on the name.
- Source:
body - Required: yes
- Type:
array<object \| object \| object \| object \| object \| object \| object>
--pass-threshold
0–1. The run passes iff its pass rate — passed items over non-errored items — is at least this. Null reports scores without gating on them.
- Source:
body - Required: no
- Type:
number \| null - Example:
0.8
soat get-eval
Get an eval
- Method:
GET - Path:
/api/v1/evals/{eval_id}
Usage
soat get-eval --eval-id <string>
Options
--eval-id
Eval ID
- Source:
path - Required: yes
- Type:
string - Example:
eval_V1StGXR8Z5jdHi6B
soat update-eval
Update an eval
- Method:
PUT - Path:
/api/v1/evals/{eval_id}
Usage
soat update-eval --eval-id <string>
Options
--eval-id
Eval ID
- Source:
path - Required: yes
- Type:
string - Example:
eval_V1StGXR8Z5jdHi6B
--name
—
- Source:
body - Required: no
- Type:
string - Example:
billing-regression-suite
--agent-id
—
- Source:
body - Required: no
- Type:
string - Example:
agent_V1StGXR8Z5jdHi6B
--dataset-id
—
- Source:
body - Required: no
- Type:
string - Example:
dset_V1StGXR8Z5jdHi6B
--scorers
Scorer configs, a discriminated union on type. Each type may appear at most once. Every scorer produces \{ score: 0–1, passed: boolean \}; binary scorers emit 0 or 1. exact_match compares the trimmed output text to expected_output. contains looks for value in the output text. json_logic evaluates expression over \{ input, output, object, expected, item.metadata \}, where object is the structured output (absent when the agent has no output_schema). output_schema validates the structured output against the scorer's own schema, falling back to the agent's; it requires the agent to carry an output_schema, because without one the platform emits no structured output and every item would score 0. llm_judge grades the output with a model completion, returning a continuous score plus its reasoning. Its pass_threshold is required: a continuous score says nothing about where "good enough" is, and a defaulted cutoff would silently decide the gate. embedding_similarity embeds the output text and expected_output with the platform's configured embedding model (EMBEDDING_PROVIDER / EMBEDDING_MODEL — the same stack document ingestion uses) and scores their cosine similarity, clamped to 0-1. Its pass_threshold is required for the same reason as the judge's. An item without an expected_output scores 0; an embedding backend failure marks the item errored, never a score of 0. tool runs a custom scoring algorithm: a server-callable project tool the engine invokes once per item with the item's context. Unlike the built-in types it may appear several times, each under a distinct name — outcomes and aggregates key on the name.
- Source:
body - Required: no
- Type:
array<object \| object \| object \| object \| object \| object \| object>
--pass-threshold
—
- Source:
body - Required: no
- Type:
number \| null - Example:
0.8
soat delete-eval
Delete an eval
- Method:
DELETE - Path:
/api/v1/evals/{eval_id}
Usage
soat delete-eval --eval-id <string>
Options
--eval-id
Eval ID
- Source:
path - Required: yes
- Type:
string - Example:
eval_V1StGXR8Z5jdHi6B
soat list-eval-runs
List eval runs
- Method:
GET - Path:
/api/v1/evals/{eval_id}/runs
Usage
soat list-eval-runs --eval-id <string>
Options
--eval-id
Eval ID
- Source:
path - Required: yes
- Type:
string - Example:
eval_V1StGXR8Z5jdHi6B
--limit
Maximum number of results to return
- Source:
query - Required: no
- Type:
integer - Default:
50
--offset
Number of results to skip
- Source:
query - Required: no
- Type:
integer - Default:
0
soat start-eval-run
Start an eval run
- Method:
POST - Path:
/api/v1/evals/{eval_id}/runs
Usage
soat start-eval-run --eval-id <string>
Options
--eval-id
Eval ID
- Source:
path - Required: yes
- Type:
string - Example:
eval_V1StGXR8Z5jdHi6B
--wait
True runs the eval synchronously (25-item cap) and returns a terminal run with its scores. False — the default — enqueues the items and returns a queued run immediately.
- Source:
body - Required: no
- Type:
boolean - Default:
false - Example:
true
--agent-version
An archived agent version to evaluate. Defaults to the active release's stable version, or the live draft version when no release is in effect.
- Source:
body - Required: no
- Type:
integer \| null - Example:
3
--baseline-run-id
A terminal run of the same eval to compare against. The finished run's aggregate_scores.baseline reports per-scorer deltas over the item intersection. A run of a different eval is rejected with 400.
- Source:
body - Required: no
- Type:
string \| null - Example:
evrun_V1StGXR8Z5jdHi6B
--metadata
Caller-supplied key/value metadata attached to the run record for attribution — what this measurement was of (the commit or release candidate being scored, the CI job that asked for it). Round-trips verbatim on every read of the run, the list included. The bag is caller-owned and no key is reserved: everything the platform decides about a run (status, agent_version, baseline_run_id, aggregate_scores, passed, the counts) is a field of its own and cannot be written from here. Nothing in the scoring path reads it. A non-object is rejected with 400 VALIDATION_FAILED and no run is created.
- Source:
body - Required: no
- Type:
object<string, unknown> - Example:
\{"commit_sha":"9f2c1ab","ci_job":"nightly-evals"\}
--tool-context
Key/value context forwarded to every item's generation, so an agent whose tools authorize through tool_context is scored against the configuration it runs in production rather than with an empty bag. Each key is forwarded as one X-Soat-Context-<key> header and resolves any \{\{context:<key>\}\} token in a bound tool's headers or preset_parameters. Stored on the run and re-read per item, since a queued run (the default) is driven by a worker with no request behind it. Write-only: no read of the run returns it, unlike metadata — a run is a report other people read, and a credential in it is not theirs to see. Cleared once the run reaches a terminal state. An eval generation has no session, so the reserved keys session_id, actor_id and actor_external_id are dropped (in any casing) rather than forwarded. Every other key becomes an HTTP header name and must match that grammar, or the request is rejected with 400 INVALID_TOOL_CONTEXT_KEY and no run is created.
- Source:
body - Required: no
- Type:
object<string, string> - Example:
\{"ocaToken":"eyJhbGciOiJIUzI1NiJ9.abc","tenant":"acme"\}
soat get-eval-run
Get an eval run
- Method:
GET - Path:
/api/v1/evals/{eval_id}/runs/{eval_run_id}
Usage
soat get-eval-run --eval-id <string> --eval-run-id <string>
Options
--eval-id
Eval ID
- Source:
path - Required: yes
- Type:
string - Example:
eval_V1StGXR8Z5jdHi6B
--eval-run-id
Eval run ID
- Source:
path - Required: yes
- Type:
string - Example:
evrun_V1StGXR8Z5jdHi6B
soat list-eval-results
List eval run results
- Method:
GET - Path:
/api/v1/evals/{eval_id}/runs/{eval_run_id}/results
Usage
soat list-eval-results --eval-id <string> --eval-run-id <string>
Options
--eval-id
Eval ID
- Source:
path - Required: yes
- Type:
string - Example:
eval_V1StGXR8Z5jdHi6B
--eval-run-id
Eval run ID
- Source:
path - Required: yes
- Type:
string - Example:
evrun_V1StGXR8Z5jdHi6B
--limit
Maximum number of results to return
- Source:
query - Required: no
- Type:
integer - Default:
50
--offset
Number of results to skip
- Source:
query - Required: no
- Type:
integer - Default:
0
soat cancel-eval-run
Cancel an eval run
- Method:
POST - Path:
/api/v1/evals/{eval_id}/runs/{eval_run_id}/cancel
Usage
soat cancel-eval-run --eval-id <string> --eval-run-id <string>
Options
--eval-id
Eval ID
- Source:
path - Required: yes
- Type:
string - Example:
eval_V1StGXR8Z5jdHi6B
--eval-run-id
Eval run ID
- Source:
path - Required: yes
- Type:
string - Example:
evrun_V1StGXR8Z5jdHi6B