Create Evaluation

View as Markdown

Path parameters

workspacestringRequired

Request

This endpoint expects an object.
namestringRequired

Producer-supplied, workspace-unique evaluation id.

dataset_namestringRequired

Producer-supplied dataset name.

experiment_idslist of stringsOptional

Entity ids of the Experiments this Evaluation belongs to (>=1). Preferred; each experiment must already exist. When omitted, the deprecated experiment_group_id is used instead.

dataset_versionstringOptional

Producer-supplied dataset version.

metadatamap from strings to stringsOptional

Free-form producer metadata.

descriptionstringOptional

Human-readable description.

parent_evaluation_idstringOptional

Entity id of the evaluation this one was derived from (e.g. a variant of a baseline), if any.

statusstringOptional

Producer-defined lifecycle status of the evaluation.

root_causestringOptional

Human- or agent-authored explanation of the evaluation’s outcome (e.g. why it was killed).

experiment_group_idstringOptionalDeprecated

Deprecated single-experiment field; provide experiment_ids instead. Coalesced into experiment_ids when experiment_ids is omitted.

Response

Successful Response
idstring
namestring
workspacestring
experiment_idslist of strings

Entity ids of the Experiments this Evaluation belongs to (>=1).

dataset_namestring
experiment_group_idstringRead-onlyDeprecated

Deprecated single-experiment alias; the first of experiment_ids. Use experiment_ids.

dataset_versionstringOptional
metadatamap from strings to stringsOptional
descriptionstringOptional
parent_evaluation_idstringOptional
statusstringOptional
root_causestringOptional
created_atdatetimeOptional
updated_atdatetimeOptional
pinned_atdatetime or nullOptional

Timestamp at which the evaluation was pinned, or null if unpinned. Managed via POST/DELETE /evaluations/{name}/pin.

evaluator_nameslist of stringsOptional
model_nameslist of stringsOptional
Distinct model names observed across ingested sessions for this evaluation.
agent_nameslist of stringsOptional
Distinct agent names observed across ingested sessions for this evaluation.
agent_versionslist of stringsOptional
Distinct agent versions observed across ingested sessions for this evaluation.
aggregate_scoresmap from strings to objectsOptional
run_countintegerOptionalDefaults to 0

Number of distinct ingested evaluation sessions; one session is treated as one run.

test_case_countintegerOptionalDefaults to 0

Number of distinct test cases in the evaluation, i.e. distinct test_case_name values (sessions with no test_case_name each count as their own). A test case run k times counts once; the rollup metrics are averaged per test case before pooling across test cases.

cost_usdobjectOptional

Aggregate statistics over evaluator scores or session-level metric values.

latency_msobjectOptional

Aggregate statistics over evaluator scores or session-level metric values.

tokensobjectOptional

Average total tokens (input + output) per test case, aggregated across the evaluation.

Errors

409
Conflict Error
422
Unprocessable Entity Error