> This page is for version 0.5.1.
> For other versions, use one of these documentation indexes:
> - Latest (default): https://docs.nvidia.com/nemo-helix/latest/llms.txt
> - 0.7.0: https://docs.nvidia.com/nemo-helix/v0.7.0/llms.txt
> - 0.6.0: https://docs.nvidia.com/nemo-helix/v0.6.0/llms.txt
> - 0.5.1: https://docs.nvidia.com/nemo-helix/v0.5.1/llms.txt
> - 0.5.0: https://docs.nvidia.com/nemo-helix/v0.5.0/llms.txt
> - 0.4.0: https://docs.nvidia.com/nemo-helix/v0.4.0/llms.txt
> - 0.3.0: https://docs.nvidia.com/nemo-helix/v0.3.0/llms.txt
> - 0.2.0: https://docs.nvidia.com/nemo-helix/v0.2.0/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo-helix/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo-helix/_mcp/server.

# List Evaluations

GET /apis/intake/v2/workspaces/{workspace}/evaluations

Reference: https://docs.nvidia.com/nemo-helix/v0.5.1/documentation/reference/api-reference/evaluations/list-evaluations-apis-intake-v-2-workspaces-workspace-evaluations-get

## Request

### Path parameters

- `workspace` (string, required)

### Query parameters

- `page` (integer, optional, default: 1) — Page number.
- `page_size` (integer, optional, default: 100) — Page size.
- `sort` (string, optional) — Comma-separated list of fields to sort by, applied in order (the first field dominates); prefix any field with '-' for descending — e.g. '-evaluators.reward.mean,cost\_usd.mean'. Each field is an evaluation attribute (name, created\_at, updated\_at, pinned\_at) or an aggregate metric: run\_count, test\_case\_count, cost\_usd.\<stat>, latency\_ms.\<stat>, tokens.\<stat>, or evaluators.\<name>.\<stat>, where \<stat> is one of mean, median, p90, p95, p99, sum, count. When omitted, defaults to -created\_at with pinned evaluations first.
- `filter` (EvaluationFilter, optional) — Filter evaluations by name, experiment\_id (experiment group membership; experiment\_group\_id is a deprecated alias), dataset\_name, dataset\_version, created\_by, created\_at, or updated\_at. Pass is\_deleted=true to return only soft-deleted evaluations; omit to see only live ones. Pass is\_pinned=true (or false) to filter by pinned state; omit to return both. Filter by a metadata key/value: filter\[metadata.\<key>]=\<value>. Filter by a rollup metric with numeric range operators ($gte/$lte/$gt/$lt/$eq): filter[run_count][$gte]=5, filter\[cost\_usd.mean]\[$lte]=0.5, filter[latency_ms.p95][$lte]=1000, filter\[tokens.mean]\[$lte]=5000, or filter[evaluators.<name>.mean][$gte]=0.8.

## Response

### 200

Successful Response

- `data` (list of EvaluationResponse, required)
- `pagination` (PaginationData, optional) — Pagination information.
- `sort` (string, optional) — The field on which the results are sorted.
- `filter` (map from string to any, optional) — Filtering information.

## Errors

### 400 Bad Request Error

Unsupported sort or filter field

- `any`

### 413 Content Too Large Error

Too many evaluations selected to sort in one request

- `any`

### 422 Unprocessable Entity Error

Validation Error

- `detail` (list of ValidationError, optional)

### 503 Service Unavailable Error

Telemetry store unavailable for a metric-based sort or filter

- `any`

## Types

### EvaluationFilter

Filter for listing Evaluations.

- `name` (string, optional) — Filter evaluations by name.
- `experiment_id` (string, optional) — Filter evaluations by experiment group membership: matches evaluations whose `experiment_ids` include this group id (legacy rows still on `experiment_group_id` also match). This is the canonical membership filter — it mirrors the write-side `experiment_ids`.
- `dataset_name` (string, optional) — Filter evaluations by dataset name.
- `dataset_version` (string, optional) — Filter evaluations by dataset version.
- `created_by` (string, optional) — Filter evaluations by the principal that created them.
- `created_at` (DatetimeFilter, optional) — Filter evaluations by creation timestamp; supports `$gte` and `$lte` for ranges.
- `updated_at` (DatetimeFilter, optional) — Filter evaluations by last-updated timestamp; supports `$gte` and `$lte` for ranges.
- `is_deleted` (boolean, optional) — When true, returns only soft-deleted evaluations. Omit (or false) to see only live evaluations.
- `is_pinned` (boolean, optional) — When true, returns only pinned evaluations. When false, returns only unpinned evaluations. Omit to return both.
- `metadata` (map from string to string, optional) — Filter by a metadata key/value pair, e.g. filter[metadata.model]=claude-opus-4-8.
- `agent_name` (string, optional) — Filter evaluations that observed this agent name in any ingested session.
- `agent_version` (string, optional) — Filter evaluations that observed this agent version in any ingested session.
- `model_name` (string, optional) — Filter evaluations that observed this model name in any ingested session.
- `run_count` (NumberFilter, optional) — Filter by run count, e.g. filter[run_count][$gte]=5.
- `cost_usd` (MetricStatFilters, optional) — Filter by a cost_usd rollup stat, e.g. filter[cost_usd.mean][$lte]=0.5.
- `latency_ms` (MetricStatFilters, optional) — Filter by a latency_ms rollup stat, e.g. filter[latency_ms.p95][$lte]=1000.
- `tokens` (MetricStatFilters, optional) — Filter by a tokens rollup stat, e.g. filter[tokens.mean][$lte]=5000.
- `evaluators` (map from string to MetricStatFilters, optional) — Filter by an evaluator rollup stat, e.g. filter\[evaluators.\<name>.mean]\[\$gte]=0.8.
- `experiment_group_id` (string, optional, deprecated) — Deprecated alias for `experiment_id`; filter evaluations by experiment group membership.

### EvaluationResponse

Evaluation as served by the API, including ClickHouse-hydrated rollups.

- `id` (string, required)
- `name` (string, required)
- `workspace` (string, required)
- `experiment_ids` (list of string, required) — Entity ids of the Experiments this Evaluation belongs to (>=1).
- `dataset_name` (string, required)
- `experiment_group_id` (string, required, deprecated) — Deprecated single-experiment alias; the first of experiment_ids. Use experiment_ids.
- `dataset_version` (string, optional)
- `source_link` (string, optional)
- `metadata` (map from string to string, optional)
- `description` (string, optional)
- `parent_evaluation_id` (string, optional)
- `status` (string, optional)
- `root_cause` (string, optional)
- `created_at` (datetime, optional)
- `updated_at` (datetime, optional)
- `pinned_at` (datetime, optional, nullable) — Timestamp at which the evaluation was pinned, or null if unpinned. Managed via POST/DELETE /evaluations/\{name}/pin.
- `evaluator_names` (list of string, optional)
- `model_names` (list of string, optional) — Distinct model names observed across ingested sessions for this evaluation.
- `agent_names` (list of string, optional) — Distinct agent names observed across ingested sessions for this evaluation.
- `agent_versions` (list of string, optional) — Distinct agent versions observed across ingested sessions for this evaluation.
- `aggregate_scores` (map from string to EvaluatorAggregate, optional)
- `run_count` (integer, optional, default: 0) — Number of distinct ingested evaluation sessions; one session is treated as one run.
- `test_case_count` (integer, optional, default: 0) — Number of distinct test cases in the evaluation, i.e. distinct test_case_name values (sessions with no test_case_name each count as their own). A test case run k times counts once; the rollup metrics are averaged per test case before pooling across test cases.
- `cost_usd` (EvaluatorAggregate, optional) — Aggregate statistics over evaluator scores or session-level metric values.
- `latency_ms` (EvaluatorAggregate, optional) — Aggregate statistics over evaluator scores or session-level metric values.
- `tokens` (EvaluatorAggregate, optional) — Average total tokens (input + output) per test case, aggregated across the evaluation.

### PaginationData

- `page` (integer, required) — The current page number.
- `page_size` (integer, required) — The page size used for the query.
- `current_page_size` (integer, required) — The size for the current page.
- `total_pages` (integer, required) — The total number of pages.
- `total_results` (integer, required) — The total number of results.

### ValidationError

- `loc` (list of ValidationErrorLocItems, required)
- `msg` (string, required)
- `type` (string, required)
- `input` (any, optional)
- `ctx` (map from string to any, optional)

### DatetimeFilter

- `$gte` (datetime, optional) — Filter for results greater than or equal to this datetime.
- `$lte` (datetime, optional) — Filter for results less than or equal to this datetime.

### NumberFilter

- `$gte` (double, optional) — Filter for results greater than or equal to this value.
- `$lte` (double, optional) — Filter for results less than or equal to this value.
- `$gt` (double, optional) — Filter for results greater than this value.
- `$lt` (double, optional) — Filter for results less than this value.
- `$eq` (double, optional) — Filter for results equal to this value.

### MetricStatFilters

Numeric range filters keyed by rollup aggregate stat. Declaring each stat explicitly (rather than an open ``dict[str, NumberFilter]``) makes the valid stats visible in the OpenAPI schema, e.g. ``filter[cost_usd.mean][$lte]=0.5``. These stats must stay in sync with the runtime sort/filter grammar (``_METRIC_STATS`` in the evaluations endpoints); a unit test guards the parity.

- `sum` (NumberFilter, optional)
- `mean` (NumberFilter, optional)
- `median` (NumberFilter, optional)
- `p90` (NumberFilter, optional)
- `p95` (NumberFilter, optional)
- `p99` (NumberFilter, optional)
- `count` (NumberFilter, optional)

### EvaluatorAggregate

Aggregate statistics over evaluator scores or session-level metric values.

- `sum` (double, optional)
- `mean` (double, optional)
- `median` (double, optional)
- `p90` (double, optional)
- `p95` (double, optional)
- `p99` (double, optional)
- `count` (integer, optional, default: 0)
- `failed_count` (integer, optional, default: 0) — Sessions whose evaluator recorded a FAILED result (ran, produced no value). They are already counted as 0 in `mean`; this says how many of the attempts behind that mean were failures.

### ValidationErrorLocItems

## Examples

**Response**

```json
{
  "data": [
    {
      "id": "string",
      "name": "string",
      "workspace": "string",
      "experiment_ids": [
        "string"
      ],
      "dataset_name": "string",
      "experiment_group_id": "string",
      "dataset_version": "string",
      "source_link": "string",
      "metadata": {},
      "description": "string",
      "parent_evaluation_id": "string",
      "status": "string",
      "root_cause": "string",
      "created_at": "2024-01-15T09:30:00Z",
      "updated_at": "2024-01-15T09:30:00Z",
      "pinned_at": "2024-01-15T09:30:00Z",
      "evaluator_names": [
        "string"
      ],
      "model_names": [
        "string"
      ],
      "agent_names": [
        "string"
      ],
      "agent_versions": [
        "string"
      ],
      "aggregate_scores": {},
      "run_count": 0,
      "test_case_count": 0,
      "cost_usd": {
        "sum": 1.1,
        "mean": 1.1,
        "median": 1.1,
        "p90": 1.1,
        "p95": 1.1,
        "p99": 1.1,
        "count": 0,
        "failed_count": 0
      },
      "latency_ms": {
        "sum": 1.1,
        "mean": 1.1,
        "median": 1.1,
        "p90": 1.1,
        "p95": 1.1,
        "p99": 1.1,
        "count": 0,
        "failed_count": 0
      },
      "tokens": {
        "sum": 1.1,
        "mean": 1.1,
        "median": 1.1,
        "p90": 1.1,
        "p95": 1.1,
        "p99": 1.1,
        "count": 0,
        "failed_count": 0
      }
    }
  ],
  "pagination": {
    "page": 1,
    "page_size": 1,
    "current_page_size": 1,
    "total_pages": 1,
    "total_results": 1
  },
  "sort": "string",
  "filter": {}
}
```

**SDK Code**

```python
import requests

url = "https://api.example.com/apis/intake/v2/workspaces/workspace/evaluations"

response = requests.get(url)

print(response.json())
```

```javascript
const url = 'https://api.example.com/apis/intake/v2/workspaces/workspace/evaluations';
const options = {method: 'GET'};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
```

```go
package main

import (
	"fmt"
	"net/http"
	"io"
)

func main() {

	url := "https://api.example.com/apis/intake/v2/workspaces/workspace/evaluations"

	req, _ := http.NewRequest("GET", url, nil)

	res, _ := http.DefaultClient.Do(req)

	defer res.Body.Close()
	body, _ := io.ReadAll(res.Body)

	fmt.Println(res)
	fmt.Println(string(body))

}
```

```ruby
require 'uri'
require 'net/http'

url = URI("https://api.example.com/apis/intake/v2/workspaces/workspace/evaluations")

http = Net::HTTP.new(url.host, url.port)
http.use_ssl = true

request = Net::HTTP::Get.new(url)

response = http.request(request)
puts response.read_body
```

```java
import com.mashape.unirest.http.HttpResponse;
import com.mashape.unirest.http.Unirest;

HttpResponse<String> response = Unirest.get("https://api.example.com/apis/intake/v2/workspaces/workspace/evaluations")
  .asString();
```

```php
<?php
require_once('vendor/autoload.php');

$client = new \GuzzleHttp\Client();

$response = $client->request('GET', 'https://api.example.com/apis/intake/v2/workspaces/workspace/evaluations');

echo $response->getBody();
```

```csharp
using RestSharp;

var client = new RestClient("https://api.example.com/apis/intake/v2/workspaces/workspace/evaluations");
var request = new RestRequest(Method.GET);
IRestResponse response = client.Execute(request);
```

```swift
import Foundation

let request = NSMutableURLRequest(url: NSURL(string: "https://api.example.com/apis/intake/v2/workspaces/workspace/evaluations")! as URL,
                                        cachePolicy: .useProtocolCachePolicy,
                                    timeoutInterval: 10.0)
request.httpMethod = "GET"

let session = URLSession.shared
let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in
  if (error != nil) {
    print(error as Any)
  } else {
    let httpResponse = response as? HTTPURLResponse
    print(httpResponse)
  }
})

dataTask.resume()
```