> This page is for version 26.07 · v1.3.0.
> For other versions, use one of these documentation indexes:
> - Latest · v1.4.0 (26.09) (default): https://docs.nvidia.com/nemo/curator/latest/llms.txt
> - Main · preview: https://docs.nvidia.com/nemo/curator/main/llms.txt
> - 26.09 · v1.4.0: https://docs.nvidia.com/nemo/curator/v26.09/llms.txt
> - 26.07 · v1.3.0: https://docs.nvidia.com/nemo/curator/v26.07/llms.txt
> - 26.04 · v1.2.0: https://docs.nvidia.com/nemo/curator/v26.04/llms.txt
> - 26.02 · v1.1.0: https://docs.nvidia.com/nemo/curator/v26.02/llms.txt
> - 25.09 · v1.0.0: https://docs.nvidia.com/nemo/curator/v25.09/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/curator/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/curator/_mcp/server.

# nemo_curator.tasks.document

## Module Contents

### Classes

| Name                                                          | Description                                    |
| ------------------------------------------------------------- | ---------------------------------------------- |
| [`DocumentBatch`](#nemo_curator-tasks-document-DocumentBatch) | Task for processing batches of text documents. |

### API

```python
class nemo_curator.tasks.DocumentBatch(
    dataset_name: str,
    data: pyarrow.Table | pandas.DataFrame = pa.Table(),
    _stage_perf: list[nemo_curator.utils.performance_utils.StagePerfStats] = list(),
    _metadata: dict[str, typing.Any] = dict()
)
```

Dataclass

**Bases:** [Task\[Table | DataFrame\]](/nemo/curator/nemo-curator/nemo_curator/tasks/tasks#nemo_curator-tasks-tasks-Task)

Task for processing batches of text documents.
Documents are stored as a dataframe (PyArrow Table or Pandas DataFrame).

**`data`** `Table | DataFrame = field(default_factory=(pa.Table))`

---

**`num_items`** `int`

Get the number of documents in this batch.

---

```python
nemo_curator.tasks.DocumentBatch.get_columns() -> list[str]
```

Get column names from the data.

```python
nemo_curator.tasks.DocumentBatch.to_pandas() -> pandas.DataFrame
```

Convert data to Pandas DataFrame.

```python
nemo_curator.tasks.DocumentBatch.to_pyarrow() -> pyarrow.Table
```

Convert data to PyArrow table.

```python
nemo_curator.tasks.DocumentBatch.validate() -> bool
```

Validate the task data.