> This page is for version 26.07 · v1.3.0.
> For other versions, use one of these documentation indexes:
> - Latest · v1.4.0 (26.09) (default): https://docs.nvidia.com/nemo/curator/latest/llms.txt
> - Main · preview: https://docs.nvidia.com/nemo/curator/main/llms.txt
> - 26.09 · v1.4.0: https://docs.nvidia.com/nemo/curator/v26.09/llms.txt
> - 26.07 · v1.3.0: https://docs.nvidia.com/nemo/curator/v26.07/llms.txt
> - 26.04 · v1.2.0: https://docs.nvidia.com/nemo/curator/v26.04/llms.txt
> - 26.02 · v1.1.0: https://docs.nvidia.com/nemo/curator/v26.02/llms.txt
> - 25.09 · v1.0.0: https://docs.nvidia.com/nemo/curator/v25.09/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/curator/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/curator/_mcp/server.

# nemo_curator.core.utils

## Module Contents

### Functions

| Name                                                                                  | Description                                                              |
| ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ |
| [`_logger_custom_deserializer`](#nemo_curator-core-utils-_logger_custom_deserializer) | -                                                                        |
| [`_logger_custom_serializer`](#nemo_curator-core-utils-_logger_custom_serializer)     | -                                                                        |
| [`_ray_serve_haproxy_source`](#nemo_curator-core-utils-_ray_serve_haproxy_source)     | Return the configured Ray Serve HAProxy source, if available.            |
| [`check_ray_responsive`](#nemo_curator-core-utils-check_ray_responsive)               | -                                                                        |
| [`get_free_port`](#nemo_curator-core-utils-get_free_port)                             | Checks if start\_port is free.                                           |
| [`ignore_ray_head_node`](#nemo_curator-core-utils-ignore_ray_head_node)               | Return True if `CURATOR_IGNORE_RAY_HEAD_NODE` is set to a truthy value.  |
| [`init_cluster`](#nemo_curator-core-utils-init_cluster)                               | Initialize a new local Ray cluster or connects to an existing one.       |
| [`split_table_by_group`](#nemo_curator-core-utils-split_table_by_group)               | Split an Arrow table without reordering or splitting consecutive groups. |

### API

```python
nemo_curator.core.utils._logger_custom_deserializer(
    _: None
) -> loguru.Logger
```

```python
nemo_curator.core.utils._logger_custom_serializer(
    _: loguru.Logger
) -> None
```

```python
nemo_curator.core.utils._ray_serve_haproxy_source() -> str | None
```

Return the configured Ray Serve HAProxy source, if available.

```python
nemo_curator.core.utils.check_ray_responsive(
    timeout_s: int = RAY_CLUSTER_START_VERIFICAT...
) -> bool
```

```python
nemo_curator.core.utils.get_free_port(
    start_port: int,
    get_next_free_port: bool = True,
    bind_host: str = 'localhost'
) -> int
```

Checks if start\_port is free.
If not, it will get the next free port starting from start\_port if get\_next\_free\_port is True.
Else, it will raise an error if the free port is not equal to start\_port.
bind\_host controls the interface used for the probe.

```python
nemo_curator.core.utils.ignore_ray_head_node() -> bool
```

Return True if `CURATOR_IGNORE_RAY_HEAD_NODE` is set to a truthy value.

Used by both the pipeline executors (to skip the head node when scheduling
stage actors) and the inference-server backends (to emit a worker-only
bundle-label selector on placement groups).

```python
nemo_curator.core.utils.init_cluster(
    ray_port: int,
    ray_temp_dir: str,
    ray_dashboard_port: int,
    ray_metrics_port: int,
    ray_client_server_port: int,
    ray_dashboard_host: str,
    num_gpus: int | None = None,
    num_cpus: int | None = None,
    object_store_memory: int | None = None,
    enable_object_spilling: bool = False,
    block: bool = True,
    ip_address: str | None = None,
    stdouterr_capture_file: str | None = None
) -> subprocess.Popen
```

Initialize a new local Ray cluster or connects to an existing one.

```python
nemo_curator.core.utils.split_table_by_group(
    table: pyarrow.Table,
    group_column: str,
    max_batch_bytes: int | None = None,
    max_batch_rows: int | None = None
) -> list[pyarrow.Table]
```

Split an Arrow table without reordering or splitting consecutive groups.

Rows for each group must be consecutive. If a single group exceeds a batch
limit, it is still emitted as one chunk.