nemo_gym.orchestration.executors.otel
nemo_gym.orchestration.executors.otel
The OpenTelemetry collector that runs beside every benchmark job.
One collector per benchmark job, in its own srun --overlap step like any other service. It
scrapes each model service’s Prometheus /metrics on localhost, accepts OTLP from the job’s own
processes, stamps the resource attributes the shared dashboards filter on (user, run_id,
slurm_job_id, plus benchmark, cluster, model), and exports to the configured OTLP/HTTP
endpoint and to <job dir>/otel/*.jsonl at the same time.
The ingest token travels twice, as the HTTP Authorization header and as an Authorization
resource attribute: a routing proxy in front of the backend resolves the tenant from the
attribute and silently drops payloads without it while still answering 200.
Module Contents
Functions
Data
API
Whether a collector step is added to this job: enabled, and there is something to scrape.
The collector’s YAML for one benchmark job.
run_id is the submission’s gym job id, which is the job directory’s parent by construction
(<output_path>/<gym_job_id>/<benchmark>); cluster is the sole compute key, the same value
SubmissionRecord.cluster records. Values only known inside the job (SLURM_JOB_ID, the
token) are left as ${env:...} for the collector to expand at startup.
The ingest token from the submitting environment; a missing one fails the submit.
Service name to serving port for every service that exposes a model (and so /metrics).
An enabled collector needs somewhere to send to; a bare default config has none.