nemo_gym.orchestration.executors.otel
nemo_gym.orchestration.executors.otel
The OpenTelemetry collector that runs beside every benchmark job.
One collector per benchmark job, in its own srun --overlap step like any other service. It
scrapes each model service’s Prometheus /metrics on localhost, accepts OTLP from the job’s own
processes, stamps the resource attributes the shared dashboards filter on (user, run_id,
slurm_job_id, plus benchmark, cluster, model), and exports to the configured OTLP/HTTP
endpoint and to <job dir>/otel/*.jsonl at the same time.
The ingest token travels twice, as the HTTP Authorization header and as an Authorization
resource attribute: a routing proxy in front of the backend resolves the tenant from the
attribute and silently drops payloads without it while still answering 200.
Module Contents
Functions
Data
API
Environment that switches on Gym’s Lens instrumentation and points it at the collector.
Gym reads these in every server process (NEMO_GYM_OTEL_* are Gym’s own, the OTEL_* ones
are the SDK’s); an explicit value in driver.env wins over these.
Whether the driver is instrumented with nemo-lens and pointed at the collector.
Enabled is enough: Gym’s own servers push telemetry even when there is no local model to scrape.
The collector’s YAML for one benchmark job.
run_id is the submission’s gym job id, which is the job directory’s parent by construction
(<output_path>/<gym_job_id>/<benchmark>); cluster is the sole compute key, the same value
SubmissionRecord.cluster records. Values only known inside the job (SLURM_JOB_ID, the
token) are left as ${env:...} for the collector to expand at startup.
The ingest token from the submitting environment; a missing one fails the submit.
Service name to serving port for every service that exposes a model (and so /metrics).
An enabled collector needs somewhere to send to; a bare default config has none.