NVCF Long Requests (h2-ping-sidecar)
NVCF Long Requests (h2-ping-sidecar)
NVCF Long Requests (h2-ping-sidecar)
NVCF endpoints (*.invocation.api.nvcf.nvidia.com, integrate.api.nvidia.com) sit behind AWS Global
Accelerator, which drops a connection that carries no application data for 340 seconds. The limit
is fixed by AWS and TCP keepalive does not reset it. A request that takes longer than 340 seconds to
produce its first response byte, such as a large-prompt judge or reasoning model, fails with
ServerDisconnectedError after about 340 seconds.
Which fix do I need?
Gym has more than one defence against connections that go silent. They fix different drops:
If a request dies at almost exactly 340 seconds with ServerDisconnectedError against NVCF, you need the
sidecar. If connections die much sooner, or somewhere that is not Global Accelerator, look at the TCP
keepalive settings first.
How it works
HTTP/2 PING frames are application data and do reset the timer, but Gym’s aiohttp client speaks only
HTTP/1.1. The h2-ping-sidecar closes the gap. It listens on 127.0.0.1, accepts plain HTTP/1.1
from Gym, forwards each request to the endpoint over HTTP/2, and sends a PING whenever the connection has
been idle for ping_interval. Gym needs no client changes: its model URL points at localhost and the
sidecar forwards the call.
The sidecar holds no credentials. Your API key stays in Gym’s config and the Authorization header passes
through unchanged. The sidecar has no authentication of its own, so keep it on a loopback address; Gym
prints a warning at startup if an instance’s listen is not one.
Enable it
In training the model that needs the sidecar is usually a judge such as GenRM, served from NVCF while
the policy runs locally over plain http://. List the judge’s host as an instance. Gym starts the sidecar
before any server, rewrites the judge’s base_url to the local address, and stops the sidecar after the
servers.
instances is required. Each entry names one upstream and gets its own sidecar process on its own
listen port. To send a second endpoint through the sidecar too, such as a policy model that is also on
NVCF, add another entry with a different port and host.
Gym rewrites every *base_url key (a string or a list) of a server under responses_api_models, such as
openai_base_url or the base_url of a genrm_model, whose host matches an instance’s upstream. It
matches the resolved URL, so a URL that comes from a resolver such as ${oc.env:JUDGE_URL}, or from an
alias such as openai_base_url: ${policy_base_url}, is routed too, and becomes a literal local address in
Gym’s in-memory config for that run. Top-level keys such as policy_base_url are not rewritten, and a list
that aliases another list is rewritten as a copy, so the list it points at never changes. Headers such as
NVCF-POLL-SECONDS in default_headers are untouched. Only scheme://host[:port] of an upstream is used;
the path stays on each request.
Where it runs
Gym starts the sidecar as a child process of the same process that starts the Gym servers, so it is always
on the node the model servers run on. That includes a Gym run inside a NeMo-RL Ray actor: wherever the
actor lands, its servers and its sidecar land together. There is no per-node setting. The binary only has
to exist on that node.
Gym waits for each sidecar to bind its port before starting servers, then checks them on every poll of
gym env start. If one dies, the run stops with the tail of its log. A port that is already in use fails
startup with a message that says so. If something else already starts a sidecar on those ports (for example
a Slurm launcher script), remove one of them.
Settings
All keys are optional except enabled and instances.
Durations use Go syntax: 60s, 15m, 1m30s.
Sizing memory
A request body is kept in memory from the moment it is read until the transport is done with the request, so a refused request can be re-sent. For this feature that is the whole wait for the first response byte, which can be minutes, and the garbage collector cannot free any of it meanwhile. Peak memory is therefore about:
For example 2,000 concurrent judge calls with 0.5 MB bodies hold about 1 GB. Size retry_body_limit for
the largest body you expect to retry; larger bodies are streamed through unbuffered and are not retried, so
they never count against it.
gomemlimit is a soft limit. When live memory is above it Go does not fail: it keeps collecting, and can
spend a large share of the sidecar’s CPU doing so. Set it above the working set computed above, with
some headroom. A value below it slows the sidecar down instead of capping its memory.
Try it
When the servers are up you should see lines like:
In another terminal, run your evaluation (gym eval run --no-serve ...). To confirm that PINGs keep a long
request alive, send one that takes longer than 340 seconds to produce its first byte and check that it
returns instead of failing with ServerDisconnectedError. The sidecar log is
<log_dir>/h2ping-<instance>-<host>-<run id>.log. Ctrl-C drains in-flight requests (up to
shutdown_grace) and exits.
Limits
- HTTP/2 only. The sidecar never falls back to HTTP/1.1, which has no PING frames.
- Load balancer idle timeout. The load balancer behind Global Accelerator has its own idle timeout
that PINGs do not reset (3600 s on per-function
*.invocation.api.nvcf.nvidia.comhosts, lower onintegrate.api.nvidia.com, 1200 s at the time of writing). A request silent for longer than that fails with or without the sidecar; use the per-function hostname for the longest requests. - Connection rotation. The load balancer closes a client connection with a GOAWAY once it is about an
hour old. The sidecar retires its upstream connection first (
max_conn_age), so new requests move to a fresh connection and requests in flight finish on the old one. A request that is still running when the limit is reached can still meet the GOAWAY; a body withinretry_body_limitis re-sent on a fresh connection, a larger one is not. A very long request can therefore still fail once, and Gym’s normal retries apply. - One late PING reply fails every request on that connection. Gym’s aiohttp client spreads requests over
many connections, but the sidecar folds them into a few HTTP/2 connections, and a load balancer allows
many concurrent requests on each (128 on an AWS ALB). If a PING reply arrives later than
ping_timeout, Go closes the connection and every request on it fails at once, each possibly minutes into its wait; a retry then resends the whole batch. Under heavy load a PING reply can queue behind response data on the same connection. Raiseping_timeoutif you see this, and watch forupstream errorlines in the sidecar log. This has not been measured at training scale yet. - Do not replace aiohttp. Gym stays on aiohttp and HTTP/1.1 (see aiohttp vs httpx); the sidecar is what speaks HTTP/2.
- Only
gym env start,gym eval runand the other commands that start servers launch the sidecar. Servers started some other way (for example onesrunstep per server) must start their own, as described innemo_gym/tools/sidecar/README.md.