nemo_rl.data.captured_media#
vLLM media extraction and placeholder remapping for shared token capture.
Pixels use the same packed imgs bundle as Megatron inference. Only the
small media_spans extras are vLLM-specific: they restore placeholder
positions when expanded-token history is spliced into a rendered prompt.
Module Contents#
Classes#
One new placeholder occurrence in the expanded token sequence. |
|
New placeholder metadata and owned tensors for the common staging sink. |
Functions#
Losslessly match Bridge Omni’s dynamic-image patch layout. |
|
Remap vLLM placeholders and snapshot per-call media in Megatron’s layout. |
API#
- exception nemo_rl.data.captured_media.MediaCaptureRejected(
- message: str,
- *,
- code: str = 'media_capture_rejected',
Bases:
ValueErrorA captured call’s media cannot be staged; the request is rejected before inference.
codeis a stable, machine-readable reason surfaced in the HTTP 400 body and the worker log.retained_media_changedmarks a retained image or video whose geometry or placeholder tokens differ from the staged occurrence (e.g. vLLM re-tiled it under a tighter token budget); every other capture-time validation failure usesmedia_capture_rejected.Initialization
Initialize self. See help(type(self)) for accurate signature.
- nemo_rl.data.captured_media._token_digest(tokens: list[int]) str#
- class nemo_rl.data.captured_media.CapturedMediaItem#
One new placeholder occurrence in the expanded token sequence.
- modality: Literal[image, video]#
None
- placeholder_offset: int#
None
- placeholder_length: int#
None
- placeholder_digest: str#
None
- token_id: int#
None
- embedding_spans: tuple[tuple[int, int], ...]#
None
- imgs_sizes: tuple[tuple[int, int], ...]#
None
- __post_init__() None#
- property end: int#
- to_dict() dict[str, Any]#
- classmethod from_dict(
- value: dict[str, Any],
- verify_tokens(tokens: list[int], *, origin: int) None#
- class nemo_rl.data.captured_media.CapturedMedia#
New placeholder metadata and owned tensors for the common staging sink.
- items: tuple[nemo_rl.data.captured_media.CapturedMediaItem, ...]#
None
- tensors: dict[str, torch.Tensor] | None#
None
- nemo_rl.data.captured_media._geometry_tensor(value: Any) torch.Tensor#
- nemo_rl.data.captured_media.pack_images(
- pixels: torch.Tensor,
- sizes: torch.Tensor,
- *,
- patch_size: int,
Losslessly match Bridge Omni’s dynamic-image patch layout.
This only rearranges processed pixels, preserving dtype and values. Frames remain in order and their temporal grouping travels separately in num_frames.
Mirrors
NemotronOmniModel._patchify_dynamic_imagesin Megatron-Bridge (the vLLM worker env has no Bridge);test_captured_media_pack_images_matches_bridge_patchify(mcore lane) guards the two against drift.
- nemo_rl.data.captured_media._processed_omni_tensors(
- data: dict[str, Any],
- modality: str,
- *,
- patch_size: int,
- nemo_rl.data.captured_media.capture_processed_media(
- engine_prompt: dict[str, Any],
- *,
- prev_len: int,
- retained: tuple[nemo_rl.data.captured_media.CapturedMediaItem, ...] = (),
- splice: nemo_rl.models.generation.openai_server_utils.PrefixSplice | None = None,
- image_token_id: int | None = None,
- patch_size: int | None = None,
Remap vLLM placeholders and snapshot per-call media in Megatron’s layout.