bridge.models.bagel.modeling#
BAGEL vision and diffusion modality modules.
Module Contents#
Classes#
Wrap BAGEL’s packed SigLIP encoder without changing its inputs. |
|
Project packed SigLIP embeddings and add BAGEL positions. |
|
Combine noisy latents with BAGEL timestep and position embeddings. |
API#
- class bridge.models.bagel.modeling.OfficialBagelVisionEncoder(
- *,
- bagel_config: Any,
- vision_model_path: str | None,
- dtype: torch.dtype,
- recompute: bool,
Bases:
torch.nn.ModuleWrap BAGEL’s packed SigLIP encoder without changing its inputs.
Initialization
- _enable_recompute() None#
Checkpoint each official SigLIP encoder layer.
- forward(
- packed_vit_tokens: torch.Tensor,
- packed_vit_position_ids: torch.Tensor,
- vit_token_seqlens: torch.Tensor,
Encode packed patches and return BAGEL’s post-connector positions.
- class bridge.models.bagel.modeling.BagelVisionSubmodule#
Bases:
megatron.core.models.mimo.submodules.vision.VisionModalitySubmodulesProject packed SigLIP embeddings and add BAGEL positions.
- forward(encoder_inputs: dict[str, Any]) torch.Tensor#
Run the one BAGEL vision encoder and connector.
- class bridge.models.bagel.modeling.BagelDiffusionSubmodule(
- *args: Any,
- dtype: torch.dtype = torch.float32,
- **kwargs: Any,
Bases:
megatron.core.models.mimo.submodules.base.ModalitySubmodulesCombine noisy latents with BAGEL timestep and position embeddings.
Initialization
- encode(
- encoders_data_batch: dict[str, torch.Tensor],
Encode packed timesteps and latent positions.
- abstractmethod decode(
- embeddings: torch.Tensor,
- data_batch: dict[str, Any],
Reject decoding because training predicts velocity directly.
- forward(encoder_inputs: dict[str, torch.Tensor]) torch.Tensor#
Build the visual-generation token embeddings.
- llm2vae(embeddings: torch.Tensor) torch.Tensor#
Project language hidden states into latent-patch velocity.