FLUX Pipeline Tuning
This example demonstrates how to use NVIDIA AITune to tune the Flux text-to-image model from Hugging Face’s diffusers library.
Environment Setup
You can use either of the following options to set up the environment:
Option 1 - virtual environment managed by you
Activate your virtual environment and install the dependencies:
Option 2 - virtual environment managed by uv
Install dependencies:
Usage
Tuning the model
To tune the Flux model, run:
You can customize the following parameters:
--model-name: HuggingFace model name or path (default: “black-forest-labs/FLUX.1-dev”)--prompt: Text prompt for image generation--sizes: Space-separatedwidth,heightimage sizes (default:512,512 1024,1024)--steps: Number of inference steps (default: 28)--guidance-scale: Guidance scale (default: 3.5)--max-sequence-length: Maximum sequence length (default: 128)--tuned-model-path: Path to save or load the tuned model (default:flux-dev.ait)
Generating images with the tuned model
After tuning, generate images with:
The generated image will be saved in AITUNE_OUTPUT_DIR, or output when the environment variable is not set.
Logging hardware metrics
If you would like to log hardware metrics during tuning or inference, export AITUNE_HARDWARE_METRICS=True environment variable, e.g.
AI Dynamo FLUX Deployment
To run FLUX as an AI Dynamo service, you need to first tune your model and then launch a test run using the provided script.
This script starts the frontend and backend services, waits for them to be ready, then runs a test client to send a sample request. Once the test completes, all services are automatically shut down. This is meant as a functional check, not to provide a permanent server.
Dynamic batching
The service uses dynamic batching — requests are grouped and processed together for efficiency. Currently, there is one frontend and one worker. To support multiple workers, move batching to a separate service that handles request grouping.
Model Details
The Flux model is a text-to-image diffusion model that generates high-quality images from text descriptions. The model is trained on a large dataset of images and text, and can generate realistic images across various domains.
For more information, visit the Flux model page on HuggingFace.