- Output-preserving / lossless-style: system settings that should preserve model behavior while changing residency, parallelism, kernels, or scheduling.
- Quality-tradeoff / lossy or approximate: techniques that can change the denoising path, numerical representation, or generated output.
Start Here
- Pick a serving or generation mode from Deployment and Performance Modes.
--performance-mode autois the default; usespeedwhen the model fits in GPU memory and latency matters most,memorywhen GPU memory is the bottleneck, andmanualwhen every performance flag should be explicit. - Choose the right attention backend from Attention Backends.
- Use Sequence Parallelism only when the model and video shape benefit from sequence splitting.
- Use Inference Batching for concurrent compatible requests during serving.
- Use Profiling before changing several levers at once.
Choose a request quality tier
--quality is cumulative: a broader tier never drops an optimization from a
stricter tier.
Use
extra-high when you want to isolate fusion wins from approximate
acceleration. A tier may be a no-op when the active model has no eligible path.
Separately configured quantization, attention, or caching options still apply.
See Fused Kernels for the current
request-gated families and their numerical contracts.
Output-Preserving / Lossless-Style Levers
These settings should preserve model behavior while changing residency, parallelism, kernels, or scheduling. They are the first choices for production tuning.| Lever | Use when | Docs |
|---|---|---|
—performance-mode | You want a safe preset for speed or memory without overriding explicit flags. | Deployment and Performance Modes |
| Breakable CUDA graph | A supported pipeline serves a fixed set of shapes and eager execution is launch-bound. | CLI reference |
| Offload, FSDP, CFG parallelism | GPU memory, multi-GPU residency, or CFG branch splitting is the main bottleneck. | Deployment and Performance Modes |
| Sequence parallelism | Long image/video sequences need sequence-level parallelism. | Sequence Parallelism |
—encoder-parallel | Text/image encoding is a visible share of the request and the DiT replica sits idle during it. | Encoder Parallelism |
| Attention backend | Kernel choice dominates DiT latency or memory. | Attention Backends |
| Fused kernels | You want to know which elementwise chains are already fused, or to opt into the request-gated set. | Fused Kernels |
| Dynamic batching | Serving many compatible requests concurrently. | Inference Batching |
Quality-Tradeoff / Lossy Or Approximate Levers
These techniques can change the denoising path, numerical representation, or generated output. They are useful after you have a baseline and an acceptance criterion for quality.| Lever | Tradeoff | Docs |
|---|---|---|
| Cache-DiT | Skips selected DiT block or step computation based on cache decisions. | Cache-DiT |
| TeaCache | Reuses residuals when consecutive denoising steps are similar enough. | TeaCache |
| Progressive resolution | Runs early denoising at lower latent resolution for supported pipelines. | Progressive Resolution Generation |
| Quantization | Uses lower-precision transformer weights or activations. | Quantization |
Practical Order
- Establish a baseline with the target model, resolution, frame count, step count, and GPU type.
- Select
--performance-modeand explicit residency or parallelism flags. - Compare
quality=losslesswithquality=extra-highto isolate the request-gated fusion set. - Compare breakable CUDA graph against eager execution for supported fixed-shape pipelines. Pass every served resolution to
--warmup-resolutionsand confirm capture in the server log. Models with request-gated DiT fusions cannot combine those fusions with a graph captured from the lossless branches. - Tune attention backend and batching for the deployment pattern.
- Profile if the bottleneck is unclear.
- Add
quality=high, caching, progressive resolution, or quantization only after comparing output quality against your acceptance target.
Per-model tuning starting points
Warmup and breakable CUDA graph (BCG) solve different problems.--warmup-mode request runs a warmup copy derived from the first request to prime one-time compilation and caches; it does not remove recurring Python launches from each denoising step. BCG captures supported DiT segments and can reduce that recurring launch overhead for captured shapes. Use the table below as a first experiment, then keep a lever only when profiling confirms that it addresses the active bottleneck.
The example assignments come from single-GPU NVIDIA B300/GB300 profiles and are directional rather than an exhaustive compatibility list. A model can match more than one row as resolution, frame count, step count, parallelism, or GPU type changes; the measured bottleneck takes precedence over the model name.
BCG is enabled only for model and pipeline configurations accepted by the runtime support check. It captures the default warmup shape automatically. Use
--warmup-resolutions for additional served resolutions; for video and variable prompt lengths, set --warmup-num-frames and --bcg-text-buckets to cover the intended workload. Requests that do not match a captured signature fall back to eager execution.
Warmup and BCG do not intentionally trade output quality for speed, but that is not a guarantee of byte-identical output. Different execution paths can introduce numerical differences, and some pipelines customize the scheduler or schedule used by a synthetic warmup request. Before production rollout, compare output hashes when bitwise stability is required, otherwise run the project’s quality acceptance check.
Performance results are configuration-dependent. Record the exact checkpoint revision, GPU, precision, resolution, frame count, step count, parallelism, command line, and whether the measurement includes warmup. Re-profile with Profiling on the target workload rather than transferring a percentage from another model or shape.
