Skip to content

Repository files navigation

Diffusers Dynamic Offloader

diffusers-dynamic-offloader (DDO) is a model-agnostic offload router for Diffusers-style inference on low-VRAM systems.

DDO can either apply its own dense-linear dynamic offload path or route a component to the official Diffusers group-offload implementation. The intent is to give runners a small API surface:

from diffusers_dynamic_offloader import enable_offload

result = enable_offload(transformer, preset="auto", component="transformer")
transformer = result.module

The library resolves the preset, platform behavior, component policy, RAM/VRAM budget, quantized-backend compatibility, and optional Windows standby-list cleanup helpers. Host scripts should not need to decide manually whether a given component should use DDO dynamic offload, Diffusers group offload, or no offload.

Install

From a host project:

uv add git+https://github.com/DEVAIEXP/diffusers-dynamic-offloader.git

For local development:

git clone https://github.com/DEVAIEXP/diffusers-dynamic-offloader.git
cd diffusers-dynamic-offloader
uv pip install -e .

Or, from another project environment, install the local checkout by path:

uv pip install -e C:\path\to\diffusers-dynamic-offloader

Minimal Full-Pipeline Usage

Use enable_pipeline_offload(...) when you have a complete Diffusers pipeline and want the shortest integration.

import torch
from diffusers import DiffusionPipeline
from diffusers_dynamic_offloader import enable_pipeline_offload

pipe = DiffusionPipeline.from_pretrained(
    "your/model",
    torch_dtype=torch.bfloat16,
)

enable_pipeline_offload(pipe, preset="auto", execution_device="cuda:0")

image = pipe(prompt="a small robot holding a lantern", output_type="pil").images[0]
image.save("out.png")

This is the simplest path. It is also the least memory-aware path because the pipeline owns the full call and DDO cannot insert cleanup between prompt encoding, denoise, and VAE decode.

Minimal Modular Usage

For Modular Diffusers, load the custom blocks and components, then apply DDO to the exposed modules:

import torch
from diffusers import ModularPipeline
from diffusers_dynamic_offloader import enable_pipeline_offload

pipe = ModularPipeline.from_pretrained(
    "your/custom_blocks",
    trust_remote_code=True,
)
pipe.load_components(
    pretrained_model_name_or_path="your/model",
    names=["text_encoder", "tokenizer", "transformer", "vae", "scheduler"],
    torch_dtype=torch.bfloat16,
)

enable_pipeline_offload(pipe, preset="auto", execution_device="cuda:0")

image = pipe(
    prompt="a small robot holding a lantern",
    width=1024,
    height=1024,
    num_inference_steps=8,
    output="images",
)[0]
image.save("out.png")

When a component cannot be accelerated safely, DDO routes it through the compatibility path instead of replacing its forward.

Usage Guides

Recommended Patterns

Use a full pipeline when:

  • you want the smallest code change;
  • the model already fits comfortably enough;
  • you mainly want one preset API over Diffusers group offload and DDO dynamic offload.

Use phase-staged execution when:

  • VRAM or system RAM is tight;
  • prompt encoding, denoise, and VAE decode have very different memory shapes;
  • you want Windows standby-list cleanup and CUDA cleanup between phases;
  • you need apples-to-apples timing between full-pipeline and phase-by-phase execution.

Use official Diffusers group offload through DDO when:

  • the model uses a quantized backend such as SDNQ, bitsandbytes, TorchAO, or Quanto;
  • the backend has custom packed weights or custom forward logic;
  • compatibility matters more than dense-linear streaming speed.

Presets

Preset Main use Transformer route Notes
auto Default entry point Resolves to the recommended policy for the current DDO release Current policy prefers one_shot_fast, with platform and backend safety checks.
one_shot_fast Best Windows low-VRAM BF16 path found so far DDO dynamic offload for dense transformer linears Uses RAM-aware CPU pinning and a VRAM-headroom-aware resident-module budget; when usable RAM fits all eligible weights and the model is at least 4x larger than VRAM, it instead selects true full pin with zero CUDA-resident modules to preserve activation headroom.
low_ram_safe Explicit constrained-memory fallback DDO dynamic offload with lower resident pressure Slower, but useful when peak accelerator memory matters more than latency.
wsl_compat WSL stability fallback DDO dynamic offload with WSL-safe settings Avoids pinned CPU memory and stream combinations that were unstable in testing.
warm_process Server/repeated generations DDO dynamic offload with process-lifetime caches Not a cold one-image latency preset. Useful for warm runners and services.
diffusers_offload_compat Official block-level baseline Diffusers group offload Compatibility path and benchmark baseline.
diffusers_leaf_offload_compat Official leaf-level baseline Diffusers group offload Often best for quantized backends and some native-Linux cold runs.
off Disable offload routing None Caller must move modules/devices manually.

Platform Behavior

auto is intentionally model-agnostic. It does not hardcode LTX, FLUX, Z-Image, or any other model family. It inspects the component, available RAM, CUDA VRAM, platform, and preset policy.

Platform Recommended first try Why
Windows + low VRAM + enough system RAM auto / one_shot_fast Best measured DDO win: faster denoise than official Diffusers group offload while keeping memory bounded.
Windows + limited system RAM auto, then low_ram_safe if needed auto skips pinned CPU weights when RAM headroom is not enough; denoise is slower but safer.
WSL wsl_compat WSL showed unstable pin/stream behavior; the compatibility preset disables the risky parts.
Native Linux compare auto and diffusers_leaf_offload_compat DDO can improve denoise, but official Diffusers leaf offload may win cold one-shot totals when setup cost dominates.
Quantized backends diffusers_leaf_offload_compat or diffusers_offload_compat DDO preserves backend-specific forwards and does not replace packed quantized matmuls.

Quantized Models

DDO's acceleration path targets standard dense PyTorch nn.Linear modules with 2D weights. Quantized integrations frequently store packed weights and implement their own dequantization, layout transform, Hadamard transform, or matmul path.

For those modules, DDO preserves the backend forward and routes to the Diffusers compatibility path. This is deliberate: replacing a quantized forward may be faster in one local experiment but fragile across backend versions and hardware.

Configuration

Application code can pass explicit settings. See the configuration reference for every settings field and DDO-recognized environment variable:

from diffusers_dynamic_offloader import DynamicOffloadSettings, enable_offload

settings = DynamicOffloadSettings.from_env(
    default_preset="auto",
    execution_device="cuda:0",
    offload_device="cpu",
)

result = enable_offload(
    transformer,
    settings=settings,
    component="transformer",
    max_resident_module_budget_gb=6.0,
)

Environment variables are optional overrides, useful for runners and benchmarks:

$env:DDO_PRESET="one_shot_fast"
$env:DDO_MAX_RESIDENT_MODULE_BUDGET_GB="6"
$env:DDO_SHOW_PROFILE="1"

The automatic full-pin/zero-resident path is reported during setup. Its default 4x model-to-VRAM threshold can be changed with DDO_AUTO_FULL_PIN_MIN_MODEL_TO_VRAM_RATIO; set it to 0 to disable that automatic choice. A stored capacity profile may internally apply a conservative residual CUDA module budget. DDO_RESIDENT_MODULE_BUDGET_GB is the explicit manual override and takes precedence.

Named workload profiles keep calibration inside DDO:

offload = enable_offload(model, profile_name="my-workload", build_profile=calibrate)
with offload.profile_run():
    output = pipeline(**inputs)

Use calibrate=True once, then False to reuse the recommendation. Choose a different name for different workloads or recalibrate after changing them. See named workload profiles for staged execution, loading integration and measurement details.

Windows standby-cache purge

one_shot_fast enables optional standby-list purge at selected stage boundaries through DDO_PURGE_WINDOWS_STANDBY_*. On Windows, the process token must be allowed to enable SeProfileSingleProcessPrivilege, which normally means starting the terminal as Administrator. Without that privilege DDO reports purged: false and continues safely; the purge is optional and does not release DDO/PyTorch pinned weight memory.

Runner-specific variables should use a separate prefix such as DDO_RUNNER_*; DDO itself only owns the DDO_* library settings.

Results

The current report is in experiments/dynamic_offload_results.md. It includes:

  • Windows, WSL, and native-Ubuntu BF16 benchmark comparisons;
  • DDO dynamic offload versus official Diffusers block/leaf group offload;
  • full-pipeline versus phase-staged execution comparisons;
  • quantized SDNQ compatibility notes;
  • exploratory FLUX.2 Klein and Z-Image observations for auto-budget behavior.

Project Scope

DDO intentionally avoids native low-level memory hooks. The priority is a portable PyTorch/Diffusers layer that is easy to integrate, easy to disable, and safe across different hardware and driver stacks.

About

Model-agnostic dynamic offload helpers for Diffusers-style low-VRAM inference.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages