Skip to content

fix(minimax): enable autocast on all accelerators in VAE decoding - #14884

Closed
li-lizhe wants to merge 2 commits into
huggingface:mainfrom
li-lizhe:fix/minimax-autocast-device-v2
Closed

li-lizhe wants to merge 2 commits into
huggingface:mainfrom
li-lizhe:fix/minimax-autocast-device-v2

Conversation

@li-lizhe

Copy link
Copy Markdown

Fixes #14882

Description: The VAE decode block in the MiniMax H3 modular pipeline only enables fp16 autocast on CUDA devices (enabled=device.type == "cuda"), which disables autocast on other accelerators such as Ascend NPU, Intel XPU, or Apple MPS.

Without autocast on NPU, the VAE decode runs in fp32 (or the tensor's dtype without fp16 acceleration), causing suboptimal performance and potential dtype mismatches when the pipeline expects half-precision computation.

Change:

  • enabled=device.type == "cuda" -> enabled=device.type != "cpu"
  • Autocast fp16 is now enabled on all accelerator devices while remaining disabled on CPU

Verification on Ascend 910B NPU (torch 2.14.0a0 + torch_npu):

  • torch.autocast(device_type="npu", dtype=torch.float16, enabled=True) produces fp16 output (out.dtype=torch.float16)
  • torch.autocast(device_type="npu", dtype=torch.float16, enabled=False) produces fp32 output

Notes: (1) end-to-end NPU decode speed-up of this change is not measured - only the dtype path above was checked; (2) CUDA behaviour is unchanged; (3) the branch also carries the get_i2v_mask device-default commit shared with #14765; (4) this is unrelated to the separate CPU/VRAM trade-off discussed in #14746.

(Replaces #14766, which the issue-link auto-close bot closed on 2026-09-25 and which GitHub refused to reopen - same commit, unchanged content.)

The `get_i2v_mask` method had a hardcoded `device="cuda"` default,
which crashes on non-CUDA accelerators (Ascend NPU, etc.) with
"Torch not compiled with CUDA enabled" when called without an
explicit device argument.

Change the default to None and resolve via `self._execution_device`,
matching the pattern used across other pipeline methods.

Verified on Ascend 910B NPU: torch.zeros(device="cuda") crashes,
fix with device-agnostic resolution creates tensors on the correct
device.
The VAE decode block in the MiniMax H3 pipeline only enables fp16
autocast on CUDA (`enabled=device.type == "cuda"`), which disables
autocast on other accelerators such as Ascend NPU, causing
suboptimal performance and potential dtype mismatches.

Change to `enabled=device.type != "cpu"` so autocast is enabled on
any accelerator device while remaining disabled on CPU.

Verified on Ascend 910B NPU: torch.autocast(device_type="npu",
dtype=torch.float16, enabled=True) correctly computes in fp16.
@li-lizhe

Copy link
Copy Markdown
Author

Note: like #14785/#14786, this PR's CI runs will sit in action_required until a maintainer approves them once for this contributor (fork PR workflow approval).

@li-lizhe

Copy link
Copy Markdown
Author

Closing in favour of #14883.

Two things changed since this was opened:

  1. src/diffusers/modular_pipelines/minimax_h3/decoders.py — @asomoza's [MInimax H3] Fix VAE decode #14754 (merged as 51a454b) already resolved this by dropping the fp16 autocast around the VAE decode (and narrowing decoder in _keep_in_fp32_modules), with an accuracy table showing that the previous "fp32 weights + fp16 autocast" was the wrong combination. My change here went the other way (enabled=device.type != "cpu"), so re-applying it would re-introduce exactly the autocast [MInimax H3] Fix VAE decode #14754 removed. I'm closing MiniMax-H3 video decode only enables fp16 autocast on CUDA, so the block silently runs in fp32 on other accelerators #14882 along with this PR for that reason.

  2. src/diffusers/pipelines/wan/pipeline_wan_animate.py — my commit 5fbfe8c ("fix(wan): use device-agnostic default for get_i2v_mask") is the same commit object that fix(wan): use device-agnostic default for get_i2v_mask #14883 carries, and fix(wan): use device-agnostic default for get_i2v_mask #14883 also covers modular_pipelines/wan_animate_2/encoders.py. So the wan fix (WanAnimatePipeline.get_i2v_mask() defaults device to "cuda", which raises on non-CUDA accelerators (NPU/XPU/MPS) #14881) lives in fix(wan): use device-agnostic default for get_i2v_mask #14883; keeping it here as well would only be a duplicate.

Thanks!

@li-lizhe li-lizhe closed this Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

MiniMax-H3 video decode only enables fp16 autocast on CUDA, so the block silently runs in fp32 on other accelerators

1 participant