Skip to content

Segmented Checkpointing

Segmented checkpointing is the middle gear between checkpointing every block and checkpointing nothing.

It uses PyTorch's activation checkpointing backend. SimpleTuner runs a contiguous group of transformer blocks under one checkpoint call, then carries the returned hidden state into the next group. Wider groups save fewer boundary states but still recompute the checkpointed blocks during backward; measure the resulting speed and peak-memory tradeoff.

For CPU offload and FFN-only checkpointing, use Unsloth-style checkpointing. The short rule of thumb still lives in Decision Rule.

Controls

{
  "gradient_checkpointing": true,
  "gradient_checkpointing_backend": "torch",
  "gradient_checkpointing_interval": 2
}

On supported whole-block paths, gradient_checkpointing_interval is the segment width. 2 means checkpoint blocks 0-1, 2-3, 4-5, and so on.

For finer VRAM control, add a stride:

{
  "gradient_checkpointing": true,
  "gradient_checkpointing_backend": "torch",
  "gradient_checkpointing_interval": 2,
  "gradient_checkpointing_segment_stride": 4
}

That checkpoints blocks 0-1, runs 2-3 normally, checkpoints 4-5, runs 6-7 normally, and repeats. The stride must be at least the interval; overlapping schedules are not valid.

Supported segmented whole-block paths: Flux.1, Flux.2, HunyuanVideo, Krea 2, LongCat Image, LongCat Video, LTXVideo 0.9, LTXVideo2, Lumina2, MageFlow, MiniMax H3, PixArt, Qwen Image 2.1, SD3, SanaVideo, Z-Image, ZLab I1, and Wan.

Stable Cascade stage C also supports interval and stride control, but it applies the schedule to the UNet Res/Timestep/Attention micro-block sequence instead of transformer whole-block groups.

Some model families use model-specific interval semantics:

Family gradient_checkpointing_interval gradient_checkpointing_segment_stride
Sana Checkpoint every N-th block Ignored
Stable Cascade stage C Checkpoint UNet micro-blocks by interval Stride alternates checkpointed and non-checkpointed UNet micro-block windows
SD1x, SDXL No segmented whole-block support Ignored

Do not compare stride rows for families where stride is ignored. If the benchmark numbers look identical there, that is usually the option being ignored, not a useful performance result.

When To Use It

Use it after normal per-block checkpointing fits but costs too much step time. Start with 2. If VRAM allows, try 2 with stride 4 on very deep models.

Do not expect it to help when the peak is mostly trainable weights, optimizer state, validation, VAE caching, block swapping, or routing. SimpleTuner falls back to the safer per-block path when a model feature needs per-block control.

dynamo_use_regional_compilation is not a universal win. It helped or stayed neutral on several image-model runs, but it was a bad fit for the Wan/RamTorch and LTXVideo2 profiles below. Treat compile settings as part of the benchmark, not as background noise.

Benchmarks

Measured with real SimpleTuner examples on single-GPU H100, H200 and L40S pods. Validation and checkpoint saves were disabled, cache preparation was excluded, and first-step compile/setup inside the train loop is excluded when post-warmup timing is available.

Each measured cell is post-warmup sec/step / peak VRAM GiB. Status-only cells mean: OOM ran out of GPU memory, failed did not reach measured training steps, unsupported means that option was not wired for that family, and not run means the sweep did not include that combination.

Compare modes within a family first. Cross-family comparisons are rough because resolution, frame count, attention backend, model depth, trainable adapter type, and dataset shape differ.

The matrix below is the source of truth for this sweep. Model-specific notes call out caveats when a row is coverage data rather than a recommendation.

Family Sweep Results

ACE Step 1.5

Example: ace_step-v1-5.peft-lora. Resolution: 512.

Note: This sweep did not produce a usable ACE Step throughput row. The status-only entries below should be treated as coverage gaps, not as a recommendation.

Precision Mode H100 L40S
bf16 none OOM OOM
bf16 activation-offload unsupported unsupported
bf16 layer OOM OOM
bf16 interval2 OOM OOM
bf16 seg2-stride4 OOM OOM
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none OOM OOM
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer OOM OOM
int8-sdnq-hadamard interval2 OOM OOM
int8-sdnq-hadamard seg2-stride4 OOM OOM
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none OOM OOM
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer OOM OOM
fp8-torchao interval2 OOM OOM
fp8-torchao seg2-stride4 OOM OOM
fp8-torchao seg2-stride4-offload unsupported unsupported

AnyFlow Distillation on Anima

Example: anima-anyflow-stage1.peft-lora. Resolution: 1024x1024. This row measures AnyFlow distillation using Anima, not the plain Anima LoRA example. Use anima.peft-lora for plain 1024x1024 Anima image training.

Precision Mode H100 L40S
bf16 none 1.022 / 17.26 0.719 / 17.21
bf16 activation-offload unsupported unsupported
bf16 layer 1.232 / 5.60 0.903 / 5.56
bf16 interval2 1.252 / 5.60 0.897 / 5.56
bf16 seg2-stride4 1.244 / 5.60 0.898 / 5.56
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 5.417 / 18.61 4.974 / 18.57
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 4.562 / 4.36 4.019 / 4.31
int8-sdnq-hadamard interval2 3.723 / 4.36 3.196 / 4.31
int8-sdnq-hadamard seg2-stride4 3.658 / 4.36 3.140 / 4.31
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none 2.242 / 45.71 OOM
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 2.810 / 5.51 2.576 / 5.46
fp8-torchao interval2 2.846 / 5.51 2.581 / 5.46
fp8-torchao seg2-stride4 2.766 / 5.51 2.567 / 5.46
fp8-torchao seg2-stride4-offload unsupported unsupported

AuraFlow

Example: auraflow.peft-lora. Resolution: 1024x1024.

Note: AuraFlow supports SDNQ and TorchAO quantization. Quantized none rows are included below; quantized checkpoint rows need fresh full-length benchmark coverage before they get numbers here.

Precision Mode H100 L40S
bf16 none 0.180 / 19.19 0.233 / 19.12
bf16 activation-offload unsupported unsupported
bf16 layer 1.824 / 13.37 0.833 / 13.32
bf16 interval2 1.764 / 16.21 0.877 / 16.14
bf16 seg2-stride4 1.771 / 16.21 0.887 / 16.14
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 0.642 / 12.88 0.610 / 12.87
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer not run not run
int8-sdnq-hadamard interval2 not run not run
int8-sdnq-hadamard seg2-stride4 not run not run
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none 0.621 / 24.53 0.757 / 24.44
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer not run not run
fp8-torchao interval2 not run not run
fp8-torchao seg2-stride4 not run not run
fp8-torchao seg2-stride4-offload unsupported unsupported

Boogu Image

Example: boogu-image-v0.1.peft-lora. Resolution: 1024x1024.

Precision Mode H100 L40S
bf16 none 0.694 / 59.14 OOM
bf16 activation-offload unsupported unsupported
bf16 layer 0.907 / 23.44 2.641 / 23.39
bf16 interval2 0.912 / 23.44 2.649 / 23.39
bf16 seg2-stride4 0.911 / 23.44 2.648 / 23.39
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 1.488 / 53.24 OOM
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 1.878 / 15.20 3.309 / 15.15
int8-sdnq-hadamard interval2 1.630 / 34.12 2.577 / 34.07
int8-sdnq-hadamard seg2-stride4 1.656 / 34.11 2.574 / 34.06
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none OOM OOM
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 1.713 / 18.48 3.778 / 18.44
fp8-torchao interval2 1.731 / 18.48 3.777 / 18.44
fp8-torchao seg2-stride4 1.721 / 18.48 3.779 / 18.44
fp8-torchao seg2-stride4-offload unsupported unsupported

Chroma

Example: chroma.peft-lora. Resolution: 1024x1024.

Note: Checkpointed Chroma rows use attention_mechanism=native-efficient, which was the stable attention path for this sweep.

Precision Mode H100 L40S
bf16 none 0.454 / 26.18 0.559 / 26.13
bf16 activation-offload 4.873 / 18.74 4.793 / 18.69
bf16 layer 1.276 / 17.67 1.430 / 17.63
bf16 interval2 1.204 / 21.80 1.349 / 21.75
bf16 seg2-stride4 1.200 / 21.78 1.382 / 21.74
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 1.083 / 21.42 1.061 / 21.37
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 1.714 / 10.41 1.646 / 10.36
int8-sdnq-hadamard interval2 1.443 / 15.72 1.391 / 15.68
int8-sdnq-hadamard seg2-stride4 1.428 / 15.71 1.323 / 15.67
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none 1.122 / 45.44 OOM
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 1.821 / 10.64 2.104 / 10.60
fp8-torchao interval2 1.447 / 27.49 1.871 / 27.44
fp8-torchao seg2-stride4 1.431 / 27.49 1.877 / 27.44
fp8-torchao seg2-stride4-offload unsupported unsupported

Cosmos 2 Image

Example: cosmos2image.lycoris-lokr. Resolution: 512x512.

Precision Mode H100 L40S
bf16 none 0.336 / 8.05 0.316 / 8.00
bf16 activation-offload unsupported unsupported
bf16 layer 0.567 / 4.09 0.559 / 4.04
bf16 interval2 0.595 / 4.09 0.544 / 4.04
bf16 seg2-stride4 0.598 / 4.09 0.546 / 4.04
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 0.831 / 6.56 0.783 / 6.56
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 1.315 / 2.44 1.288 / 2.39
int8-sdnq-hadamard interval2 1.346 / 2.44 1.321 / 2.39
int8-sdnq-hadamard seg2-stride4 1.413 / 2.44 1.274 / 2.39
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none 0.850 / 13.30 0.840 / 13.25
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 1.692 / 2.70 1.560 / 2.65
fp8-torchao interval2 1.626 / 2.70 1.544 / 2.65
fp8-torchao seg2-stride4 1.607 / 2.70 1.555 / 2.65
fp8-torchao seg2-stride4-offload unsupported unsupported

Cosmos 3

Example: cosmos3-edge-image-24g.lycoris-lokr. Resolution: 1024 px.

Precision Mode H100 L40S
bf16 none 2.747 / 8.90 2.904 / 8.86
bf16 activation-offload unsupported unsupported
bf16 layer 2.602 / 8.90 2.606 / 8.86
bf16 interval2 2.658 / 8.90 2.567 / 8.86
bf16 seg2-stride4 2.628 / 8.90 2.899 / 8.86
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 17.112 / 6.21 17.965 / 6.16
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 3.238 / 6.21 2.891 / 6.16
int8-sdnq-hadamard interval2 3.253 / 6.21 2.916 / 6.16
int8-sdnq-hadamard seg2-stride4 3.254 / 6.21 2.923 / 6.16
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none 1.849 / 17.84 1.559 / 17.80
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 1.869 / 17.84 1.599 / 17.79
fp8-torchao interval2 1.855 / 17.84 1.600 / 17.80
fp8-torchao seg2-stride4 1.859 / 17.84 1.523 / 17.80
fp8-torchao seg2-stride4-offload unsupported unsupported

ERNIE 4.5 Image

Example: ernie.peft-lora. Resolution: 512x512.

Precision Mode H100 L40S
bf16 none 0.711 / 15.61 1.282 / 15.56
bf16 activation-offload unsupported unsupported
bf16 layer 1.110 / 4.94 1.840 / 4.90
bf16 interval2 0.876 / 10.16 1.536 / 10.12
bf16 seg2-stride4 0.874 / 10.16 1.532 / 10.12
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 2.782 / 13.75 2.722 / 13.70
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 4.262 / 2.98 4.393 / 2.94
int8-sdnq-hadamard interval2 3.547 / 8.26 3.380 / 8.21
int8-sdnq-hadamard seg2-stride4 3.366 / 8.26 3.457 / 8.21
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none 2.432 / 14.03 2.303 / 13.98
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 3.920 / 3.26 4.008 / 3.22
fp8-torchao interval2 3.316 / 8.54 2.973 / 8.49
fp8-torchao seg2-stride4 2.994 / 8.54 3.046 / 8.49
fp8-torchao seg2-stride4-offload unsupported unsupported

HeartMula

Example: heartmula.peft-lora. Audio-token training.

Precision Mode H100 L40S
bf16 none failed failed
bf16 activation-offload unsupported unsupported
bf16 layer failed failed
bf16 interval2 unsupported unsupported
bf16 seg2-stride4 unsupported unsupported
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none OOM OOM
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer failed failed
int8-sdnq-hadamard interval2 unsupported unsupported
int8-sdnq-hadamard seg2-stride4 unsupported unsupported
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none failed failed
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer failed failed
fp8-torchao interval2 unsupported unsupported
fp8-torchao seg2-stride4 unsupported unsupported
fp8-torchao seg2-stride4-offload unsupported unsupported

HiDream

Example: hidream.peft-lora. Resolution: 512x512.

Precision Mode H100 L40S
bf16 none 0.515 / 44.58 OOM
bf16 activation-offload unsupported unsupported
bf16 layer 0.928 / 33.57 1.043 / 33.52
bf16 interval2 0.971 / 33.57 1.002 / 33.52
bf16 seg2-stride4 0.952 / 33.57 1.010 / 33.52
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none not measured not measured
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 2.497 / 17.99 2.058 / 17.93
int8-sdnq-hadamard interval2 unsupported unsupported
int8-sdnq-hadamard seg2-stride4 unsupported unsupported
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none not measured not measured
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 4.920 / 18.10 4.338 / 18.05
fp8-torchao interval2 unsupported unsupported
fp8-torchao seg2-stride4 unsupported unsupported
fp8-torchao seg2-stride4-offload unsupported unsupported

SDNQ Hadamard numbers use sdnq_compile_mode=eager. The compiled SDNQ path quantizes HiDream, but this sweep spent the first training step in Inductor dequantizer compilation, so it is not listed as a throughput row.

HunyuanVideo

Example: hunyuanvideo-1.5-t2v.peft-lora. Training shape: 480 pixel-area video buckets, 48 frames, batch 2.

Precision Mode H100 L40S
bf16 none OOM OOM
bf16 activation-offload not run not run
bf16 layer 7.682 / 26.35 22.816 / 26.30
bf16 interval2 7.398 / 26.11 22.772 / 26.06
bf16 seg2-stride4 OOM OOM
bf16 seg2-stride4-offload 10.679 / 58.37 not run
int8-sdnq-hadamard none not run not run
int8-sdnq-hadamard activation-offload not run not run
int8-sdnq-hadamard layer 11.765 / 25.96 34.464 / 25.92
int8-sdnq-hadamard interval2 not run not run
int8-sdnq-hadamard seg2-stride4 not run not run
int8-sdnq-hadamard seg2-stride4-offload not run not run
fp8-torchao none not run not run
fp8-torchao activation-offload not run not run
fp8-torchao layer 10.516 / 33.55 32.003 / 33.53
fp8-torchao interval2 not run not run
fp8-torchao seg2-stride4 not run not run
fp8-torchao seg2-stride4-offload not run not run

HunyuanVideo is activation-heavy at this training shape. Per-block and interval-2 checkpointing both fit cleanly; leaving checkpointing off does not fit on an 80 GB H100. seg2-stride4 only fit in this sweep when attention activation offload was enabled, and that row is a fallback rather than a speed recommendation. SDNQ Hadamard works, but variable conditioning shapes still trigger dynamic-kernel compilation in the measured window.

Ideogram 4.0

Example: ideogram-fp8.peft-lora. Resolution: 1024x1024. The fp8 flavour uses Ideogram 4's native weight-only fp8 checkpoint (base_model_precision=no_change); bf16-upcast sets ideogram_fp8_base_upcast=true to dequantize the base weights to bf16 at load. Requesting any other base_model_precision (e.g. int8-sdnq) also dequantizes first so the quantizer operates on real weights.

Earlier revisions of this table were measured before Ideogram honored gradient_checkpointing_interval, gradient_checkpointing_segment_stride, or gradient_checkpointing_backend (its custom loader skipped the shared wiring), and before int8-sdnq actually quantized the fp8 checkpoint — all checkpointing rows were silently full-layer torch checkpointing over fp8-native weights. The H100 numbers below are post-fix; L40S rows await re-measurement.

The torch-ffn/unsloth-ffn backends are unsupported: Ideogram 4 does not expose an attention/FFN checkpointing boundary.

Precision Mode Backend H100 speed (s/step) H100 VRAM (GiB)
fp8-native layer torch 1.068 12.33
fp8-native layer unsloth 1.098 11.17
fp8-native seg2-stride4 torch 0.915 36.28
fp8-native seg2-stride4 unsloth 0.929 35.70
bf16-upcast layer torch 0.962 20.51
bf16-upcast layer unsloth 0.992 19.35
bf16-upcast seg2-stride4 torch 0.842 36.85
bf16-upcast seg2-stride4 unsloth 0.848 36.27
int8-sdnq-hadamard layer torch 1.735 11.88
int8-sdnq-hadamard layer unsloth 1.535 10.72
int8-sdnq-hadamard seg2-stride4 torch 0.981 28.19
int8-sdnq-hadamard seg2-stride4 unsloth 0.992 27.61

Takeaways: seg2-stride4 is ~14% faster than full-layer at ~24 GiB more retained activations; unsloth offload saves ~1.2 GiB for a 1-3% step-time cost (and is faster than torch under int8 full-layer, where offload overlaps the quantized matmul); bf16-upcast is the throughput winner when VRAM allows.

Kandinsky 5 Image

Example: kandinsky5-image-6b-t2i.lycoris-lokr. Resolution: 1024x1024.

Note: batch 3 at 1024x1024 needs full checkpointing on both cards. SDNQ with Hadamard is the best low-VRAM row; H100 can also use partial checkpointing with SDNQ, but only near the top of the card.

Precision Mode H100 L40S
bf16 none OOM OOM
bf16 activation-offload 6.956 / 25.52 10.189 / 25.42
bf16 layer 6.590 / 25.58 9.458 / 25.55
bf16 interval2 OOM OOM
bf16 seg2-stride4 OOM OOM
bf16 seg2-stride4-offload OOM OOM
int8-sdnq-hadamard none OOM OOM
int8-sdnq-hadamard activation-offload 7.186 / 20.00 10.141 / 19.95
int8-sdnq-hadamard layer 6.830 / 20.12 9.362 / 20.08
int8-sdnq-hadamard interval2 5.746 / 75.67 OOM
int8-sdnq-hadamard seg2-stride4 OOM OOM
int8-sdnq-hadamard seg2-stride4-offload 6.057 / 71.46 OOM
fp8-torchao none OOM OOM
fp8-torchao activation-offload 8.319 / 24.29 14.716 / 24.24
fp8-torchao layer 7.976 / 24.40 13.949 / 24.35
fp8-torchao interval2 OOM OOM
fp8-torchao seg2-stride4 OOM OOM
fp8-torchao seg2-stride4-offload OOM OOM

Kandinsky 5 Video

Example: kandinsky5-video-2b-t2v.peft-lora. Resolution: 768x512, 81f.

Kandinsky 5 video is activation-heavy at this frame count. Full block checkpointing is the practical baseline on both cards. On H100, interval2 and seg2-stride4 are faster when they fit; on L40S, SDNQ interval2 is the only partial-checkpoint row here that fits without attention activation offload.

Precision Mode H100 L40S
bf16 none OOM OOM
bf16 activation-offload 2.580 / 9.49 7.136 / 9.45
bf16 layer 2.267 / 9.83 6.379 / 9.79
bf16 interval2 1.967 / 44.57 OOM
bf16 seg2-stride4 1.971 / 46.62 OOM
bf16 seg2-stride4-offload 2.275 / 37.99 6.249 / 37.94
int8-sdnq-hadamard none OOM OOM
int8-sdnq-hadamard activation-offload 2.844 / 8.07 7.234 / 8.02
int8-sdnq-hadamard layer 2.460 / 8.40 6.509 / 8.35
int8-sdnq-hadamard interval2 2.126 / 43.12 5.641 / 43.08
int8-sdnq-hadamard seg2-stride4 2.125 / 45.19 OOM
int8-sdnq-hadamard seg2-stride4-offload 2.451 / 36.56 6.322 / 36.51
fp8-torchao none OOM OOM
fp8-torchao activation-offload 3.867 / 15.11 11.501 / 15.06
fp8-torchao layer 3.579 / 15.28 10.822 / 15.24
fp8-torchao interval2 OOM OOM
fp8-torchao seg2-stride4 OOM OOM
fp8-torchao seg2-stride4-offload OOM OOM

Kolors

Example: kolors.peft-lora. Resolution: 1024x1024.

Precision Mode H100 L40S
bf16 none 0.635 / 7.22 0.628 / 7.17
bf16 activation-offload unsupported unsupported
bf16 layer 1.118 / 5.36 1.065 / 5.31
bf16 interval2 1.105 / 5.36 1.068 / 5.31
bf16 seg2-stride4 1.110 / 5.36 1.072 / 5.31
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 1.860 / 5.68 1.726 / 5.63
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 2.967 / 3.38 2.803 / 3.33
int8-sdnq-hadamard interval2 2.770 / 3.32 2.655 / 3.27
int8-sdnq-hadamard seg2-stride4 2.805 / 3.32 2.742 / 3.27
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none 1.629 / 8.79 1.637 / 8.75
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 3.044 / 3.50 3.013 / 3.45
fp8-torchao interval2 3.040 / 3.50 3.019 / 3.45
fp8-torchao seg2-stride4 2.988 / 3.50 2.935 / 3.45
fp8-torchao seg2-stride4-offload unsupported unsupported

Krea 2

Example: krea2.peft-lora. Training resolution: 512 px square crop. Example validation setting: 1024x1024; validation was disabled for the benchmark.

The main table uses regional compilation, which is good for Krea2 step speed but not a clean VRAM comparison: the compiled graph/workspace keeps the peak close to the uncheckpointed peak for several checkpoint modes. A bf16 control run with regional compilation disabled showed checkpointing is wired and has the expected memory/speed shape:

The activation-offload row here means full-block checkpointing plus attention activation offload. Against full-block layer checkpointing alone, attention offload did not reduce Krea2 peak VRAM in this shape; it mostly added CPU transfer overhead.

Mode H100 no-compile L40S no-compile
none 0.272 / 40.09 0.661 / 40.01
layer 0.371 / 30.06 0.919 / 30.01
seg2-stride4 0.317 / 34.75 0.788 / 34.70
activation-offload 0.657 / 30.50 1.341 / 30.30
Precision Mode H100 L40S
bf16 none 0.275 / 40.06 0.662 / 40.01
bf16 activation-offload 0.416 / 34.62 0.822 / 34.57
bf16 layer 0.268 / 40.06 0.665 / 40.01
bf16 interval2 0.274 / 40.06 0.663 / 40.01
bf16 seg2-stride4 0.279 / 40.06 0.663 / 40.01
bf16 seg2-stride4-offload 0.404 / 34.62 0.819 / 34.57
int8-sdnq-hadamard none 0.462 / 27.01 0.807 / 26.96
int8-sdnq-hadamard activation-offload 0.773 / 21.57 1.012 / 21.53
int8-sdnq-hadamard layer 0.473 / 27.01 0.802 / 26.96
int8-sdnq-hadamard interval2 0.474 / 27.01 0.804 / 26.96
int8-sdnq-hadamard seg2-stride4 0.472 / 27.01 0.802 / 26.96
int8-sdnq-hadamard seg2-stride4-offload 0.744 / 21.57 1.007 / 21.53
fp8-torchao none 0.689 / 51.63 OOM
fp8-torchao activation-offload 0.975 / 37.77 2.058 / 37.73
fp8-torchao layer 0.689 / 51.63 OOM
fp8-torchao interval2 0.684 / 51.63 OOM
fp8-torchao seg2-stride4 0.674 / 51.63 OOM
fp8-torchao seg2-stride4-offload 0.965 / 37.77 2.053 / 37.73

LongCat Image

Example: longcat-image.peft-lora. Training resolution: 512 px square; validation resolution: 1024x1024. Rows use attention_mechanism=native-flash.

Precision Mode H100 L40S
bf16 none 0.193 / 16.47 0.262 / 16.42
bf16 activation-offload 0.544 / 12.73 0.543 / 12.69
bf16 layer 0.327 / 12.38 0.370 / 12.34
bf16 interval2 0.257 / 14.38 0.313 / 14.34
bf16 seg2-stride4 0.263 / 14.36 0.316 / 14.31
bf16 seg2-stride4-offload 0.446 / 13.42 0.492 / 13.38
int8-sdnq-hadamard none 0.578 / 12.54 0.537 / 12.45
int8-sdnq-hadamard activation-offload 1.184 / 7.43 1.185 / 7.39
int8-sdnq-hadamard layer 0.901 / 7.19 0.911 / 7.09
int8-sdnq-hadamard interval2 0.718 / 9.73 0.695 / 9.68
int8-sdnq-hadamard seg2-stride4 0.735 / 9.72 0.717 / 9.68
int8-sdnq-hadamard seg2-stride4-offload 1.056 / 8.49 1.011 / 8.45
fp8-torchao none 0.602 / 25.19 0.834 / 25.14
fp8-torchao activation-offload 1.662 / 7.75 1.844 / 7.70
fp8-torchao layer 0.984 / 7.57 1.080 / 7.53
fp8-torchao interval2 0.750 / 16.15 0.938 / 16.10
fp8-torchao seg2-stride4 0.760 / 16.13 0.961 / 16.09
fp8-torchao seg2-stride4-offload 1.287 / 13.02 1.653 / 12.98

LongCat Video

Example: longcat-video.peft-lora+ramtorch. Resolution: 832x480, 81f. Rows use attention_mechanism=native-flash.

LongCat Video is activation-heavy at this shape. Full per-block checkpointing is the practical row. The partial checkpoint rows (interval2, seg2-stride4) do not fit here, even when attention activation offload is enabled for the strided row. Plain attention activation offload fits for bf16 and SDNQ, but it is much slower than full checkpointing.

Precision Mode H100 L40S
bf16 none OOM (77.86 GiB) OOM (43.30 GiB)
bf16 activation-offload 25.774 / 37.36 49.149 / 37.14
bf16 layer 7.448 / 23.73 24.866 / 23.68
bf16 interval2 OOM (76.41 GiB) OOM (42.80 GiB)
bf16 seg2-stride4 OOM (76.42 GiB) OOM (43.06 GiB)
bf16 seg2-stride4-offload OOM (76.72 GiB) OOM (42.29 GiB)
int8-sdnq-hadamard none OOM (77.40 GiB) OOM (43.59 GiB)
int8-sdnq-hadamard activation-offload 30.887 / 35.28 61.270 / 35.24
int8-sdnq-hadamard layer 8.444 / 21.60 25.164 / 21.55
int8-sdnq-hadamard interval2 OOM (76.01 GiB) OOM (42.47 GiB)
int8-sdnq-hadamard seg2-stride4 OOM (77.01 GiB) OOM (43.02 GiB)
int8-sdnq-hadamard seg2-stride4-offload OOM (76.42 GiB) OOM (42.57 GiB)
fp8-torchao none OOM (75.87 GiB) OOM (41.28 GiB)
fp8-torchao activation-offload 30.163 / 47.88 OOM (40.11 GiB)
fp8-torchao layer 8.343 / 34.16 24.659 / 34.07
fp8-torchao interval2 OOM (74.63 GiB) OOM (40.85 GiB)
fp8-torchao seg2-stride4 OOM (75.37 GiB) OOM (41.19 GiB)
fp8-torchao seg2-stride4-offload OOM (75.55 GiB) OOM (40.94 GiB)

LTXVideo 0.9.5

Example: ltxvideo-0.9.5-t2v.peft-lora. Resolution: 768x512, 49f.

Numbers are warm seconds per step / peak GiB. The full-run average includes setup and compile overhead and is recorded in the sweep artifacts.

Precision Mode H100 L40S
bf16 none 0.274 / 8.77 0.275 / 8.72
bf16 layer 0.462 / 4.21 0.459 / 4.16
bf16 interval2 0.449 / 4.30 0.433 / 4.25
bf16 seg2-stride4 0.359 / 6.40 0.357 / 6.35
int8-sdnq-hadamard none 0.688 / 7.05 0.640 / 6.90
int8-sdnq-hadamard layer 1.094 / 2.59 1.081 / 2.44
int8-sdnq-hadamard interval2 1.112 / 2.57 1.073 / 2.53
int8-sdnq-hadamard seg2-stride4 0.887 / 4.61 0.817 / 4.56
fp8-torchao none 0.655 / 17.45 0.735 / 17.41
fp8-torchao layer 1.226 / 2.94 1.206 / 2.89
fp8-torchao interval2 1.259 / 3.40 1.233 / 3.35
fp8-torchao seg2-stride4 0.933 / 9.96 0.901 / 9.91
fp8wo-torchao none 0.328 / 10.06 0.325 / 10.01
fp8wo-torchao layer 0.567 / 2.64 0.540 / 2.59
fp8wo-torchao interval2 0.555 / 2.84 0.531 / 2.79
fp8wo-torchao seg2-stride4 0.443 / 6.20 0.432 / 6.15

Attention activation offload rows are unsupported for LTXVideo 0.9 in this sweep.

LTXVideo2 2.3

Example: ltxvideo2-2.3-dev-720p-single-gpu.peft-lora+sdnq-hadamard. Resolution: 1280x704, 49f.

Note: LTXVideo2 2.3 should be read from the no-regional-compile rows in this sweep; regional compile raised memory pressure for this model.

Precision Mode H100 L40S
bf16 none OOM OOM
bf16 activation-offload 6.350 / 58.36 OOM
bf16 layer 3.993 / 47.95 OOM
bf16 interval2 3.977 / 48.83 OOM
bf16 seg2-stride4 OOM OOM
bf16 seg2-stride4-offload 5.554 / 75.64 OOM
int8-sdnq-hadamard none OOM OOM
int8-sdnq-hadamard activation-offload 10.102 / 38.20 9.852 / 38.15
int8-sdnq-hadamard layer 7.753 / 27.78 7.579 / 27.73
int8-sdnq-hadamard interval2 7.733 / 28.66 7.288 / 28.61
int8-sdnq-hadamard seg2-stride4 6.522 / 61.50 OOM
int8-sdnq-hadamard seg2-stride4-offload 8.659 / 55.48 OOM
fp8-torchao none OOM OOM
fp8-torchao activation-offload failed 22.917 / 38.57
fp8-torchao layer 8.580 / 30.27 10.381 / 30.22
fp8-torchao interval2 8.660 / 33.71 10.661 / 33.66
fp8-torchao seg2-stride4 OOM OOM
fp8-torchao seg2-stride4-offload failed OOM

Lumina2

Example: lumina2.peft-lora. Resolution: 512x512.

Note: Lumina2 now uses the segmented whole-block path. interval2 checkpoints every two-block segment; seg2-stride4 checkpoints two blocks, lets the next two keep activations, then repeats. Attention activation offload was not part of this Lumina2 run.

Precision Mode H100 L40S
bf16 none 0.235 / 15.99 0.332 / 15.94
bf16 layer 0.384 / 6.60 0.457 / 6.56
bf16 interval2 0.356 / 6.87 0.424 / 6.82
bf16 seg2-stride4 0.295 / 11.43 0.377 / 11.38
int8-sdnq-hadamard none 0.584 / 13.59 0.541 / 13.55
int8-sdnq-hadamard layer 0.899 / 4.21 0.827 / 4.16
int8-sdnq-hadamard interval2 0.865 / 4.48 0.835 / 4.43
int8-sdnq-hadamard seg2-stride4 0.719 / 9.03 0.707 / 8.99
fp8-torchao none 0.598 / 28.03 0.763 / 27.98
fp8-torchao layer 0.950 / 5.93 0.967 / 5.88
fp8-torchao interval2 0.938 / 6.70 0.974 / 6.66
fp8-torchao seg2-stride4 0.765 / 17.36 0.901 / 17.32
fp8wo-torchao none 0.273 / 17.97 0.389 / 17.93
fp8wo-torchao layer 0.452 / 4.99 0.525 / 4.94
fp8wo-torchao interval2 0.427 / 5.40 0.522 / 5.35
fp8wo-torchao seg2-stride4 0.360 / 11.68 0.466 / 11.64

MageFlow

Example: mageflow-image-24g.peft-lora. Resolution: 1024x1024.

Note: MageFlow's 1024px variable-shape image path is mostly helped by attention activation offload and weight-only FP8. Block checkpointing modes are valid, but did not lower measured peak residency in this sweep.

Precision Mode H100 L40S
bf16 none 7.735 / 36.74 7.658 / 36.69
bf16 activation-offload 8.902 / 23.05 9.477 / 23.00
bf16 layer 7.805 / 36.74 7.504 / 36.69
bf16 interval2 8.036 / 36.74 7.762 / 36.69
bf16 seg2-stride4 7.882 / 36.74 7.531 / 36.69
bf16 seg2-stride4-offload 8.833 / 23.05 9.288 / 23.01
int8-sdnq-hadamard none 81.016 / 37.38 94.991 / 37.34
fp8wo-torchao none 5.454 / 36.86 5.772 / 36.82
fp8wo-torchao activation-offload 6.295 / 23.18 6.738 / 23.14
fp8wo-torchao seg2-stride4 5.542 / 36.86 5.595 / 36.82

OmniGen

Example: omnigen.lycoris-lokr. Resolution: 1024x1024.

Note: OmniGen uses token-ID prompts instead of cached text embeddings. These rows measure the supported no-checkpoint and full-block torch checkpointing paths; interval, segmented-stride, and attention-offload controls are not implemented for this family in this sweep.

Precision Mode H100 L40S
bf16 none 0.425 / 14.20 0.293 / 14.24
bf16 layer 0.597 / 10.13 0.389 / 10.09
int8-sdnq-hadamard none 1.523 / 11.00 1.388 / 11.06
int8-sdnq-hadamard layer 1.312 / 6.75 1.064 / 6.70
fp8-torchao none 0.690 / 19.25 0.608 / 19.30
fp8-torchao layer 1.069 / 7.09 0.824 / 7.04
fp8wo-torchao none 0.454 / 17.71 0.377 / 17.73
fp8wo-torchao layer 0.646 / 6.98 0.534 / 6.94

PixArt

Example: pixart.lycoris-lokr. Resolution: 1024x1024.

Precision Mode H100 L40S
bf16 none 1.700 / 41.73 1.734 / 41.67
bf16 activation-offload unsupported unsupported
bf16 layer 2.433 / 5.63 2.346 / 5.58
bf16 interval2 2.440 / 6.01 2.348 / 5.96
bf16 seg2-stride4 2.092 / 23.90 2.072 / 23.86
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 1.905 / 47.58 OOM
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 2.738 / 6.07 2.937 / 6.03
int8-sdnq-hadamard interval2 2.734 / 6.03 2.943 / 5.99
int8-sdnq-hadamard seg2-stride4 2.336 / 26.85 2.596 / 26.81
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none OOM OOM
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 3.254 / 8.06 4.646 / 8.01
fp8-torchao interval2 3.260 / 9.66 4.649 / 9.61
fp8-torchao seg2-stride4 2.827 / 63.27 OOM
fp8-torchao seg2-stride4-offload unsupported unsupported

Qwen Image 2.1

Measured on verified H200 and L40S GPUs with qwen_image.peft-lora, Qwen-Image 2.1, Domokun at 512x512, BF16, rank-32 LoRA, Optimi Lion and regional Inductor compilation (default mode). Each run has 20 steps; timings exclude the first five. Cells are seconds/step / peak VRAM GiB; peaks include component setup.

Runtime: torch==2.11.0+cu128, diffusers==0.40.0; 10 text tokens and 1024 image tokens. Component setup dominates some batch-1 memory peaks.

BF16 checkpoint schedules

none disables checkpointing. layer checkpoints each block. interval2 checkpoints contiguous pairs 0–1, 2–3, …; seg2-stride4 checkpoints 0–1, retains activations for 2–3, and repeats. Grouped execution requires no block swapping, hidden-state capture or KV cache; those paths retain per-block interval/stride scheduling. Attention activation offload (activation-offload, seg2-stride4-offload) is unsupported. BF16 fits, so quantization was not included in this 2.1 sweep.

Mode H200, batch 1 H200, batch 20 L40S, batch 1
none 0.081 / 20.99 1.223 / 128.42 0.238 / 20.64
layer 0.206 / 17.69 1.773 / 24.19 0.368 / 17.60
interval2 0.190 / 17.69 1.761 / 25.14 0.371 / 17.60
seg2-stride4 0.116 / 17.77 1.490 / 73.26 0.290 / 17.60

Attention backends

Attention comparisons disable checkpointing. The default native SDPA selected cuDNN on H200 and PyTorch Flash SDPA on L40S. Qwen 2.1 uses causal text attention and noncausal target-image attention over all valid keys. Equal, unpadded prompt lengths do not require varlen. FlashAttention paths currently require unpadded text-to-image prompts; use native SDPA or FlexAttention for padded prompts or interleaved conditioning. A padding-only varlen mask cannot preserve the latter causal structure.

attention_mechanism H200, batch 1 H200, batch 20 L40S, batch 1
native 0.081 / 20.99 1.223 / 128.42 0.238 / 20.64
native-flash 0.080 / 20.98 1.278 / 128.40 0.238 / 20.63
native-efficient not run not run 0.259 / 20.75
flash-attn-hub not run not run 0.238 / 20.75
flash-attn-3-hub 0.094 / 20.98 1.215 / 128.40 not run
flash-attn-3-varlen-hub 0.096 / 20.97 1.216 / 128.44 not run
flex 0.100 / 20.92 1.314 / 134.45 0.271 / 20.70
flash-attn-4-hub failed failed not run
cudnn failed not run not run

The uniform-length varlen metadata helper is vendored from upstream Diffusers: lengths come from shapes and cumulative offsets use arange, removing the released helper’s GPU .item() synchronizations and graph breaks. Redundant all-valid masks are removed during collation. Forward/gradient equivalence and graph capture are checked separately from timing.

Forcing attention_mechanism: "cudnn" failed in PyTorch 2.11 Inductor with CantSplit; automatic native SDPA still used cuDNN successfully. FA4 Hub failed to load with nvidia-cutlass-dsl==4.7.1: cutlass.cute.core.ThrMma is missing, so no timing is available. Hub backends use trust_remote_code: true. FA3 is a Hopper backend; it is not an L40S option (upstream support).

H200 batch scaling

On the same H200, native SDPA took 0.612 s at batch 10 and 1.223 s at batch 20, or about 16.35 images/s in both cases. Doubling the batch doubled the work after throughput had plateaued. The 144 GB preset keeps checkpointing disabled; its larger batch uses the extra capacity without promising higher images/s.

Batch Seconds/step Images/s Peak GiB
1 0.081 12.38 20.99
10 0.612 16.35 71.85
20 1.223 16.35 128.42
H100 with the same updated native attention path took 0.639 s/step at batch 10, with 71.83 GiB peak VRAM and checkpointing disabled.

A full-transformer check on cached Domokun inputs found identical FP32 outputs with and without the redundant mask; relative gradient L2 error was 1.97e-6. BF16 backends are not numerically interchangeable: on this single batch, aggregate LoRA-gradient cosine against FP32 was about 0.92–0.93 for native SDPA and 0.80 for FA3. These are numerical diagnostics, not convergence comparisons.

The H200 batch-20 trace had kernels active for 93.1% of the captured interval. GEMMs accounted for 70.5% of kernel time and attention for 9.7%; Inductor already fused normalization/RoPE and pointwise epilogues. The largest idle gaps were around batch preparation. Kernel active time is not SM occupancy, and profiler timings are excluded from the throughput tables.

Prefetching was slower for this 27-image dataset at batch 20: with queue length 2, host prefetch averaged 2.030 s/step and device prefetch (1 MiB threshold) 2.036 s/step, versus 1.223 s without prefetch. Both had periodic roughly 3-second steps despite medians near 1.17 s; the presets leave prefetch disabled.

The 250-step Domokun recipe is a throughput example, not a reliable convergence recipe. An earlier checkpoint produced recognizable Domokun images after reloading, but fresh 250-step runs did not reproduce that result. Controls retaining padding masks, disabling compilation, and restoring the earlier RoPE expression also failed. Cached latents decode to the correct subject. The cause of the training deterioration remains unresolved; the timing tables do not establish comparable image quality across attention backends.

Qwen Image

Historical Qwen-Image 1.0 measurements (model_flavour: "v1.0"), using the earlier qwen_image.peft-lora example at 1024x1024. The example now selects 2.1. The nearly identical interval/stride rows below do not establish that distinct checkpoint schedules were active and should not guide 2.1 configuration.

Precision Mode H100 L40S
bf16 none OOM OOM
bf16 activation-offload unsupported unsupported
bf16 layer 1.202 / 41.04 3.377 / 40.99
bf16 interval2 1.201 / 41.04 3.382 / 40.99
bf16 seg2-stride4 1.205 / 41.03 3.385 / 40.99
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 1.675 / 63.48 OOM
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 2.640 / 24.09 3.928 / 24.05
int8-sdnq-hadamard interval2 2.722 / 24.09 3.918 / 24.05
int8-sdnq-hadamard seg2-stride4 2.663 / 24.09 3.919 / 24.05
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none OOM OOM
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 3.088 / 25.34 6.172 / 25.29
fp8-torchao interval2 3.095 / 25.34 6.173 / 25.29
fp8-torchao seg2-stride4 3.125 / 25.34 6.141 / 25.29
fp8-torchao seg2-stride4-offload unsupported unsupported

Sana

Example: sana.lycoris-lokr. Resolution: 1024x1024.

Note: Sana has interval checkpointing; stride is not a separate segmented schedule for this family in the measured rows.

Precision Mode H100 L40S
bf16 none 0.529 / 23.71 0.597 / 23.67
bf16 activation-offload unsupported unsupported
bf16 layer 0.529 / 23.72 0.596 / 23.66
bf16 interval2 0.530 / 23.72 0.598 / 23.66
bf16 seg2-stride4 unsupported unsupported
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 0.554 / 22.92 0.590 / 22.88
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 0.554 / 22.92 0.589 / 22.88
int8-sdnq-hadamard interval2 0.556 / 22.92 0.591 / 22.88
int8-sdnq-hadamard seg2-stride4 unsupported unsupported
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none 0.633 / 33.67 0.753 / 33.62
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 0.633 / 33.67 0.755 / 33.62
fp8-torchao interval2 0.631 / 33.67 0.759 / 33.62
fp8-torchao seg2-stride4 unsupported unsupported
fp8-torchao seg2-stride4-offload unsupported unsupported

SanaVideo

Example: sanavideo-2b-480p.peft-lora. Resolution: 832x480, 49f.

Note: SanaVideo uses linear attention, so attention activation offload remains unsupported. Segmented whole-block checkpointing is supported for the standard path.

Precision Mode H100 L40S
bf16 none 0.599 / 59.15 OOM
bf16 activation-offload unsupported unsupported
bf16 layer 0.597 / 59.15 OOM
bf16 interval2 not run OOM
bf16 seg2-stride4 not run OOM
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 0.641 / 58.36 OOM
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 0.641 / 58.36 OOM
int8-sdnq-hadamard interval2 not run OOM
int8-sdnq-hadamard seg2-stride4 not run OOM
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none OOM OOM
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer OOM OOM
fp8-torchao interval2 OOM OOM
fp8-torchao seg2-stride4 OOM OOM
fp8-torchao seg2-stride4-offload unsupported unsupported

SD 1.x

Example: sd1x-dreamshaper.peft-lora. Resolution: 512x512.

Note: SD1x uses the diffusers UNet path. Regular layer checkpointing is supported, but interval and segmented stride controls are not wired for this family.

Precision Mode H100 L40S
bf16 none 0.181 / 2.87 0.176 / 2.83
bf16 activation-offload unsupported unsupported
bf16 layer 0.305 / 1.98 0.292 / 1.93
bf16 interval2 unsupported unsupported
bf16 seg2-stride4 unsupported unsupported
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 0.446 / 3.07 0.431 / 3.04
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 0.701 / 1.79 0.666 / 1.74
int8-sdnq-hadamard interval2 unsupported unsupported
int8-sdnq-hadamard seg2-stride4 unsupported unsupported
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none 0.410 / 4.23 0.401 / 4.18
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 0.785 / 1.96 0.756 / 1.91
fp8-torchao interval2 unsupported unsupported
fp8-torchao seg2-stride4 unsupported unsupported
fp8-torchao seg2-stride4-offload unsupported unsupported

SD3

Example: sd3.peft-lora. Resolution: 1024x1024.

Note: SD3 uses true contiguous segmented checkpointing on the plain transformer path. Attention activation offload is supported; it cuts VRAM hard, but costs throughput.

Precision Mode H100 L40S
bf16 none 0.529 / 34.59 1.189 / 34.53
bf16 activation-offload 1.335 / 9.68 2.850 / 9.63
bf16 layer 0.721 / 7.67 1.607 / 7.62
bf16 interval2 0.723 / 8.93 1.606 / 8.88
bf16 seg2-stride4 0.620 / 20.14 1.398 / 20.09
bf16 seg2-stride4-offload 1.229 / 12.21 2.639 / 12.16
int8-sdnq-hadamard none 0.858 / 33.66 1.381 / 33.61
int8-sdnq-hadamard activation-offload 1.845 / 8.90 3.297 / 8.85
int8-sdnq-hadamard layer 1.265 / 6.58 1.857 / 6.53
int8-sdnq-hadamard interval2 1.264 / 7.30 1.858 / 7.26
int8-sdnq-hadamard seg2-stride4 1.049 / 19.02 1.625 / 18.98
int8-sdnq-hadamard seg2-stride4-offload 1.527 / 12.42 3.023 / 12.37
fp8-torchao none OOM OOM
fp8-torchao activation-offload 4.707 / 10.27 9.814 / 10.23
fp8-torchao layer 1.575 / 9.13 3.279 / 9.09
fp8-torchao interval2 1.576 / 12.62 3.282 / 12.58
fp8-torchao seg2-stride4 1.376 / 45.70 OOM
fp8-torchao seg2-stride4-offload 4.184 / 26.17 8.714 / 26.12
fp8wo-torchao none 0.567 / 35.21 1.252 / 35.17
fp8wo-torchao activation-offload 1.392 / 8.71 2.921 / 8.67
fp8wo-torchao layer 0.795 / 5.69 1.736 / 5.65
fp8wo-torchao interval2 0.797 / 7.08 1.734 / 7.03
fp8wo-torchao seg2-stride4 0.679 / 19.43 1.490 / 19.38
fp8wo-torchao seg2-stride4-offload 1.257 / 12.00 2.676 / 11.96

SDXL

Example: sdxl.lycoris-lokr. Resolution: 1024x1024.

Note: SDXL has real layer checkpointing. Interval and stride rows are included as coverage data, not as segmented-support recommendations.

Precision Mode H100 L40S
bf16 none 0.606 / 13.03 0.585 / 12.98
bf16 activation-offload unsupported unsupported
bf16 layer 1.080 / 6.53 1.029 / 6.48
bf16 interval2 unsupported unsupported
bf16 seg2-stride4 unsupported unsupported
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 1.741 / 13.72 1.643 / 13.68
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 2.820 / 4.64 2.647 / 4.59
int8-sdnq-hadamard interval2 unsupported unsupported
int8-sdnq-hadamard seg2-stride4 unsupported unsupported
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none 1.608 / 26.22 1.582 / 26.16
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer 2.939 / 5.10 2.890 / 5.04
fp8-torchao interval2 unsupported unsupported
fp8-torchao seg2-stride4 unsupported unsupported
fp8-torchao seg2-stride4-offload unsupported unsupported

Stable Cascade

Example: cascade-stage-c.lycoris-lokr. Resolution: 1024x1024.

Note: Stage C is a full-precision prior path. These rows ran with mixed_precision=no and base_model_precision=no_change; quantized base precision rows are not meaningful for this model. The interval and stride modes operate over the UNet's Res/Timestep/Attention micro-block sequence.

Precision Mode H100 L40S
bf16 none 0.884 / 51.52 OOM
bf16 activation-offload unsupported unsupported
bf16 layer 1.179 / 22.68 2.135 / 22.61
bf16 interval2 1.032 / 36.99 1.871 / 36.92
bf16 seg2-stride4 1.032 / 37.20 1.870 / 37.13
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none unsupported unsupported
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer unsupported unsupported
int8-sdnq-hadamard interval2 unsupported unsupported
int8-sdnq-hadamard seg2-stride4 unsupported unsupported
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none unsupported unsupported
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer unsupported unsupported
fp8-torchao interval2 unsupported unsupported
fp8-torchao seg2-stride4 unsupported unsupported
fp8-torchao seg2-stride4-offload unsupported unsupported

Wan 2.1 T2V 1.3B

Example: wan2.1-t2v-1.3b-480p-single-gpu.peft-lora+ramtorch. Resolution: 832x480, 81f.

Note: Wan 1.3B should be read from the no-regional-compile/RamTorch rows; regional compile was not a useful throughput setting in this sweep.

Precision Mode H100 L40S
bf16 none 1.407 / 71.90 OOM
bf16 activation-offload 3.472 / 8.78 7.179 / 8.66
bf16 layer 2.099 / 4.73 4.459 / 4.68
bf16 interval2 2.139 / 6.32 4.514 / 6.27
bf16 seg2-stride4 1.806 / 39.25 3.921 / 39.21
bf16 seg2-stride4-offload 2.993 / 22.66 6.493 / 22.61
int8-sdnq-hadamard none 1.850 / 71.91 OOM
int8-sdnq-hadamard activation-offload 4.204 / 8.72 7.387 / 8.68
int8-sdnq-hadamard layer 2.790 / 4.70 4.989 / 4.65
int8-sdnq-hadamard interval2 2.874 / 6.29 5.093 / 6.24
int8-sdnq-hadamard seg2-stride4 2.558 / 39.27 4.393 / 39.22
int8-sdnq-hadamard seg2-stride4-offload 3.711 / 22.67 6.695 / 22.63
fp8-torchao none 1.727 / 73.57 OOM
fp8-torchao activation-offload 4.061 / 10.08 7.404 / 9.96
fp8-torchao layer 2.607 / 5.98 4.888 / 5.93
fp8-torchao interval2 2.744 / 7.57 4.916 / 7.52
fp8-torchao seg2-stride4 2.246 / 40.55 4.245 / 40.50
fp8-torchao seg2-stride4-offload 3.602 / 24.02 6.683 / 23.91

Wan 2.1 T2V 14B

Example: wan2.1-t2v-14b-480p-single-gpu.peft-lora+ramtorch. Resolution: 832x480, 81f.

Note: Wan 14B is mainly a fit test for activation savings. Status-only cells are still useful because they show which combinations reached the memory limit.

Precision Mode H100 L40S
bf16 none OOM OOM
bf16 activation-offload 13.144 / 36.80 OOM
bf16 layer 7.162 / 16.28 21.770 / 16.23
bf16 interval2 7.172 / 19.62 21.777 / 19.58
bf16 seg2-stride4 OOM OOM
bf16 seg2-stride4-offload OOM OOM
int8-sdnq-hadamard none unsupported unsupported
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer unsupported unsupported
int8-sdnq-hadamard interval2 unsupported unsupported
int8-sdnq-hadamard seg2-stride4 unsupported unsupported
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none failed failed
fp8-torchao activation-offload failed failed
fp8-torchao layer failed failed
fp8-torchao interval2 failed failed
fp8-torchao seg2-stride4 failed failed
fp8-torchao seg2-stride4-offload failed failed

Wan S2V

Example: wan-s2v-14b-480p.peft-lora+ramtorch. Resolution: 832x480, 81f.

Note: Wan S2V is included as coverage data for the video/audio path. Treat failed cells as implementation coverage gaps.

Precision Mode H100 L40S
bf16 none failed failed
bf16 activation-offload unsupported unsupported
bf16 layer failed failed
bf16 interval2 unsupported unsupported
bf16 seg2-stride4 unsupported unsupported
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none failed failed
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer failed failed
int8-sdnq-hadamard interval2 unsupported unsupported
int8-sdnq-hadamard seg2-stride4 unsupported unsupported
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8-torchao none failed failed
fp8-torchao activation-offload unsupported unsupported
fp8-torchao layer failed failed
fp8-torchao interval2 unsupported unsupported
fp8-torchao seg2-stride4 unsupported unsupported
fp8-torchao seg2-stride4-offload unsupported unsupported

Z-Image Turbo

Example: z-image-turbo.peft-lora. Resolution: 1024x1024.

Precision Mode H100 L40S
bf16 none 0.243 / 21.25 0.316 / 21.21
bf16 activation-offload 0.837 / 13.24 0.805 / 13.19
bf16 layer 0.479 / 12.87 0.493 / 12.83
bf16 interval2 0.452 / 13.04 0.477 / 12.99
bf16 seg2-stride4 0.349 / 16.88 0.400 / 16.83
bf16 seg2-stride4-offload 0.681 / 15.03 0.736 / 14.99
int8-sdnq-hadamard none 0.645 / 15.60 0.615 / 15.55
int8-sdnq-hadamard activation-offload 1.439 / 7.61 1.382 / 7.56
int8-sdnq-hadamard layer 1.046 / 7.25 1.021 / 7.20
int8-sdnq-hadamard interval2 1.074 / 7.41 0.996 / 7.36
int8-sdnq-hadamard seg2-stride4 0.867 / 11.25 0.841 / 11.20
int8-sdnq-hadamard seg2-stride4-offload 1.202 / 9.39 1.162 / 9.35
fp8-torchao none 1.232 / 37.50 1.476 / 37.46
fp8-torchao activation-offload 3.623 / 7.97 3.564 / 7.93
fp8-torchao layer 2.319 / 7.93 2.344 / 7.88
fp8-torchao interval2 2.336 / 8.80 2.309 / 8.75
fp8-torchao seg2-stride4 1.843 / 22.55 1.930 / 22.50
fp8-torchao seg2-stride4-offload 2.947 / 15.57 3.243 / 15.52

ZLab I1

Example: zlab-i1.peft-lora. Resolution: 1024x1024.

Note: ZLab I1 carries its U-Net-style skip tensors through the segmented checkpoint state. Attention activation offload is not wired for this family.

Precision Mode H100 L40S
bf16 none 0.462 / 22.21 0.865 / 22.16
bf16 activation-offload unsupported unsupported
bf16 layer 0.693 / 7.79 1.148 / 7.75
bf16 interval2 0.676 / 8.30 1.152 / 8.25
bf16 seg2-stride4 0.567 / 14.97 1.014 / 14.92
bf16 seg2-stride4-offload unsupported unsupported
int8-sdnq-hadamard none 0.861 / 19.21 0.926 / 19.16
int8-sdnq-hadamard activation-offload unsupported unsupported
int8-sdnq-hadamard layer 1.385 / 4.78 1.265 / 4.74
int8-sdnq-hadamard interval2 1.298 / 5.30 1.277 / 5.26
int8-sdnq-hadamard seg2-stride4 1.073 / 11.98 1.098 / 11.93
int8-sdnq-hadamard seg2-stride4-offload unsupported unsupported
fp8wo-torchao none 0.504 / 25.08 0.930 / 25.02
fp8wo-torchao activation-offload unsupported unsupported
fp8wo-torchao layer 0.772 / 5.12 1.280 / 5.07
fp8wo-torchao interval2 0.759 / 5.84 1.290 / 5.79
fp8wo-torchao seg2-stride4 0.633 / 15.12 1.115 / 15.08
fp8wo-torchao seg2-stride4-offload unsupported unsupported