ProCreations's picture
Accelerate full 40-step FP8 generation with native precision, measured quality and real-time demo
1081be0 verified
|
Raw History Blame Contribute Delete
4.52 kB

Runtime optimization with unchanged denoising precision

Full-compute runtime update

The default generate.py now uses decode-only dynamic compilation and kernel fusion, keeping all 40 denoising steps, original calibrated FP8 weights/scales, FP32 GEMM accumulation, and native BF16 attention. No step reuse, distilled adapter, further attention quantization, or reduced resolution is enabled. --eager selects the original runtime.

Resolution Original FP8, fresh control Optimized FP8 Speedup
1024×1024 6.922 s 5.932 s 1.167×
2048×2048 44.101 s 37.289 s 1.183×

Same RTX PRO 6000, batch 1, 40 steps, CFG 1, prefix KV cache. CUDA-synchronized end-to-end generation includes encoding and VAE. Warmup/model loading/PNG writes excluded. Original control has 2 measured repeats per size; optimized has 5 at 1024 and 3 at 2048. Raw measurements include first-use warmup and load costs. First-use compilation costs extra time. --warmup reports that cost separately; --prompts-json prompts.json processes a list of prompt strings in one loaded pipeline.

All 18 paired checks were visually reviewed, including primary English/Chinese titles, portraits, transparency and edits. No obvious general visual-quality loss was observed in this finite set. Outputs are not bit-identical: fine textures, decorative typography and some pottery positioning change with GPU reduction rounding. Mean LPIPS versus original FP8=0.014105, mean SSIM=0.985841; worst LPIPS=0.104154 (pottery). Relative to BF16, mean LPIPS=0.035437, compared with 0.034470 for the original FP8. These are fidelity checks, not a guarantee for every prompt. Comparisons · Detailed report.

Updated 30-second real-time image-only video. One completed warmup image is shown at time 0, then each new image appears when actual generation finishes. All waits remain at 1× speed. Frame/request timestamps.

Implementation

acceleration.py compiles the 32 repeated transformer blocks for cached-prefix decode, while keeping prefill in the upstream eager path. Dynamic lengths avoid per-prompt recompilation after warming relevant resolution shapes. Real-valued rotary arithmetic is fused with surrounding operations. Cached decode contains only target tokens, allowing direct per-sample modulation broadcasting instead of materializing a target/prefix mask across every token. Inductor emulate_precision_casts=True preserves intermediate BF16 rounding boundaries. Floating-point GPU reduction ordering can still differ.

Weights and calibration were not changed. Actual denoising-step profiling still observes 224 native CUTLASS SM120 E4M3 weight GEMMs and 32 native BF16 FlashAttention kernels. Exactly 40 transformer calls occur for 40 steps. See kernel evidence and trace.

Investigated and not adopted

Fast FP8 accumulation, cuDNN attention, FlashAttention4 and a 24-configuration Triton FP8 GEMM probe did not offer a meaningful applicable improvement. SageAttention 2 and conservative first-block residual reuse produced larger gains, but introduce additional approximation; both were excluded following the explicit quality requirement. Neither is required or enabled by this release.

Quality review

No obvious general visual quality loss was observed in all 18 paired images, including portraits, wildlife, macro, food, architecture, paintings, glass, primary English and Chinese titles, transparency, crafts, coastline, flowers, human hands, and two edits. Outputs are not bit-identical: potter arm and clay positioning, snow leopard coat and background details, decorative typography, and some fine textures differ. The pottery sample has the largest LPIPS difference (0.10415) and remains visually plausible. This finite review cannot guarantee every prompt.

The original release remains immutable at f642d1f49f031f550655d07ca4569ff729ba9cfd. --eager is available for the original execution path. Tested Torch 2.14.0+cu130, Triton 3.8.0, Transformers 5.17.0, Diffusers 80c7ed262aeffbeb43ef13ae04baeb9b84515a69, RTX PRO 6000 SM120. Research license unchanged.

The source scripts retain explicit workstation layout assumptions; adjust input/output paths for another installation.