- resolve seed=-1 once at top of run_ltx and thread the int through every
stage (StableDiffusionProcessingVideo, _latent_pass, upsample 0.9/2.x,
refine, post-refine vae_decode). get_generator(-1) reseeds globally per
call so each stage was rolling an uncorrelated value; p.seed now carries
the resolved seed so reruns reproduce.
- pre-encode prompts via shared.sd_model.encode_prompt before the latent
path; park the four tensors on CPU and pass them as prompt_embeds /
*_attention_mask kwargs through _latent_pass and refine_args. Stage 2
reuses Stage 1's embeds instead of re-running the text encoder. The
manual encode is outside pipe.__call__ so the post-forward offload hook
never fires; apply_balanced_offload(force=True) re-anchors the device
map so the text encoder doesn't stay pinned through Stage 1.
- skip audio_vae.decode + vocoder on refine when audio_enable=False:
output_type='latent' bypasses the internal audio + video decode and
hands off to the post-refine vae_decode block. Per-step audio
cross-attention still runs for video conditioning.
drop the project-specific stage 1 direct audio decode (a6870f7d2). smoke
testing showed broadband tinniness on distilled is BWE-bound, not stage 2
corruption, so the deviation didn't fix the underlying issue.
revert to canonical:
- _latent_pass returns video latents only; result.audio is unused.
- stage 2 receives audio_latents=None (default), prepare_audio_latents
generates fresh gaussian noise, audio scheduler runs the 3 stage-2
sigmas under identity guidance, video<->audio cross-attention
conditions the audio branch.
- capture stage 2 result.audio[0].float().cpu() for save.
upstream evidence: pipeline_ltx2.py:937-940 documents audio_latents as
pre-generated noisy latents (initial gaussian, not stage 1 output); no
caller in diffusers threads stage outputs into the kwarg.
non-latent path is unchanged and continues to work via p.audio_capture.
process_decode strips video pipeline output to a flat list of frames at the
PIL early-return (processing_diffusers.py:461-465), so any output.audio is
lost before processing.process_images returns. video pipelines that produce
synchronized audio (LTX-2 audio-capable models) were getting silent mp4s on
the non-latent path.
stash output.audio on p.audio_capture before process_decode runs and let
processing read it back as a fallback when samples is a flat list.
ltx_process non-latent branch strips the (B, 2, N) batch dim with [0] so
write_audio's .T+contiguous() path produces interleaved bytes for AAC s16.
stop threading stage 1 audio_latents into stage 2 refine. Lightricks/LTX-2#126
reports the two-stage pipeline degrades audio quality, confirmed locally as
clean speech with tinny ambient/foley/music on the threaded path.
root cause: stage 2 prepare_audio_latents calls _create_noised_state at
noise_scale=0.909 (pipeline_ltx2.py:704-714, 598-603), keeping ~9% of stage 1
signal. 3 refine steps recover speech via video<->audio cross-attention but
not broadband content.
new path: _latent_pass decodes audio_latents to waveform via audio_vae +
vocoder mirroring pipeline_ltx2.py:1471-1473 exactly (input cast to
audio_vae.dtype, no module dtype mutation). stage 2 result.audio is
discarded; cross-attention still runs each block for video conditioning.
relabel toggle to 'LTX save audio' (default true) since audio always
generates on 2.x audio-capable models; the toggle gates mux only. hint
added to locale_en.json.
split add_audio_stream from write_audio. avformat_write_header runs on
first container.mux() and freezes the stream set, so audio added after
video packets has time_base=0/0 and raises 'Cannot rebase to zero time.'
atomic_save_video registers the audio stream before the encode loop.
applied_layers is cleared and re-populated on every network_activate call.
With lora_apply_te=True the second activate (TE-only pass) finds all
modules already at the target state and skips them all, leaving
applied_layers empty and breaking the restore trigger on the next gen.
native_active is set from loaded_networks at the end of activate, so it
survives idempotent re-runs and only flips false after the restore call
clears loaded_networks.
Flux2/Klein loads LoRAs as native modules through the diffusers method
path. network_activate() was only called when new native modules existed,
so removing a LoRA from the prompt left backed-up weights unrestored.
Check applied_layers to detect previously active native modules and
trigger network_activate() for the restore path.
When a video script (animatediff, text2video, image2video, stablevideodiffusion)
runs via Control tab, both the script's save and control_run's end-of-run save
fired. The latter crashed silently because the local `video` cv2 capture name at
control_run:584 shadowed the modules.video import, so the duplicate was hidden
and the gallery video link never propagated.
- alias import as video_module to bypass the shadow
- p.video_saved marker set by each script
- control_run skips its end-of-run save when the marker is set
LTX refine pipe uses output_type='pil' so result.frames[0] returns a list, not a 5-D tensor. Convert via images_to_tensor before the helper sees it, mirroring what save_video already does for the same input.