by.waclaw.online / spark / video

Series 3 — Video, with Sound and Lip Sync

Short cinematic clips, animated stills, and talking-head presenters speaking in your own cloned voice — generated entirely on a 150 mm cube, in minutes rather than seconds.

Video is the workload that most tests this machine's character. The models fit — comfortably, in a way they do not on a 24 GB card — and then take minutes per clip, because generation is iterative and this hardware's memory bandwidth is modest. That trade is the theme of the whole series: you can attempt things here that a consumer GPU cannot hold at all, provided you are prepared to wait for them.

The through-line is the audio from series 2. A generated voice is what makes a talking head worth generating, and it is what turns dubbing from a novelty into something useful.

  1. What a Generated Clip Actually Costs Here lead The model landscape as it really stands — LTX-2.3 with audio generated in the same pass, Wan 2.2 for photoreal humans, HunyuanVideo for physics — plus honest wall-clock numbers, memory footprints, and a decision table matching each of your three goals to a pipeline.
  2. ComfyUI on GB10 That Doesn't Run Out of Memory build The unglamorous foundation: why ComfyUI appears capped at 64 GB on a 128 GB machine, the exact launch flags that fix it, SageAttention built for sm_121a, a model library on external storage, and access from the rest of the LAN.
  3. B-Roll and Image-to-Video with LTX-2.3 build Text-to-video and image-to-video with synchronised audio in a single diffusion pass. The resolution and duration ladder, distilled versus development weights, the quality knobs that matter, and the patch that stops audio encoding from crashing.
  4. Talking Heads and Dubbing build Driving video from the Polish and Italian voices you cloned in series 2: LipDub for re-voicing existing footage, InfiniteTalk for long-form presenters, identity preservation across minutes of speech, and wrapping ComfyUI in a job API so your own code can call it.
Prerequisites. The hub page for the hardware constraints — the unified-memory section matters more here than anywhere else — and series 2 if you want the talking-head and dubbing workflows to speak with a voice you chose.