Series 3 — Video, with Sound and Lip Sync
Short cinematic clips, animated stills, and talking-head presenters speaking in your own cloned voice — generated entirely on a 150 mm cube, in minutes rather than seconds.
Video is the workload that most tests this machine's character. The models fit — comfortably, in a way they do not on a 24 GB card — and then take minutes per clip, because generation is iterative and this hardware's memory bandwidth is modest. That trade is the theme of the whole series: you can attempt things here that a consumer GPU cannot hold at all, provided you are prepared to wait for them.
The through-line is the audio from series 2. A generated voice is what makes a talking head worth generating, and it is what turns dubbing from a novelty into something useful.
-
What a Generated Clip Actually Costs Here lead
The model landscape as it really stands — LTX-2.3 with audio generated in the same pass, Wan 2.2 for photoreal humans, HunyuanVideo for physics — plus honest wall-clock numbers, memory footprints, and a decision table matching each of your three goals to a pipeline.
-
ComfyUI on GB10 That Doesn't Run Out of Memory build
The unglamorous foundation: why ComfyUI appears capped at 64 GB on a 128 GB machine, the exact launch flags that fix it, SageAttention built for sm_121a, a model library on external storage, and access from the rest of the LAN.
-
B-Roll and Image-to-Video with LTX-2.3 build
Text-to-video and image-to-video with synchronised audio in a single diffusion pass. The resolution and duration ladder, distilled versus development weights, the quality knobs that matter, and the patch that stops audio encoding from crashing.
-
Talking Heads and Dubbing build
Driving video from the Polish and Italian voices you cloned in series 2: LipDub for re-voicing existing footage, InfiniteTalk for long-form presenters, identity preservation across minutes of speech, and wrapping ComfyUI in a job API so your own code can call it.
Prerequisites. The
hub page for the hardware constraints — the unified-memory section matters more here than anywhere else — and
series 2 if you want the talking-head and dubbing workflows to speak with a voice you chose.