Also supports training on a single 80GB GPU.
Use -1 for a random seed.
This demo uses a four-block startup buffer for smoother streaming on Hugging Face ZeroGPU; local deployments can use one block when generation keeps pace.