@ihteshamali ·
🚨 BREAKING: A research lab just released a 15B model that generates multilingual talking human videos with synced audio, beats every competitor in human evaluation, and runs in 38 seconds on one GPU. It's called daVinci-MagiHuman. The key insight is that every other model in this category stacks cross-attention, multi-stream pipelines, and separate conditioning branches to handle video and audio together. This one throws all of that out and uses a single unified self-attention stream across all modalities. Super-resolution happens in latent space rather than pixel space so there's no extra VAE decode-encode round trip. The turbo VAE decoder cuts decoding overhead even further. The distilled version runs in 8 steps with no CFG at all. Visual quality, text alignment, and word error rate all beat Ovi 1.1 and LTX 2.3 on the benchmark table. 100% Opensource. Apache 2.0. Repo and research paper links are in the comments.
- 19Replies
- 116Reposts
- 696Likes
- 43.8KViews








































