Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching
- VGen
Main:2 Pages
2 Figures
Bibliography:1 Pages
1 Tables
Abstract
We present Livatar, a real-time audio-driven talking heads videos generation framework. Existing baselines suffer from limited lip-sync accuracy and long-term pose drift. We address these limitations with a flow matching based framework. Coupled with system optimizations, Livatar achieves competitive lip-sync quality with a 8.50 LipSync Confidence on the HDTF dataset, and reaches a throughput of 141 FPS with an end-to-end latency of 0.17s on a single A10 GPU. This makes high-fidelity avatars accessible to broader applications. Our project is available atthis https URLwith with examples atthis https URL
View on arXivComments on this paper
