Description:
An open-source framework for audio-driven human video generation. LongCat-Video-Avatar 1.5 improves lip-sync with Whisper-Large, enhances production stability for long videos, generalizes to stylized domains, and enables faster 8-step inference.