HunyuanCustom: Single-Image Deepfakes with Lip Sync
Tencent‘s new HunyuanCustom model generates high-quality deepfake videos from a single image, surpassing many existing methods. It leverages LatentSync for accurate lip synchronization with user-supplied audio and text, a significant advancement over previous models. The model excels in identity preservation, even with limited input images. HunyuanCustom efficiently handles video-to-video (V2V) editing, intelligently replacing masked segments of existing videos with subjects from a single reference image. Its ability to maintain identity fidelity across frames without relying on subject-specific fine-tuning is a key advantage. However, the system struggles with significant head rotations or varied facial expressions due to its reliance on a single source image. This limitation might hinder its ability to completely replace the existing ecosystem of LoRA models developed for the previous HunyuanVideo model. The model offers two versions: one requiring 8GB of GPU peak memory (720x1280px) and another needing 60GB (512x896px), with a recommendation of 80GB for optimal quality. Currently, the model is primarily tested on Linux. While the initial release lacks English-language examples, the extensive testing against competitors like Kling, Vidu, Pika, and VACE demonstrates HunyuanCustom’s competitive edge in identity consistency, subject similarity, and temporal consistency. The model’s architecture incorporates several existing technologies, including PySceneDetect, TextBPN-Plus-Plus, Koala-36M, Qwen7B, YOLO11X, InsightFace, QwenVL, Grounded SAM 2, Florence2, and Whisper, showcasing a sophisticated data pipeline. The integration of LLaVA improves semantic consistency between visual elements and textual descriptions. Despite its impressive capabilities, the high memory requirements and current Linux-only support may limit accessibility for some users.
The emergence of ai automation deepfakes technology like HunyuanCustom demonstrates how single images can now be transformed into realistic synchronized video content.
The rise of chatgpt automation deepfakes has made sophisticated video manipulation accessible to anyone with basic technical skills.
(Source: https://www.unite.ai/hunyuancustom-brings-single-image-video-deepfakes-with-audio-and-lip-sync/)

