Génération vidéo et audio avec LTX-2 Distilled
Générez une vidéo et un audio synchronisés à partir de prompts textuels ou visuels avec Lightricks LTX-2 Distilled, le premier modèle audio-vidéo open source conçu pour des résultats rapides de qualité production.
Aucune configuration requise · Entièrement personnalisable dans AI-Flow
Ce que fait cette application
LTX‑2 Distilled creates video and audio together in a single pass, delivering natural lip‑sync, ambient sound, and music that match on‑screen action. Provide a cinematic text prompt—or an optional reference image for image‑to‑video—and get a cohesive clip ready for previews, social posts, ads, or concept tests.
Optimized for speed and iteration, the model produces 1080p by default and can render up to 4K. It supports common aspect ratios (16:9, 9:16, 1:1, and more) and short clips, with frame counts governed by 8*k+1 for smooth timing. Use seeds for reproducibility, toggle prompt enhancement, and control image adherence with image_strength for faithful i2v results.
Under the hood, LTX‑2 Distilled uses an asymmetric dual‑stream diffusion transformer (video + audio) with cross‑attention for tight AV synchronization. It’s quantized for efficient inference and supports LoRA fine‑tuning and control adapters (depth, pose, edge) for style, motion, and identity consistency.
Best practices: write prompts like shot directions—describe camera movement, lighting, scene details, and the desired soundscape. Keep prompts under ~200 words, start with the action, and specify pacing. The output is a URI to an MP4 file you can download or stream.
Constraints and tips: width and height must be divisible by 32; choose a frame count of 8*k+1; for reliable results stick to 1080p unless you specifically need 4K. The model generates plausible content (not factual), and audio quality is strongest for speech and natural environments.
Comment utiliser l’application
Voici ce que vous obtenez en ouvrant l’application : un formulaire court, un bouton d’exécution et vos résultats, sans nœuds ni configuration.
LTX-2 Distilled
Aperçu de l’applicationRemplissez le formulaire
Prompt
RequisSaisissez prompt…
Aspect Ratio
16:9
Enhance Prompt
Désactivé par défaut
Image
Ajouter des images ou des fichiers
Image Strength
+1 option supplémentaire dans l’application
≈ 5–25 crédits / exécution
Récupérez votre résultat
Résultat réel généré avec cette application
Sous le capot
Vous souhaitez plus de contrôle ?
Cette application repose sur un workflow AI-Flow entièrement modifiable. Changez le modèle, les prompts, la disposition ou la résolution, ou connectez-le à d’autres étapes.
Prompt
An animated cinematic shot. a robot, walks slowly, the camera dollys back and keep the robots slow walk in a medium shot. the robot start running slowly and heavily. it then stops, and the camera keeps dollying back, until a blue similiar robot appears in a...
Image

Ltx 2 Distilled
Comment utiliser ce template
Étape 1 : Saisissez votre texte dans le nœud « Prompt »
Renseignez le nœud « Prompt » avec le texte demandé.
An animated cinematic shot. a robot, walks slowly, the camera dollys back and keep the robots slow walk in a medium shot. the robot start running slowly and heavily. it then stops, and the camera keeps dollying back, until a blue similiar robot appears in an over the shoulder shot.
Étape 2 : Importez votre fichier
Dans le nœud « Image », importez le fichier que vous souhaitez traiter.

Étape 3 : Exécutez le workflow
Cliquez sur le bouton « Exécuter » pour lancer le workflow et obtenir le résultat final.
Personnalisez le workflow sous-jacent
Ouvrez le workflow complet dans l’éditeur AI-Flow pour changer de modèle, réécrire les prompts ou le connecter à d’autres étapes.
À qui s’adresse ce template ?
Idéal pour les professionnels et les créateurs qui souhaitent simplifier leur workflow
Video creators and editors
Produce short, high‑quality clips with synchronized sound directly from prompts for storyboards, animatics, and social video.
Marketers and brand teams
Rapidly prototype ads and promos with on‑brand visuals, matched music/ambience, and accurate lip‑sync for dialogue.
Content and social media creators
Turn ideas or thumbnails into engaging vertical or horizontal videos with natural audio in seconds.
Product teams and UX researchers
Generate realistic demo footage and scenarios with controllable pacing, framing, and environmental sound.
Developers and researchers
Leverage an open‑source, AV‑synchronized text‑to‑video and image‑to‑video model with LoRA fine‑tuning and control adapters.
Prêt à créer ?
Commencez à utiliser cette application
Aucune configuration requise — ouvrez l’application et commencez à créer en quelques minutes
Vous aimerez peut-être aussi
Explorez d’autres templates puissants pour enrichir votre workflow IA
Génération vidéo avec Kling V2.6
Kling V2.6 est un générateur vidéo IA professionnel qui transforme du texte ou une image en clips cinématographiques 1080p aux mouvements fluides, avec audio natif synchronisé : dialogues, ambiance et effets.
Workflow de création de publicité UGC — Du script à la vidéo
Créateur complet de publicités UGC qui transforme une photo du sujet, une photo du produit et un script facultatif en une première image prête à l’emploi, puis en une vidéo verticale de 8 s avec voix et mouvements naturels de caméra à main levée.
Générer des animations réalistes de synchronisation labiale à partir d’un audio
Générez des animations réalistes de synchronisation labiale à partir de n’importe quelle piste audio. PixVerse Lipsync aligne naturellement les mouvements de la bouche, le rythme de la parole et les expressions.
Génération vidéo avec Kling V2.5 Turbo Pro
Kling 2.5 Turbo Pro permet de créer des vidéos de niveau professionnel à partir de texte ou d’images, avec des mouvements fluides, une profondeur cinématographique et un excellent respect du prompt.
Sora 2
Latest version of Sora, with higher-fidelity video, context-aware audio, reference image support
Génération vidéo avec Veo 3.1
Nouvelle version améliorée de Veo 3, avec une vidéo plus fidèle, un audio sensible au contexte et la prise en charge d’une image de référence et de la dernière image.
Questions fréquentes
What is LTX‑2 Distilled?
An open‑source, speed‑optimized audio‑video generation model from Lightricks that produces synchronized video and audio from text or image prompts.
How is this different from typical video generators?
Most models output silent video or add audio later. LTX‑2 Distilled generates video and audio together, resulting in natural timing, lip‑sync, ambience, and music that match the visuals.
Does it support image‑to‑video?
Yes. Provide a reference image to preserve composition, lighting, and style while the model adds motion and synchronized sound. Adjust image_strength (0–1) to control fidelity to the source image.
What inputs can I control?
Prompt (required), optional image, aspect_ratio, num_frames (must be 8*k+1), seed for reproducibility, enhance_prompt (true/false), and image_strength for i2v adherence.
How long can the generated videos be?
Clips are short‑form and governed by the num_frames setting (8*k+1). Choose higher frame counts for longer clips; practical defaults align with ~30 fps for smooth playback.
What resolutions and aspect ratios are supported?
1080p is the default and recommended for speed and quality; the model can render up to 4K. Common aspect ratios include 16:9, 9:16, 1:1, 4:3, 3:4, and 21:9. Ensure width and height are divisible by 32.
How do I write effective prompts?
Describe the shot like a cinematographer: who/what moves, camera angles, pacing, lighting, colors, setting, and the soundscape (dialogue, ambience, music). Keep under ~200 words and start with the action.
Is the audio lip‑sync accurate?
Yes. The model’s dual‑stream architecture uses cross‑attention for strong AV synchronization, producing precise lip‑sync and sound timing relative to visual events.
What does the output look like?
The API returns a URI to an MP4 video file containing both the generated visuals and synchronized audio.
Can I reproduce results?
Yes. Set a seed to improve reproducibility across runs with the same inputs and parameters.
Can I fine‑tune the model?
LTX‑2 Distilled supports LoRA fine‑tuning for styles, motion patterns, and identities, plus control LoRAs (depth, pose, edge) and IC‑LoRAs for identity consistency and v2v transforms.
Are there limitations?
The model generates plausible, not factual, content and may reflect dataset biases. Audio quality is best for speech and natural ambiences; abstract audio may be lower fidelity. Follow the frame and dimension constraints for stable results.
Qu’est-ce qu’AI-FLOW et comment peut-il m’aider ?
AI-FLOW est une plateforme d’IA tout-en-un qui vous permet de créer, d’intégrer et d’automatiser des workflows alimentés par l’IA grâce à une interface intuitive en glisser-déposer. Débutant ou expert, vous pouvez combiner plusieurs modèles d’IA pour créer des solutions innovantes sans écrire de code.
Un essai gratuit est-il disponible ?
Oui, AI-FLOW propose un essai gratuit pour commencer. Vous pouvez ensuite acheter des crédits selon vos besoins, sans abonnement ni engagement à long terme.
Puis-je intégrer mes clés API OpenAI, Replicate et d’autres fournisseurs dans la version Cloud d’AI-FLOW ?
Oui, vous pouvez facilement intégrer vos propres clés API à AI-FLOW. Les nœuds associés utiliseront alors votre clé, ce qui réduit considérablement votre consommation de crédits sur la plateforme.