Application IA prête à l’emploiWorkflow personnalisable

Synthèse vocale avec Speech 2.6 Turbo

Synthèse vocale multilingue à faible latence avec plus de 300 voix et des émotions expressives, grâce à MiniMax Speech 2.6 Turbo sur Replicate.

Aperçu de Synthèse vocale avec Speech 2.6 Turbo

Aucune configuration requise · Entièrement personnalisable dans AI-Flow

Ce que fait cette application

MiniMax Speech 2.6 Turbo delivers fast, natural text-to-speech optimized for real-time and interactive applications. Choose from 300+ curated voices (or your own cloned voice), switch emotions on the fly, and synthesize lifelike audio in 40+ languages—all with predictable, usage-based pricing.

Key capabilities

  • Low-latency synthesis ideal for chat agents, voice bots, and UI feedback

  • 300+ voices plus support for voice cloning via minimax/voice-cloning

  • Emotions set to auto or explicitly (happy, calm, surprised, neutral, and more)

  • 40+ languages with optional language boosting for better pronunciation

  • Flexible audio controlsspeed, pitch, volume, sample rate, bitrate, mono/stereo

  • Export to mp3, wav, flac, or pcm; optional sentence-level subtitles (non-streaming)

Inputs at a glance

  • text (required)Up to 10,000 characters. Insert pauses with markers like <#0.5#>

  • voice_idAny MiniMax system voice or a cloned voice ID

  • emotionauto or a specific emotion

  • audio_formatmp3, wav, flac, or pcm

  • sample_rate8000–44100 Hz; bitrate options for mp3

  • channelmono or stereo

  • speed, pitch, volumeFine-tune speaking rate, semitone shift (−12 to +12), and loudness

  • language_boostAutomatic or a specific language

  • english_normalizationImproves number/date reading in English (adds slight latency)

  • subtitle_enableReturns sentence timestamps in non-streaming mode

Output: - A hosted audio file URL (mp3/wav/flac/pcm), with MiniMax metadata such as character count and optional subtitles.

Pricing

  • $0.06 per 1,000 input tokens (roughly one token ≈ one character)

  • Audio output is not billed—only input text is counted

When to choose Turbo vs HD

  • Use Speech 2.6 Turbo for low-latency, interactive experiences

  • Prefer Speech 2.6 HD for maximum fidelity in long-form content like audiobooks and high-end voiceovers

Sound GenerationFast
Rapide à configurerEntièrement personnalisablePrêt à l’emploi

Comment utiliser l’application

Voici ce que vous obtenez en ouvrant l’application : un formulaire court, un bouton d’exécution et vos résultats, sans nœuds ni configuration.

Speech 2.6 Turbo

Aperçu de l’application
1

Remplissez le formulaire

Text

Requis

Saisissez text…

Audio Format

mp3

mp3wavflacpcm

Bitrate

128000

3200064000128000256000

Channel

mono

monostereo

Emotion

auto

autohappysadangryfearfuldisgusted+4 de plus

+8 options supplémentaires dans l’application

3

Récupérez votre résultat

Vos résultats apparaîtront ici

Sous le capot

Vous souhaitez plus de contrôle ?

Cette application repose sur un workflow AI-Flow entièrement modifiable. Changez le modèle, les prompts, la disposition ou la résolution, ou connectez-le à d’autres étapes.

Voice

Wise_Woman

Text

Once upon a quiet evening, in a small town that most people passed without noticing, something unexpected was about to happen.

Speech 2.6 Turbo

Comment utiliser ce template

1

Étape 1 : Saisissez votre texte dans le nœud « Texte »

Dans le nœud « Texte », saisissez vos instructions.

Exemple :
Once upon a quiet evening, in a small town that most people passed without noticing, something unexpected was about to happen.
2

Étape 2 : Configurez le nœud « Voice »

Configurez le nœud « Voice » selon vos besoins.

Exemple :
Wise_Woman
3

Étape 3 : Exécutez le workflow

Cliquez sur le bouton « Exécuter » pour lancer le workflow et obtenir le résultat final.

Personnalisez le workflow sous-jacent

Ouvrez le workflow complet dans l’éditeur AI-Flow pour changer de modèle, réécrire les prompts ou le connecter à d’autres étapes.

À qui s’adresse ce template ?

Idéal pour les professionnels et les créateurs qui souhaitent simplifier leur workflow

Developers building real-time voice experiences

Create chat agents, live voice assistants, and interactive apps that need low-latency, natural speech in many languages.

Product teams and UX designers

Add dynamic audio prompts, confirmations, and brand voices to product flows with consistent quality and fast response.

Support, IVR, and operations teams

Power multilingual IVR menus, helpdesk handoffs, status updates, and automated customer communications.

Content creators and marketers

Generate quick voiceovers for demos, explainers, and social content; clone voices for consistent brand sound.

Educators and training teams

Produce interactive tutorials and localized lessons across 40+ languages with emotion control and timestamped subtitles.

Prêt à créer ?

Commencez à utiliser cette application

Aucune configuration requise — ouvrez l’application et commencez à créer en quelques minutes

Questions fréquentes

What makes Speech 2.6 Turbo different from the HD variant?

Turbo is tuned for low latency and real-time interactions, like voice bots or live app feedback. HD prioritizes maximum fidelity for long-form, production-grade narration such as audiobooks and premium voiceovers.

How many languages and voices are supported?

MiniMax supports 40+ languages and dialects with optional language boosting. You can choose from 300+ curated voices or use a voice_id from the minimax/voice-cloning model.

How do I control style and emotions?

Set emotion to auto for inferred tone or pick a specific emotion (e.g., happy, calm, surprised, neutral). You can further adjust speed, pitch (−12 to +12 semitones), and volume to fine-tune delivery.

What audio formats and settings are available?

Export audio as mp3, wav, flac, or pcm. Configure sample_rate (8000–44100 Hz), channel (mono or stereo), and for mp3 select bitrate (32k, 64k, 128k, 256k).

Can I insert pauses in the speech?

Yes. Use inline markers like <#0.5#> within your text to pause for the specified number of seconds.

Does it support subtitles or timestamps?

Enable subtitle_enable to return sentence-level timestamps and subtitle metadata (available in non-streaming mode).

What’s the character limit for input text?

You can pass up to 10,000 characters in a single request.

Should I enable english_normalization?

Enable it if your English content contains numbers, dates, or structured text that benefits from improved normalization. It may add a small amount of latency.

Qu’est-ce qu’AI-FLOW et comment peut-il m’aider ?

AI-FLOW est une plateforme d’IA tout-en-un qui vous permet de créer, d’intégrer et d’automatiser des workflows alimentés par l’IA grâce à une interface intuitive en glisser-déposer. Débutant ou expert, vous pouvez combiner plusieurs modèles d’IA pour créer des solutions innovantes sans écrire de code.

Un essai gratuit est-il disponible ?

Oui, AI-FLOW propose un essai gratuit pour commencer. Vous pouvez ensuite acheter des crédits selon vos besoins, sans abonnement ni engagement à long terme.

Puis-je intégrer mes clés API OpenAI, Replicate et d’autres fournisseurs dans la version Cloud d’AI-FLOW ?

Oui, vous pouvez facilement intégrer vos propres clés API à AI-FLOW. Les nœuds associés utiliseront alors votre clé, ce qui réduit considérablement votre consommation de crédits sur la plateforme.