MiniMax Speech 2.8 Turbo Async Text-to-Speech
Audio
MiniMax Speech 2.8 Turbo Async Text-to-Speech
POST
MiniMax Speech 2.8 Turbo Async Text-to-Speech
MiniMax asynchronous text-to-speech API, supports various voice, emotion, speed and other parameter settings, text length limit up to 50,000 characters, supports file input (up to 100,000 characters)
Request Headers
string
required
Supports:
application/jsonstring
required
Bearer authentication format, for example: Bearer {{API Key}}.
Request Body
string
Text to synthesize into audio, maximum length is 50,000 characters. Either
• Interjection tags: Only supported when model is
text or text_file_id is required.• Interjection tags: Only supported when model is
speech-2.8-hd or speech-2.8-turbo. Supported interjections: (laughs) (laughter), (chuckle) (light laugh), (coughs) (cough), (clear-throat) (clear throat), (groans) (groan), (breath) (normal breathing), (pant) (panting), (inhale) (inhale), (exhale) (exhale), (gasps) (gasp), (sniffs) (sniff), (sighs) (sigh), (snorts) (snort), (burps) (burp), (lip-smacking) (lip smacking), (humming) (humming), (hissing) (hissing), (emm) (um), (whistles) (whistle), (sneezes) (sneeze), (crying) (crying), (applause) (applause)integer
Text file ID for audio synthesis, single file length limit is less than 100,000 characters, supported file formats: txt, zip. Either
• txt file: Length limit <100000 characters. Supports custom pause using
• zip file:
• Compressed package must contain txt or json files of the same format.
• json file format: Supports [
text or text_file_id is required, format will be automatically validated.• txt file: Length limit <100000 characters. Supports custom pause using
<#x#> tag. x is pause duration (in seconds), range [0.01, 99.99], up to 2 decimal places. Pause must be set between two pronounceable text segments, cannot use multiple pause tags consecutively• zip file:
• Compressed package must contain txt or json files of the same format.
• json file format: Supports [
title, content, extra] three fields, representing title, body, and additional information. If all three fields exist, 3 groups of results will be produced, 9 files in total, stored in one folder. If a field does not exist or is empty, no corresponding result will be generatedobject
object
object
required
boolean
default:false
Controls whether to add audio rhythm identifier at the end of synthesized audio, default is False. This parameter is only valid for non-streaming synthesis
string
Whether to enhance recognition ability for specified minor languages and dialects. Default is
null, can be set to auto to let the model decide automatically.Optional values: Chinese, Chinese,Yue, English, Arabic, Russian, Spanish, French, Portuguese, German, Turkish, Dutch, Ukrainian, Vietnamese, Indonesian, Japanese, Italian, Korean, Thai, Polish, Romanian, Greek, Czech, Finnish, Hindi, Bulgarian, Danish, Hebrew, Malay, Persian, Slovak, Swedish, Croatian, Filipino, Hungarian, Norwegian, Slovenian, Catalan, Nynorsk, Tamil, Afrikaans, autoobject
Response
string
Use the task_id to request the Task Result API and retrieve the generated output.
object
Additional task details.
Last modified on August 3, 2026