Skip to main content
POST
Text to Speech · Workflow: TTS quickstart · Models: TTS Models & Migration
string
required
Your API key. See Authentication.

Body

string
required
Voice ID returned by GET /sound_clone/api/v1/voices.
For myvocal_v3_accent_enhance, use a voice returned by GET /sound_clone/api/v1/voices?modelId=myvocal_v3_accent_enhance.
string
Optional title for the generated record.
string
required
Text to synthesize.
string
Language selector.
Default/V2 path: required and must use existing mapped language values.
myvocal_v3: optional. If provided, it is forwarded to upstream as language. If omitted, upstream auto-detect is used.
See supported V3 codes in V3 Language Codes (98 languages).
myvocal_v3_accent_enhance: optional. Omit it (or send auto) for automatic detection. The Accent model has its own code set (40 language codes plus auto); see TTS Models & Migration. A plain V3 code the Accent model does not support (for example hmn) is rejected with code = 10004. Legacy mapped enum values such as US_ENGLISH are not Accent codes; use en.
string
Optional explicit model selector.
Omit it to use the default model path.
Set myvocal_v3 to request the V3 path.
Set myvocal_v3_accent_enhance to request the Accent model path.
When the Accent model is requested but its execution path is unavailable, the response is the standard generation error envelope (code = 10003). The request is never silently served by a different model.
object
For advanced settings, use suggested values unless you have strong tuning needs.
For myvocal_v3, the currently publicly documented stability values are:
- 0.5: natural pronunciation, closer to normal human speech.
- 1: more stable output with less variation.
For myvocal_v3_accent_enhance, plain text is the supported minimal path. The inline [tag] prompt control documented for myvocal_v3 is not documented as equivalent, and the service does not rewrite your text. Send plain text for the Accent model.

Prompt Control (myvocal_v3)

Reference capability source: v3.myvocal.ai. Use [tag] in text to guide style and delivery.
Basic syntax:
Examples:
Representative categories:
  • Emotion control: happy, excited, sad, angry, surprised
  • Non-speech sounds: laughter, throat clearing, yawning, shushing, screaming
  • Sound-related effects: breathing, sighing, gasping, crying, murmuring

Response

string
Success usually returns audio/mpeg.
stream
Binary audio stream payload.
An HTTP 200 does not by itself mean the request succeeded. Read Content-Type before trusting the body:
  • audio/* (normally audio/mpeg): the audio payload described above.
  • any JSON content type: the standard result envelope with a non-success code (see Common errors below).
The two paths fail differently. Before any audio is sent, a rejected or failed request is answered with the JSON envelope above. After audio has started, the connection is terminated instead of appending a JSON body, because appending JSON to a partial audio body would corrupt it. Treat a short or interrupted download as a failed request, not as a complete audio file.

Common errors

  • Invalid language option. (code = 10004): non-V3/default path is missing or using an unsupported language value; or the Accent model received a code outside its own language list.
  • VOICE_NOT_FOUND (code = 10005): voiceId is invalid or not available under this API key.
  • Invalid stability option. (code = 10006): voiceSettings.stability is outside the legal range.
  • code = 10003: the Accent model path was requested but its execution path is unavailable. The request is not silently served by another model.
  • code = -1: upstream or backend failed to complete generation.