Skip to main content
POST
Create a transcription
Submits one transcription from an uploadId or a public mediaUrl. The call quotes the work, reserves Characters and starts the asynchronous task in a single request.
string
required
API key for authentication.
string
required
Binds this request. The same key and the same body return the same task; the same key with a different body is refused as IDEMPOTENCY_CONFLICT.

Body

string
An upload from POST /uploads. Use exactly one of uploadId or mediaUrl.
string
A public http/https URL. The server fetches it over a real connection and re-checks every redirect; private, loopback, link-local and cloud-metadata addresses are refused. There is no domain or port allowlist.
number
Which audio track to transcribe. Required when the source contains more than one audio track; use the index from the upload’s media.audioTracks.
string
The spoken language, or omit for automatic detection.
string
A title for the History row.
string
Optional model identifier; defaults to myvocal_stt_v1.
object
Processing options: separateSpeakers, maxSpeakerCount, alignmentLevel, entityCategories, redactCategories, redactionStyle, rewriteInstruction, removeDisfluencies, identifyConversationRoles, vocabularyHints, separateChannels, channelResultMode, exportFormats, inputEncoding, processingContentStorage and more. Option combinations that the model does not support are refused at submit with INPUT_INVALID rather than silently dropped. Completion notifications are options too: options.notifyOnCompletion (true to enable), options.notifyUrl (a public URL MyVocal posts the signed event to), options.clientMetadata (opaque caller data echoed back) and options.notificationSigningSecret (the write-only HMAC key). Polling stays the authoritative result. See capabilities for the live list.
boolean
When true, this call waits for a terminal state inside the same request.
number
The wait budget in milliseconds. If the task is still running at the budget, the request returns the task with its identity; it is not cancelled.

Batch options

All fields below belong inside options. Omit fields you do not need.

Option combinations

  • rewriteInstruction cannot be combined with entity detection, redaction or separateChannels.
  • Multi-channel processing does not separate speakers. Do not combine separateChannels with maxSpeakerCount, speakerMergeSensitivity, identifyConversationRoles or pcm_s16le_16 input. The selected audio track can have at most five channels for transcription.
  • Multi-channel combined results require alignment and cannot be combined with entity detection or redaction.
  • matchKnownSpeakers is unavailable. Do not request it.
Each exportOptions entry has a required format and optional controls: Only relevant formats use each control; json is the public transcription view. Without alignment, SRT cannot produce timed cues. These options do not change recognition or billing.
Some invalid entity categories currently fail asynchronously as MEDIA_UNREADABLE even for valid audio. See known limitations before retrying.

Response

Successful REST calls return { code: 1, message, data }. The fields below are inside data; JSON downloads are the documented exception and return the transcript view directly.
number
1 for success.
string
Result message.
object