Skip to main content
The real-time API streams audio over a WebSocket and returns partial and final transcripts as they are produced. Like Batch, it uses the existing MyVocal API key, the same Characters balance and the same History. Read availability and known limitations, especially the connection recovery and processing-storage notes.

1. Create a session

Create a session with an Idempotency-Key and the options you want. The response returns a sessionId, a MyVocal transcriptionId and a streamUrl.

POST /sound_clone/api/v1/stt/realtime/sessions

Set the input encoding, the commit mode and the requested capabilities.

2. Send your audio with the right encoding

inputEncoding must match the bytes you send. pcm_s16le_16000 is signed 16-bit little-endian mono at 16 kHz; mulaw_8000 is one μ-law byte per sample. Always declare the microphone’s real sample rate, or a second of audio is interpreted as the wrong duration.

3. Connect the socket

A server client sends the accessKey header. A browser must not put its long-lived key in a URL: it uses a short-lived, single-use ticket and connects with ?ticket=. In a customer-facing app, your backend keeps the API key and creates the session and ticket; send only the ticket and session information to the browser. The supplied browser page is a local developer demo.

POST /sound_clone/api/v1/stt/realtime/sessions/{sessionId}/tickets

Issues a single-use browser ticket bound to the owner and the session.
Send audio with a MyVocal event envelope and read the results:
The server answers with session.ready, transcript.delta, transcript.final, transcript.alignment, transcript.revision, transcript.entities, usage.updated, session.notice, session.completed and session.error, each with an eventId, sessionId and a monotonic sequence.

4. Pause, resume or commit

  • segment.commit submits the audio so far and keeps the session recording.
  • session.pause / session.resume close and reopen a recording segment (a new epoch).
  • session.finish flushes the tail, settles once and completes the task.

POST /sound_clone/api/v1/stt/realtime/sessions/{sessionId}/finish

The terminal call; a repeated finish returns the same task and never charges twice.

5. Events in detail

Client to server, each as { "eventType", "epoch", "payload" }: Server to client, each as { "eventId", "sessionId", "sequence", "eventType", "payload" }: All 64-bit quantities are decimal strings; accepting, rotate and epoch stay boolean/number.

6. Rotation, recovery and errors

  • A usage.updated with rotate: true and accepting: false means the server wants a new segment: call session.pause, then session.resume (a new epoch), and continue sending from the new epoch. session.ready reports the live epoch and cursors so a reconnect knows where to resume.
  • A server session.error carries a MyVocal code (INPUT_INVALID, PROVIDER_UNAVAILABLE, INSUFFICIENT_CHARACTERS, TEMPORARILY_UNAVAILABLE, …). A plain disconnect is not a successful finish. Read the session over REST, then call finish to settle the audio the server can confirm. The result may be PARTIAL; unsent or unconfirmed audio is not restored by reconnecting. Full pause/resume across a service replacement is not yet verified.
  • A repeated frame at the same sampleOffset, a repeated segment.commit and a repeated session.finish are idempotent; they neither double-count nor create a second task.
The production socket can disconnect after about 60 seconds without traffic. Server clients can send a WebSocket ping every 20 seconds while idle; browser clients should handle disconnects and read the session before deciding what to do next. heartbeatIntervalMs does not keep the client-to-MyVocal connection alive on its own.

Complete examples

Read Characters and quotes for the billing rule.