Real-Time API Documentation
// Overview
The ElevateAI Real-Time API provides clients with low latency, highly accurate, real-time transcription utilizing the Secure WebSocket protocol.
The Client Application establishes a WebSocket connection with the Real-Time API for each channel in a session:
- One (1) WebSocket for a Mono (Single Channel) Audio Stream
- Two (2) WebSockets for a Stereo (Dual Channel) Audio Stream
Communication occurs between the Client Application and the Real-Time API via messages passed along the bidirectional WebSocket(s).
At a high level, the Client Application sends the audio, in real-time, across the bidirectional WebSocket(s). The Real-Time API processes the audio as soon as a slight pause in speech is detected. The Real-Time API then sends the resulting transcription back to the Client Application, across each connected WebSocket, as a response message.
The Client Application should have a listener waiting to receive these response messages and consume them in the desired manner.
Audio data is sent in binary form with PCM, G711u or G711a encoding. All other WebSocket communication is done via text requests/responses in JSON format.
// Connecting to the ElevateAI Real-Time API
The request to establish WebSocket communication is done via a request to a wss: URI.
The details required to establish the connection and properly communicate the binary audio data are included as Route and Query String parameters within the URI, with details provided below.
- For authentication, the client's ElevateAI API Token is passed via the encrypted header of the request.
The elements that make up the ElevateAI Real-Time API WebSocket URI are below:
Base URI: wss://api.elevateai.com
// Full Template
wss://api.elevateai.com/v1/{interactionType}/{languageTag}/{vertical}?session_identifier=<string>&channels=<int>&channel_index=<int>&participant_role=<string>&participant_name=<string>&codec=<string>&sample_rate=<int>&bit_depth=<int>&bit_rate=<int>
Example: Agent Channel
wss://api.elevateai.com/v1/audio/en/default?session_identifier=fe5f75b8-fa06-428b-bxy2-7f0a0b6bf169&channels=2&channel_index=0&participant_role=Agent&codec=pcm&sample_rate=16000
Example: Customer Channel
wss://api.elevateai.com/v1/audio/en/default?session_identifier=fe5f75b8-fa06-428b-bxy2-7f0a0b6bf169&channels=2&channel_index=1&participant_role=Customer&codec=pcm&sample_rate=16000
// Connection Parameters
Headers
| | | |
|---|---|---|---|
Parameter Name | Path Text | Valid Inputs | Required? |
X-API-TOKEN | Attach as Header | Any Valid, Activated ElevateAI API Token | Required |
Route Parameters
All of these route parameters are of type string:
| | | |
|---|---|---|---|
Parameter Name | Path Text | Valid Inputs | Required? |
Interaction Type | v1/{interactionType} | audio | Required |
Language Tag | v1/audio/{languageTag} | Any Supported BCP 47 Language Identifier (w/o localization), i.e.: en | Required |
Vertical | v1/audio/en/{vertical} | default | Required |
Query String Parameters
| | | |
|---|---|---|---|
Parameter Name | Path Text / Description | Valid Inputs | Required? |
Session Identifier | ?session_identifier=<string> Used to correlate one or more client connections related to the same interaction. All connections for a single interaction must have matching session identifier. GUID recommended | Unique String | Required |
Channels | &channels=<int> The expected number of channels for a given session. | 1, 2 | Required |
Channel Index | &channel_index=<int> Used to indicate the channel associated with the client connection. | 0, 1 | Required |
Participant Role | &participant_role=<string> The role label of the participant on the channel. This label is used to identify the participant's role when processing with Generative AI. Examples: Agent, Customer, Salesperson, Client | Descriptive string | Required |
Participant Name | &participant_name=<string> The name of the participant on the channel. Reserved for future use. | String | Optional |
Codec | &codec=<string> The codec of the audio being transmitted. The codec input is case insensitive. | PCM, G711U, G711A | Required |
Sample Rate | &sample_rate=<integer> Sample rate of the audio being transmitted. | Maximum value: 48000 | Required |
Bit Depth | &bit_depth=<integer> Bit depth of the audio being transmitted in bits. | 8, 16, 24 or 32 | Optional |
Bit Rate | &bit_rate=<integer> Audio bit rate of the audio being transmitted in bits per second (bps). | Any positive integer | Optional |
// Communication
Communications between the Client Application and the ElevateAI Real-Time API occur via messages.
- The Client Application sends request messages
- The ElevateAI Real-Time API sends response messages
The ElevateAI Real-Time API processes the Client Application requests asyncronously.
As soon as the Real-Time API has a response for the Client Application, it sends it along the WebSocket – meaning that a response will not necessarily be sent/received in response to the most recent request.
Sequence Diagram
The following sequence diagram outlines the order and flow of the communication between the Client Application and ElevateAI Real-Time API.
NOTE: The diagram is exemplary and is not meant to indicate all possible scenarios.
sequenceDiagram
autonumber
Client Application->>+ElevateAI RT API: Request: establish WebSocket channel 0,1
ElevateAI RT API -->>Client Application: Response: channelConnected 0,1
ElevateAI RT API -->>Client Application: Response: sessionStarted
Client Application->>ElevateAI RT API : Binary audio
Client Application->>ElevateAI RT API : Request: send metadata (optional)
Client Application->>ElevateAI RT API : Binary audio
loop Every 30 seconds
ElevateAI RT API-->>Client Application: Response: AutoSummary & Sentiment
end
Client Application->>ElevateAI RT API : Binary audio
Client Application->>ElevateAI RT API : Request: sessionEnd
ElevateAI RT API -->>Client Application: Response: sessionEnding
ElevateAI RT API -->>Client Application: Response: sessionEnded
Sequence # | Format | Description |
|---|---|---|
1 | | The initial request to establish a WebSocket connection with the Real-Time API. This is done once per channel – either 1 or 2. Query parameters should be adjusted in the URI to indicate the details of each channel. |
2 | TEXT / JSON | The Real-Time API responds after a WebSocket connection request with a response indicating the connection has been established. This is sent after each connection to the connected channels at that time. |
3 | TEXT / JSON | Once all channels for a session – 1 or 2 – have been established, the Real-Time API responds with a sessionStarted response, which includes the assigned interactionIdentifier. |
4, 6, 8 | BINARY | The Client sends raw binary audio data to the Real-Time API via the WebSocket in binary form. Data transmission continues throughout the duration of the session. |
5 | TEXT / JSON | If desired,metadata can be submitted at any point after the session is established and before the sessionEnd request is sent. Metadata is retained when the session is migrated to the Post-Call API environment after Real-Time processing completes. This can be sent on any session channel. |
7 | TEXT / JSON | Approximately every 30 seconds throughout the call, the Real-Time API will provide an AutoSummary of the interaction up to that point. The Real-Time API will also generate a sentiment score across the interaction, indicating negative, neutral or positive sentiment. A brief description of why the sentiment score was received is also included. |
9 | TEXT / JSON | When the client is ready to terminate the transcription session, the sessionEnd request is submitted. This can be sent on any session channel. |
10 | TEXT / JSON | In response to the sessionEnd request, the Real-Time API will reply with a sessionEnding response. |
11 | TEXT / JSON | Once the final AutoSummary is complete, the Real-Time API will send a sessionEnded response. This response will include the transcript and AutoSummary of the complete session. |
// Request Messages
Sending Audio
Send the audio data as a binary message across the WebSocket.
The encoding and bit rate of the binary data must match the specifics included in the query parameters of the URI used to connect to the Real-Time API.
- Typical blocks of audio are 20 - 250ms in duration.
All other request messages are of type text in JSON format.
Metadata
Send this message as a text message in the format below once the session has been connected and initialized.
- Provided metadata is stored with the session and ingested into our post-call system after the session is finalized for use in ElevateAI Explore.
- Only the first metadata request message received by the Real-Time API is processed.
- Additional metadata requests are ignored.
This message applies at the session level and only needs to be sent on one channel.
{
"type": "interactionMetadata",
"content": {
"category": "Call Centers",
"direction": "inbound",
"recorded": "2023-08-03T15:27:09Z",
"audio": {
"extension": "3",
"DNIS": "dnis123"
},
"agent": {
"id": "716",
"name": "Stephen Jones",
"supervisor": {
"id": "49",
"name": "Sabrina Jones"
},
"group": {
"id": "93",
"name": "Jones Group"
}
},
"site": {
"id": "7741",
"name": "Atlanta"
},
"customer": {
"id": "844135",
"name": "Xander",
"city": "Los Angeles",
"state": "California"
},
"custom": {
"dateTimes": {
"dateTime1": "2023-08-03T15:27:09Z"
},
"texts": {
"text1": "California Callers"
},
"integers": {
"integer1": 84416
},
"decimals": {
"decimal1": 344.15
}
}
}
}Maintaining a stream during inactivity
When the audio stream is inactive for any event, such as a hold or transfer, holdStart and holdEnd need to be used to keep the interaction active. During this period keepAlive needs to be sent periodically to keep the interaction active.
Hold Start
- To pause activity in an audio stream, use the holdStart function
- keepAliveshould be sent during the hold at the recommended interval.
- Any audio data sent while the holdStart is active will be ignored.
{
"type": "holdStart"
}Keep Alive
Send this message in the format below during periods of no audio to ensure a timeout or disconnect does not occur.
- keepAlive must be used to keep the audio stream from timing out.
- The ElevateAI Real-Time API will timeout when there is no audio data transmitted on a WebSocket for 1 minute.
- Best practice would be to send the keepAlive message every 30 seconds.
- It may be desireable to send the keepAlivemessage more frequently to ensures that network components along the route do not timeout due to inactivity.
- The keepAlivemessage applies at the channel level.
{
"type": "keepAlive"
}Hold End
Send this message in the format below to resume sending audio data in the interaction.
- This message should only be used in conjuction with holdStart .
{
"type": "holdEnd"
}Session End
Send this message as a text message in the format below during the session to initiate the session end procedure.
- ElevateAI will respond to the client with a sessionEnding response indicating we are finalizing the session.
- ElevateAI will finish transcribing any audio sent thus far, gather final transcript and summary, and send the client a sessionEnded response message.
- This message applies at the session level and only needs to be sent on one channel.
{
"type": "sessionEnd"
}// Response Messages
The response messages are the means by which the ElevateAI Real-Time API sends information to the Client Application.
The Client Application should have a listener on the WebSocket with a message handler for these various message response types.
Channel Connected
This message is sent as soon as the parameters are validated and the connection is established.
- This response messages is sent on all connected channels of a session.
{
"type": "channelConnected",
"content": {
"channelIndex": "0",
"connected": "2025-04-23T06:42:45.3914884Z"
}
}Session Started
The session started message is sent to the client immediately the client connects all expected channels to the ElevateAI Real-Time API.
- All channels for a session must be connected before this message will be sent.
- This response messages is sent on all channels of a session.
{
"type": "sessionStarted",
"content": {
"interactionIdentifier": "e7b61ef8-4f92-4fc4-a75d-d1b6309f2d83",
"started": "2025-04-23T06:42:45.3914884Z"
}
}Transcript Results
These results messages are sent as a text JSON message via the WebSocket as soon as the results have been processed by the ElevateAI Real-Time API.
- This response messages is sent on all channels of a session.
{
"type": "punctuatedTranscript",
"content": {
"sentenceSegments": [
{
"participant": "participantOne",
"startTimeOffset": 3397,
"endTimeOffset": 8487,
"phrase": "Thank you for calling the customer service help line. How may I be of service?",
"score": 1
}
],
"redactionSegments": []
}
}Generative AI AutoSummary
The Generative AI-powered AutoSummary results are returned every 30 seconds once the transcription has begun.
- This response message is sent across all channels of a session.
{
"type": "summary",
"content": {
"summary": "The AutoSummary of the call to this point appears here.",
"sentimentScore": "Positive/Neutral/Negative",
"sentimentReason": "Brief reason for the assigned sentiment score."
}
}Session Ending
The sessionEnding message is sent as a response to the sessionEnd request from the Client Application.
The response indicates that the audio streaming and transcription are terminating for the session – and indicates that the ElevateAI Real-Time API is starting the AutoSummary of the final transcript.
Session Ended
The session finalized message is sent as the final message before the ElevateAI Real-Time API closes the websocket connection.
- The message is in response to the SessionEnd message and contains the full final transcript and AutoSummary.
- This response messages is sent on all channels of a session.
{
"type": "sessionEnded",
"content": {
"interactionIdentifier": "e7b61ef8-4f92-4fc4-a75d-d1b6309f2d83",
"started": "2025-04-23T06:42:45.3914884Z",
"ended": "2025-04-23T06:42:45.3914884Z",
"durationInMs": 19832,
"punctuatedTranscript": {
"sentenceSegments": [
{
"participant": "participantOne",
"startTimeOffset": 3397,
"endTimeOffset": 8487,
"phrase": "Thank you for calling the customer service help line. How may I be of service?",
"score": 1
},
{
"participant": "participantTwo",
"startTimeOffset": 9538,
"endTimeOffset": 14487,
"phrase": "Hello, I would like to change the features of my account.",
"score": 1
},
{
"participant": "participantOne",
"startTimeOffset": 15203,
"endTimeOffset": 19784,
"phrase": "I would be happy to review those options with you.",
"score": 1
}
],
"redactionSegments": []
},
"summary": "The final AutoSummary of the entire call appears here."
}
}Need more help? Contact the ElevateAI Support team.