Transcription Features
// Introduction
Accurate transcription is essential for downstream analytics, making speech-to-text (STT) one of the most valuable capabilities in modern contact center solutions. Yet call center audio presents unique challenges – such as background noise, overlapping conversations, and telephony limitations that affect quality.
NiCE ElevateAI addresses these challenges with decades of research and billions of real-world contact center interactions. Our platform delivers high-accuracy conversational transcripts through two purpose-built models:
- CX Model – optimized for contact center environments
- Echo Model – next-generation transcription for global, multilingual use cases
Transcript formats available:
- Phrase-by-phrasePhrase-by-phrase
- Punctuated, sentence-by-sentencePunctuated, sentence-by-sentence
This flexibility ensures seamless integration into analytics, compliance, and customer experience (CX) workflows.
// Transcription Models
What's New
- November 2024 → We introduced Echo, our next-generation transcription model. Built for accuracy at scale, Echo delivers up to 40% higher accuracy than our original CX model, ensuring superior transcription quality across a wide range of use cases.
- July 2025 → We expanded Echo with the launch of Echo Real-Time Transcription, enabling enterprises to capture conversations instantly with the same industry-leading accuracy.
Use our new Fast Start Program to try Echo today – no contract, no credit card required. Explore both real-time and post-call transcription with 10 free interactions per day.
Transcription Models at a Glance
Choose the ElevateAI transcription model that best suits your workflow:
Model | Best For | Key Strengths |
|---|---|---|
CX Model | Contact center environments |
|
Echo Model | Global, multilingual transcription |
|
- When to use CX → Customer service conversations, compliance/QA, and real-time insights monitoring
- When to use Echo → Multilingual transcription, global media (interviews, podcasts, videos), and enterprise workflows at scale
PRO TIP: Many enterprises combine both models – using Echo for fast, accurate real-time transcription across global operations and CX for complex contact center interactions enriched with Enlighten AI and CX AI insights.
Selecting the Echo Model
To use the Echo Transcription model, add the body parameter model to your POST declare and set the value to echo when declaring an interaction.
NOTE: Currently, the Echo model does not perform redaction.
Future releases will continue to move Echo towards feature parity with the CX transcription model; access the latest ElevateAI Release Notes.
Switching Between Echo and CX Models
- Our original, purpose-built CX model can be used by omitting the model parameter or setting it to cx
// PII Redaction
Personally Identifiable Information (PII) such as social security numbers, credit card numbers, and CVV/CVC numbers is automatically redacted before a transcript is stored or returned.
IMPORTANT: PII redaction is available only with the CX transcription model.
The Echo model does not currently support PII redaction, and any media processed with Echo will not be redacted.
When retrieving either a Phrase-by-phrasePhrase-by-phrase or Punctuated TranscriptPunctuated Transcript, details on each redacted segment – including timing, type, and confidence score – are also returned.
Remember: All source data is deleted immediately and promptly upon processing.
Please refer to our Data Retention & Security Guidelines for additional details. uidelines for additional details.
{
...
"redactionSegments": [
{
"startTimeOffset": 1467720,
"endTimeOffset": 1468030
"result": "CVV",
"score": 0.98745
},
...
]
}// Speaker Labels
ElevateAI automatically associates each phrase in a transcript with a participant (up to two speakers) in both Phrase-by-Phrasephrase-by-phrase and Punctuated punctuated transcript transcript formats.
Speaker Diarization
With Speaker Diarization, ElevateAI distinguishes between participants in a conversation – delivering cleaner transcripts and enabling more reliable analytics.
CX Model → Supports automatic speaker diarization for both Mono (single-channel) and Stereo (dual-channel) audio inputs.
Echo Model → Recently expanded to support automatic diarization for Stereo (dual-channel) audio, with additional capabilities in progress to bring feature parity with the CX model.
The system distinguishes between speakers – labeling them as participantOne and participantTwo – and tags each phrase accordingly.
Key benefits of Speaker Diarization include:
- Improved transcript clarity → Each speaker's statements are clearly attributed, making transcripts easier to read and review.
- Simplified post-processing → Accurate labels enable downstream automation for analytics, summarization, and role assignment.
- Enhanced compliance & QA → Clear speaker separation helps meet regulatory requirements and streamlines quality assurance workflows.
NOTE: Diarization currently supports two speakers (participantOne, participantTwo) and is automatically applied when using supported models and audio formats.
Channel Labels
For dual channel audio files – e.g., a phone recording with the agent and customer on separate channels – ElevateAI automatically associates each phrase with the correct channel.
Using stereo recordings does not affect processing speed or fees, making it a flexible option for contact center workflows.
Leveraging stereo recordings does not impact processing speed or fees.
Need more help? Contact the ElevateAI Support team.