What Is Text-to-Speech?
Text-to-speech (TTS) is the technology that converts written text into spoken audio. It is used in a wide range of applications, from accessibility tools that help visually impaired and low vision users, to navigation systems, virtual assistants, content creation, telephony systems, and interactive voice response (IVR) platforms.
Over the past several decades, TTS technology has evolved dramatically. Early systems used robotic-sounding formant synthesis, which generated speech entirely from mathematical rules. Modern approaches, like the unit selection synthesis that powers Voice Forge, produce remarkably natural speech by using recordings of real human voices.
The Technology Behind Voice Forge
Voice Forge is powered by Cepstral's speech synthesis engine, a technology rooted in pioneering research at Carnegie Mellon University. Cepstral was founded in 2000 by CMU researchers Alan W. Black and Kevin Lenzo, who had been at the forefront of speech synthesis research. Their work led to the development of a high-quality, efficient TTS system that could produce natural-sounding voices with personality and style.
The core of this technology is called unit selection synthesis, a concatenative approach that is widely considered one of the most natural-sounding methods of speech generation. Rather than generating audio from scratch using algorithms, unit selection stitches together small segments of pre-recorded human speech to create new utterances.
This is the approach used by Voice Forge. By maintaining a large database with many instances of each sound recorded in different contexts, the system can select the version that best fits the current sentence. This produces highly natural speech because the output consists almost entirely of real human speech recordings. The speaker's identity, personality, and vocal characteristics are preserved, making each voice distinctive and engaging.
The Voice Database
The quality of a unit selection TTS system depends heavily on the size and diversity of its speech database. Creating a voice requires a professional voice actor to record several hours of carefully scripted speech in a controlled studio environment. The scripts are designed to cover a wide range of phonetic contexts, ensuring the database contains units for virtually every sound combination that could occur in natural speech.
Voice Forge offers 42 unique voices, each with its own distinct personality and character. Every voice was recorded by a different speaker, giving each one a unique timbre, speaking style, and personality. This diversity allows content creators to find the perfect voice for their specific application, whether that is a warm, conversational tone for a podcast, a clear and authoritative delivery for a presentation, or a fun, character-driven voice for creative projects.
How Unit Selection Synthesis Works
Unit selection synthesis works through a multi-stage pipeline that transforms your written text into natural-sounding speech. Here is a step-by-step look at what happens when you type text into Voice Forge and click generate:
Step 1: Text Analysis
The system first analyzes your text to understand its linguistic structure. This involves breaking the text into sentences, words, and phonemes (the individual sounds of speech). It also determines prosodic targets such as pitch, rhythm, stress patterns, and duration for each sound. For example, a question naturally rises in pitch at the end, while emphasis on certain words changes their loudness and duration. This analysis creates a detailed "recipe" for how the final speech should sound.
Step 2: Searching the Voice Database
Each voice in Voice Forge's library is backed by a large database of recorded speech from a real human speaker. These recordings have been carefully segmented into small units, typically individual phonemes or pairs of phonemes called diphones. Each unit is annotated with acoustic information including its pitch, energy, spectral characteristics, and the phonetic context in which it was originally spoken. When you generate speech, the engine searches through thousands of these recorded units to find the best candidates that match what the text analysis stage specified.
Step 3: Optimal Unit Selection
This is the heart of the technology and where "unit selection" gets its name. The system needs to pick the best sequence of units from the database. For each sound, there may be dozens or even hundreds of candidate units recorded in different contexts. The engine evaluates every candidate using two key cost functions:
- Target Cost – How closely does this unit match the desired pitch, duration, and phonetic context? A unit recorded in a similar context to what is needed will have a lower target cost.
- Join Cost – How smoothly will this unit connect with the units before and after it? This is measured using mel-frequency cepstral coefficients (MFCCs), which capture the spectral characteristics of speech in a way that closely mirrors human hearing. Units that have similar spectral properties at their boundaries join seamlessly with minimal audible artifacts.
The engine uses a Viterbi search algorithm, a form of dynamic programming, to efficiently find the globally optimal path through all possible unit combinations. This ensures the best overall result rather than just picking the best unit at each individual position.
Step 4: Concatenation and Smoothing
Once the optimal sequence of units is selected, they are concatenated (joined end-to-end) to form the complete audio output. Minimal signal processing is applied at the boundaries to smooth any remaining discontinuities in pitch, energy, or spectral content. Because the system uses actual human speech recordings, this post-processing is kept to a minimum, preserving the natural quality, warmth, and personality of the original speaker's voice.
Glossary of Key Terms
Phoneme – The smallest unit of sound in a language. For example, the word "cat" contains three phonemes: /k/, /ae/, and /t/.
Diphone – A speech unit that spans the transition from the middle of one phoneme to the middle of the next. Diphones capture the crucial coarticulation effects that occur between adjacent sounds.
Prosody – The patterns of rhythm, stress, and intonation in speech. Prosody conveys meaning beyond the words themselves, such as whether a sentence is a question or a statement.
MFCC (Mel-Frequency Cepstral Coefficients) – A compact representation of the spectral characteristics of a sound, modeled to approximate human auditory perception. MFCCs are used to measure how smoothly two speech units will join together.
Concatenative Synthesis – Any TTS approach that works by joining together pre-recorded speech segments. Unit selection is the most advanced form of concatenative synthesis.
Viterbi Algorithm – A dynamic programming algorithm used to find the optimal sequence of units from the speech database. It efficiently evaluates all possible combinations to minimize the total cost of the output.
SSML (Speech Synthesis Markup Language) – An XML-based markup language that provides detailed control over speech output, including pronunciation, pauses, emphasis, and audio effects.
Ready to Try It?
Experience the power of unit selection speech synthesis for yourself. Voice Forge gives you access to 42 unique voices, batch processing, and SSML controls.