CrePal Releases Workflow Guide for AI Video Voiceovers
CrePal has published a detailed technical guide for utilizing the open-source Breeze TTS 2 text-to-speech model to generate AI video voiceovers locally.

CrePal has released a comprehensive technical guide detailing the workflow for using the open-source Breeze TTS 2 text-to-speech model to create AI-generated voiceovers for video. The guide specifically focuses on a pipeline designed for local audio production, eliminating reliance on external services.
Authored by Marcus Vance, Lead Audio Pipeline Engineer at CrePal's Video Systems Architecture Lab, the documentation outlines the process from script sanitization to aligning the final audio with an editorial timeline. It highlights the advantages of self-hosting, such as removing recurring API overhead, network latency, and third-party data retention concerns.
The guide also addresses the challenges inherent in local hardware processing, including potential prosodic errors from unformatted text prompts, incompatible raw PCM outputs, and subtle cadence drifts that can disrupt visual editing. The Breeze TTS 2 model is noted for its expressive delivery capabilities, supporting both reference-free and reference-guided voice design.
The workflow involves script sanitization to remove visual cues and scene directions, phonetic respelling for abbreviations and brand names, and pacing segmentation to ensure natural speech rhythms. CrePal emphasizes that meticulous input preparation is crucial for achieving high-integrity audio.
This guide provides professionals with a detailed resource for streamlining AI-powered voiceover production within their own environments, underscoring the utility of open-source models and local processing in video post-production.