
Europe Wants Its Voice Back: Tech Fund Moves on ElevenLabs
Cass Iverson · 1h ago · 4 min read
Respeecher laid down the law on voice training data. If your recording space sounds like a tiled bathroom, your synthetic model will sound like a broken appliance.

By Cass Iverson, Synthetic voice correspondent · Reported off Respeecher
Audio engineer turned reporter. Hears the artefacts you do not.

Human hearing lies to you. Your brain ignores the hum of an air conditioner, the whine of a laptop fan, and the slap of sound off a bare drywall ceiling. An artificial intelligence model does not forgive any of it. On September 3, 2026, synthetic speech firm Respeecher published an engineering checklist that stripped the glamour out of AI voice creation and laid the blame for defective clones squarely where it belongs: on lazy audio engineers.
The Ukrainian voice cloning company detailed exact studio requirements for anyone attempting to train speech engines. The baseline specifications ditch standard studio post-production completely. Engineers must deliver raw, uncompressed 24-bit, 48 kHz WAV or FLAC files with no EQ, no dynamic compression, and no noise gates. Peak levels must sit between -12 dBFS and -6 dBFS. Push past that ceiling into digital clipping, and the take is dead on arrival.
Time requirements split by architecture. A standard text-to-speech model demands at least one full hour of spotless, rhythmically varied audio across narration, dialogue, and technical terminology. Speech-to-speech conversion pipelines can function on 20 to 30 minutes of clean source performance, but they require a separate pitch calibration file to map the vocal actor's fundamental frequency. Any studio tracking over multiple days must keep the identical room, hardware gain, and mic capsule position to prevent spectral drift.
``` Target Peak Levels: -12 dBFS to -6 dBFS Master Export Format: 24-bit / 48 kHz uncompressed (WAV or FLAC) TTS Raw Audio Volume: 60+ minutes STS Raw Audio Volume: 20 to 30 minutes Microphone Distance: 12 inches (approx. 1 foot) fixed on a stand ```
The model does not understand ambience. It ingests the HVAC compressor, the room flutter, and the microphone pre-amp noise, then bakes them directly into the synthetic vocal chords.
Respeecher has built a reputation on salvaging broken audio. Their team of more than 15 sound professionals previously extracted voice models from degraded historical sources. They rebuilt the voice of Puerto Rican television icon Tommy Muñiz, resurrected basketball legend Wilt Chamberlain, and pulled usable vocal prints from the fragile, century-old wax cylinders recovered from Ernest Shackleton's Endurance expedition. Those projects proved that machine learning can rescue degraded mono recordings from narrow bandwidths when history demands it.
Yet historical salvage is an emergency procedure, not an enterprise business model. Salvaging ancient tape requires proprietary noise isolation, manual spectral repair, and long model training iterations. Commercial synthetic voice production, tailored for real-time text-to-speech APIs and Pro Tools plugin environments, operates under the opposite premise: speed, consistency, and scale. Feeding broken audio into automated synthetic speech pipelines generates models riddled with phasing artifacts, robotic flutter, and smeared consonants.
Standard commercial voice-over protocols routinely sabotage voice models before the files ever reach the training cluster. Traditional broadcast tracking uses aggressive hardware limiters, high-pass filters, and close-proximity vocal booth techniques to cut through a busy music bed. For training a neural acoustic model, those standard radio tricks act like structural damage. Filtering low frequencies discards vocal formants that the algorithm relies on to map chest resonance, while aggressive compression flattens dynamic range until the neural network cannot tell a whisper from a shout.
Respeecher's production mandates directly echo the internal recording doctrine published by Microsoft for its custom neural voice platform. Both engineering groups tell producers to abandon acoustic vanity setups in favour of dead, non-reflective boxes:
Audio producers entering the synthetic voice pipeline often protest the demand for raw audio. Studio engineers spend decades mastering dynamic chain insertion, running expensive analog channel strips, and dialling in surgical equalization curves to make voices sound larger than life. Respeecher's team explicitly demands that engineers park their creative ego at the door.
Machine learning researchers maintain that pre-processing alters the harmonic integrity of human speech. When a sound engineer activates a noise gate, it clips subtle room reflections and natural breathing decays, forcing the neural network to interpret the sudden drop in noise floor as a bizarre vocal artifact. When an engineer adds algorithmic reverb to sweeten a voice actor's dry take, the network learns the room's reverberation profile as an intrinsic physical component of the performer's throat.
Production contracts for voice actors and commercial recording studios are facing an aggressive overhaul. Traditional voice-over riders that govern hourly billing, session retakes, and post-production delivery formats fail when applied to synthetic asset generation. Agencies building enterprise voice clones must now mandate rigid capture protocols inside talent agreements to avoid burning capital on ruined neural runs.
The operational bottlenecks will move from algorithm design to vocal director competence. Recording one hour of usable clean speech requires capturing multiple emotional registers without allowing the performer to lapse into vocal fatigue or uncharacteristic pitch drift. Directors must explicitly force performers through diverse linguistic gymnastics:
As synthetic voice engines move from slow offline batch processing into sub-second conversational voice bots, input hygiene defines the final latency. Models trained on clean, raw, isolated acoustics converge faster, consume fewer computational training epochs, and respond reliably when deployed across edge networks. Teams that continue to record training datasets through cheap USB headsets in tiled conference rooms will discover that algorithmic compute cannot clean up bad audio discipline.
Voice cloning failures are rarely algorithmic; they are acoustic. Studios that bypass recording fundamentals will burn thousands in compute on unusable, buzzy models that fall apart in production. Clean data is the only moat that matters.
Reported off Respeecher. Original reporting and analysis by Cass Iverson for Vox Roboti.
You must record in raw, uncompressed 24-bit, 48 kHz WAV or FLAC. Never submit MP3s, AACs, or lossy formats for model training.
Full text-to-speech engines need at least one hour of clean, diverse speech. Speech-to-speech systems can work with 20 to 30 minutes.
Target peak levels between -12 dBFS and -6 dBFS. This preserves essential dynamic headroom and prevents destructive digital clipping.
No. Never apply EQ, compression, reverb, or noise gates before training. Raw signals preserve the actual harmonic structure the model needs.
Position the performer roughly one foot from the microphone capsule. Keep the microphone on a locked stand to prevent distance changes.
No. Machine learning models treat room tone, HVAC hum, and acoustic flutter as intrinsic parts of the speaker's vocal profile.

Cass Iverson · 1h ago · 4 min read

Devon Achebe · 2d ago · 5 min read

Marla Quinn · 3d ago · 5 min read