Voice Cloning

Create a high-fidelity digital twin of any voice from just a short audio sample.

How it works

VoiceForge uses a zero-shot voice cloning technique. It extracts a "speaker embedding" (a 192-dimensional vector) from your reference audio and conditions the text-to-speech model on this vector.

Audio Sample (WAV/MP3) → [Speaker Encoder] → Speaker Embedding → [TTS Engine] → Cloned Speech

Creating a Voice

  1. Navigate to the Voice Lab tab.
  2. Click New Voice.
  3. Upload an audio file (details below) or record directly.
  4. Name your voice and click Clone Voice.

Best Practices

  • • Use high-quality audio (no background noise).
  • • A 30-60 second sample is usually sufficient.
  • • Ensure the speaker is speaking clearly and naturally.
  • • Mono WAV files at 22050Hz or 24000Hz are ideal.

Avoid

  • • Audio with music or heavy reverb.
  • • Multiple speakers in the same clip.
  • • Extremely short clips (under 5 seconds).
  • • Overly processed audio (heavy compression).