Docs/Guides/Voice & Phone

    Voice Cloning

    Last updated · MAR 2026·Read as Markdown

    A cloned voice is the difference between an AI that sounds generic and one that sounds like your firm. A managed-services AV company can clone the voice of its lead programmer for a tier-1 troubleshooting line. A higher-ed AV team can clone the campus AV director for a classroom support hotline. AVCodex uses ElevenLabs Instant Voice Cloning. No training phase. Ready in seconds.

    Warning: Voice cloning requires a Studio plan or higher.
    1. Record or upload a short audio sample (1 to 2 minutes).
    2. ElevenLabs analyzes vocal characteristics (pitch, timbre, prosody, accent).
    3. A custom voice is generated instantly.
    4. Select it as your voice agent's voice.

    Cloned voices are shared across all agents in your AVCodex organization.

    1. Open voice settings

    Go to your agent's Build page and open the Voice card. Voice mode must be enabled.

    2. Add custom voice

    Find the Custom Voices section and click Add Custom Voice.

    3. Record or upload audio

    Choose one method.

    Browser recording (recommended):

    • Click Record and speak naturally for 1 to 2 minutes.
    • Re-record if needed.

    File upload:

    • Upload a pre-recorded audio file.
    • Supported formats: WAV, MP3, M4A, OGG, WebM, FLAC.
    • Maximum file size: 50MB.

    4. Name and create

    Give your voice a memorable name ("Service Manager Rachel", "Programming Lead Danny", "Front Desk Persona"). Click Create Voice. The voice is available immediately.

    5. Select the voice

    In the voice selection dropdown, your custom voices appear at the top. Select one to use with the agent.

    The quality of the clone depends entirely on the quality of your sample.

    Duration#

    • Optimal: 1 to 2 minutes.
    • Too short (under 30 seconds): may lack vocal variety.
    • Too long (over 5 minutes): can introduce instability.
    • The AI captures characteristics best from concise, focused samples.

    Environment#

    • Record in a quiet room with soft furnishings. Curtains and carpets reduce echo.
    • Turn off fans, AC, and notifications.
    • Close windows to block outside noise.
    • The AI replicates everything it hears. Background noise becomes part of the voice. A treated voiceover booth or a small carpeted office is ideal. The rack room is not.

    Microphone technique#

    • Position the mic about 20cm away (two fists).
    • Speak slightly off-axis to reduce plosives (hard P's and B's).
    • Use a pop filter if you have one.
    • Avoid breathing directly into the mic.
    • A Shure MV7 or similar dynamic broadcast mic is a good baseline. The conference room ceiling array is not.

    Audio quality#

    • Peak levels: -6dB to -3dB. Loud parts must not clip.
    • Avoid clipping and distortion at all costs. The AI cannot recover from it.
    • Standard sample rate (44.1kHz or 48kHz) works well.

    Delivery#

    • Hold a consistent tone and energy throughout.
    • Don't switch between animated and subdued delivery.
    • Read with natural pacing. Not robotic, not theatrical.
    • Use your own writing for a natural rhythm. A page from one of your SOPs or a section of a programming guide read aloud works well.
    • Delete a voice. Remove it from your organization's library at any time. This also removes it from ElevenLabs.
    • Multiple voices. Create as many custom voices as your plan allows.
    • Cross-agent usage. Every agent in the organization can use any custom voice.
    • Audio privacy. Your recording is stored encrypted and never shared publicly.

    Voice doesn't sound right?

    • Re-record with cleaner audio (less background noise, no clipping).
    • Try a longer sample (aim for 1 to 2 minutes of natural speech).
    • Hold consistent tone throughout the recording.

    Voice not appearing in selection?

    • Confirm creation completed successfully.
    • Refresh the voice settings page.
    • Verify you're on a Studio plan or higher.

    Upload failing?

    • Check file size (max 50MB).
    • Verify the file format (WAV, MP3, M4A, OGG, WebM, FLAC).
    • Try a different format or re-export the audio.
    Note: For voice agent configuration, see the Voice Agents guide.

    *AVCodex · Your AV expertise. Amplified by AI.*

    Was this helpful?
    Edit this page →