AI Voice Cloning: How It Works and Consent Rules (2026)

Vorec Team · 2026-10-02 · About 9 min read

Google's Gemini API asks for a reference clip of 10 to 30 seconds to replicate a voice. Microsoft's personal voice asks for 5 to 90 seconds.

Short samples are why voice cloning comes up for training and tutorial videos. The appeal is simple: record a subject-matter expert once, then narrate later updates in their voice without booking them again. In practice, reuse is limited by how long the vendor keeps the voice, what the speaker agreed to, whether the feature stays available to you, and output quality. All three products covered here also build a consent or verification step into cloning.

This post covers how AI voice cloning works and what Google, ElevenLabs and Microsoft require, using only their own documentation. It also covers a 2024 FCC ruling that names voice cloning, and a practical consent checklist for teams. Vendor terms change often, so check each linked page before you rely on it.

Checked on 2 October 2026. Sources: Google's voice replication guide ("Last updated 2026-09-24 UTC"), ElevenLabs' voice cloning guide (no date shown), Microsoft's personal voice overview (updated 16 September 2026), and the FCC's Declaratory Ruling FCC 24-17 (adopted 2 February 2024, released 8 February 2024) with its news release. This post is not legal advice.

What is AI voice cloning?

AI voice cloning creates a synthetic voice that sounds like a specific real person, from recordings of that person speaking. Once the voice exists, you type text and the model speaks it in that voice.

Vendors use different names for the same idea. Google calls it voice replication, Microsoft calls it personal voice, and ElevenLabs calls it voice cloning. ElevenLabs offers two kinds:

Cloning is different from choosing a preset or stock voice: a ready-made voice you select from a provider's library, rather than one created from a sample of someone on your team. It's also different from voice design, which some vendors offer to generate a new voice from a text description.

Reference audio passes through a voice model to generated speech, with consent required

How much audio does voice cloning need?

Vendor (feature)Audio neededSource wording
Google Gemini API (voice replication)10–30 seconds"A 10–30 second clip of clean, natural speech"
Microsoft Azure (personal voice)5–90 seconds"A 5–90-second human speech sample"
ElevenLabs (Instant Voice Cloning)1–2 minutes"1–2 minutes of good audio"
ElevenLabs (Professional Voice Cloning)30–180 minutes"30–180 minutes of good audio"

The clip lengths are short, but all three vendors stress quality. Google says to record in a quiet room and to record the reference and consent clips "on the same microphone in the same acoustic setting so the speaker verification check succeeds reliably." ElevenLabs recommends MP3 at 192 kbps or higher, and says uncompressed formats such as WAV usually don't improve clone quality.

What consent do the vendors require?

Each of the three products documents a consent or verification step for cloning, and each does it differently. The rules below are what each vendor documents; they sit alongside the vendor's terms, eligibility rules and approved uses.

Google: a spoken consent statement

Every replication request needs two recordings from the same adult speaker: the reference clip and a consent clip. In the consent clip, the speaker reads a fixed statement. The English version is: "I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model." The guide lists the statement in several supported languages.

Storage limits also apply. Stored voices are capped at 200 per project, "shared across prompted and replicated voices", and kept for one year. Stateless voice keys, which your app stores itself, last seven days. Google lists Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS as supporting replication.

Microsoft: a verbal statement and restricted access

Microsoft says "every voice be created with explicit consent from the user." A recorded statement from the speaker must acknowledge that the Azure resource owner will create and use their voice. Both personal voice and professional voice list "Speaker's verbal statement required. No unapproved use case allowed."

Access is gated too. API access is "restricted to eligible customers and approved use cases", and you apply through an intake form. Microsoft says a personal voice can speak in more than 90 languages, across more than 100 locales.

ElevenLabs: your own voice only, for Professional clones

For Professional Voice Cloning, ElevenLabs states: "You can only create a Professional Voice Clone of your own voice. Even with their consent, you cannot clone someone else's voice." Every Professional clone goes through a verification process to confirm the voice belongs to the account holder. If you fail every attempt, you wait 24 hours to retry. If a colleague wants you to use their voice, ElevenLabs says they can create and verify a clone on their own account and share it with you.

ElevenLabs lists Professional Voice Cloning as available "on our Creator plan or above." The account-holder verification described here is documented for Professional clones; we didn't find the same verification documented for Instant Voice Cloning, so don't assume it applies. For current plan prices, see our ElevenLabs pricing breakdown.

Voice consent card covering who can use a voice, its purpose and duration
GoogleMicrosoftElevenLabs (Professional)
Consent mechanismSpeaker reads a fixed consent statementSpeaker records a verbal consent statementVoice verification against the account holder
Clone someone else's voice?Documented, with their recorded consent statement, subject to Google's termsDocumented, with their recorded statement, for approved use cases and eligible customers onlyNo, even with their consent
AccessGemini API, 3.8 TTS modelsApplication and approval required for the APICreator plan or above

What did the FCC rule about cloned voices?

In a 2024 Declaratory Ruling, FCC 24-17 (adopted 2 February, released 8 February 2024), the FCC confirmed that the Telephone Consumer Protection Act's restrictions on "artificial or prerecorded voice" cover "current AI technologies that generate human voices." The ruling names voice cloning directly: such technologies fall within the TCPA because they artificially simulate a human voice.

The practical effect, in the ruling's words: calls using these technologies "require the prior express consent of the called party to initiate such calls absent an emergency purpose or exemption." The ruling adds that artificial or prerecorded voice messages must identify the entity responsible for the call, and that telemarketing messages must offer opt-out methods. The FCC's news release describes the effect as making voice cloning "used in common robocall scams targeting consumers illegal."

The scope is phone calls, with emergency and other exemptions. The ruling doesn't address narration in a training video. State laws and laws outside the US also apply, and they vary. If you plan to clone a voice for public or commercial content, check with a lawyer where you operate.

Should you clone a voice for tutorial narration?

This section is our recommendation, not a legal requirement.

Where cloning helps. You have a well-known presenter, such as a founder, an instructor or a support lead, and viewers expect to hear them. You publish frequent updates and can't keep booking recording sessions.

Where a preset voice is simpler. Many product tutorials and onboarding videos don't need a specific person's voice. Choosing a preset voice means you don't collect a colleague's voice sample or cloning consent for that workflow (the provider's usage terms still apply). It also avoids a hard question: what happens to the narration when that employee leaves?

A consent checklist for teams

If you do clone a colleague's voice, get these points in writing before you record:

  1. Who. The speaker's name, and that they're an adult agreeing for themselves.
  2. What for. Specific uses, such as "internal training videos" or "public product tutorials". Avoid "any purpose".
  3. Which vendor. The vendor's own step (Google's consent statement, Microsoft's verbal statement, ElevenLabs' Professional Voice Clone verification) is still required on top of your agreement.
  4. How long. An end date, or a review date.
  5. Withdrawal. How the speaker can withdraw consent, and what happens to existing videos if they do.
  6. Departure. What happens when the speaker leaves the company.
  7. Disclosure. Whether viewers are told the narration is AI-generated. We'd recommend saying so in the video description.

Where Vorec fits

With Vorec, you choose a preset voice for tutorial narration, so you don't record your own narration into a microphone.

Vorec records your screen with its own macOS recorder, and an AI agent can drive it for you through the Claude Code plugin. Capture runs locally so you can review the take before anything is uploaded. Vorec drafts narration matched to the workflow it captured and generates the voiceover. Edit a line and regenerate the voiceover for just that segment. On eligible plans, narration can be regenerated in supported languages without re-recording. You can also upload an existing recording.

For more on AI voices in tutorials, see how to add an AI voice to a screen recording, how to make a tutorial video without a microphone, and our explainer on Gemini TTS.

FAQ

How does AI voice cloning work?

You give a model recordings of a person speaking, and it creates a synthetic voice that sounds like them. You then type text and the model speaks it in that voice. Some services create the clone almost instantly from a short sample. Others train a dedicated model on much more audio.

How much audio do you need to clone a voice?

It depends on the vendor. Google's Gemini API asks for 10–30 seconds and Microsoft's personal voice 5–90 seconds. ElevenLabs asks for 1–2 minutes for an Instant clone and 30–180 minutes for a Professional clone.

Is AI voice cloning legal?

It depends on where you are and what you use it for. In the US, the FCC's February 2024 ruling (FCC 24-17) confirmed that AI-generated voices in calls count as "artificial" under the TCPA, so such calls need the called party's prior express consent unless an emergency purpose or exemption applies. That ruling covers phone calls, not videos. Other laws vary by state and country. Get legal advice for commercial use.

Can I clone someone else's voice?

Google and Microsoft document cloning another person's voice with that person's recorded consent statement, subject to their terms. Microsoft also limits API access to eligible customers and approved use cases. On ElevenLabs, a Professional clone can only be of your own voice, "even with their consent".

Do I need to clone a voice for tutorial videos?

No. Preset text-to-speech voices are an option for many tutorials and onboarding videos. Cloning makes sense when viewers expect a specific presenter's voice.

Want narrated tutorials without recording a voice? Record with Vorec, or upload one you already have, and get a narrated tutorial plus a written guide. Start free. 7-day trial, 100 credits, no credit card required. Trial includes up to 3 projects; exports carry a watermark. Paid plans start at $9/month.

← Back to blog