Set up your ElevenLabs account
This is the account your voice will live in. It stays yours — you own it, you control it, and you can switch it off at any time.
ElevenLabs is the service that builds and stores the voice model. There is one rule that shapes everything below: a voice clone can only be made by the person whose voice it is. Partway through, ElevenLabs will ask you to record a short verification passage while signed in, and it compares that against your recordings. Cloning someone else's voice is prohibited on the platform even with written permission.
So the account has to be in your name. We work inside it with a key you create and can revoke, which is Step 5.
What to do
- Go to elevenlabs.io/sign-up and create an account in your own name, using a clinic email you will keep access to.
- Open elevenlabs.io/pricing and subscribe to the Creator plan.
Creator — $22/month
| Professional voice clone | 1 — this is the one that matters |
| Credits | 121,000/month, roughly 2 hours of finished voiceover |
| First month | Often 50% off at $11 |
Two hours a month is far more than a monthly video schedule uses. You will not need to think about credits.
- ElevenLabs account created in my name
- Creator plan active on the account
Choose your script topics
Before you record anything, we need to know what the scripts should be about. Tick what fits, cross out what does not, and add anything we have missed. Then we write the scripts and send them to you.
You are not writing anything here — you are pointing. We turn what you pick into about 100 short scripts, each 22 to 40 seconds long. That comes to roughly an hour of speech, which is what makes a voice clone sound like you rather than like a machine doing an impression of you.
What we are suggesting
Anything we have missed
One per line. A few words is plenty — "why people think Botox is one thing when it is really three" is a perfectly good brief.
Saves as you type.
Once you send them, we write your scripts and they appear on the next step. That usually takes us a couple of days.
Review your scripts
Here are your scripts. Read each one and tell us: happy to say it, or not. Edit anything that does not sound like you — your version wins, always.
Two different jobs on this page
| Keep or drop | Everything you keep gets recorded. The volume is what teaches the system your voice, so keep anything you are willing to say, even if it is not your favourite. |
| Star about 10 | Of everything you keep, star the ones you actually want made into videos first. That is where we start producing. |
A script you keep but do not star still earns its place — it trains the voice, and it stays in the pool for later.
Your scripts are not written yet
Record your scripts
You read your scripts once. From that recording we build a model of how you speak, and after that we can produce the voiceover for any script in your own voice without you reading it.
That means no more recording sessions. When you want a new video you send us the idea and we build it. If a word changes, we change it in seconds instead of asking you to sit back down.
You are producing two things from one sitting
| One continuous recording | Press record once and work down the list without stopping. This is what ElevenLabs trains your voice on. |
| One file per script | About 100 separate clips, each named after its script. These are what we dub onto your footage later. |
You do not record twice. You record once, drop a marker as you start each script, and one export gives you all hundred files. The steps are below.
| How much speech | About an hour, set by your script list |
| Absolute minimum | 30 minutes — below this a clone is noticeably worse |
| Time to set aside | 2 to 3 hours, with breaks |
| How often | Once |
That is talking time, not clock time. With pauses, re-reads and a coffee, an afternoon covers it comfortably. You only ever do this once, so it is worth doing properly rather than quickly.
What you will need
| Item | Why |
|---|---|
| Audacity — free, Mac and Windows | The recorder we recommend, and the reason the hundred files are not a hundred jobs: it can split one long recording into separate named files in a single action. Please use this one rather than a substitute — audacityteam.org/download |
| A USB or clip-on microphone | Your laptop mic is not usable for this. Blue Yeti, Audio-Technica AT2020USB, Shure MV7, or any lavalier |
| A quiet room | Background noise gets learned along with your voice |
| A glass of water | You are going to be talking for a while |
Your room
Somewhere quiet. No fan, no air conditioning running, no traffic through an open window, no music. A room with soft furnishings sounds far better than a hard empty one — a walk-in closet with clothes in it genuinely beats a boardroom.
Setting up Audacity
You only do this once.
- Open Audacity. It will ask permission to use your microphone — say yes.
- Audio Setup → Recording Device → choose your microphone by name. Check it is not the built-in one.
- Audio Setup → Recording Channels → 1 (Mono).
- Bottom left of the window → set Project Rate to 48000 Hz.
- Mac only: open Control Centre → Microphone Mode and set it to Standard. On Voice Isolation your Mac processes your voice before Audacity ever hears it, and that hurts the result.
Do a thirty second test first
Record yourself talking normally for thirty seconds, then play it back and listen. Watch the level meter while you talk. Peaks should sit around -12 and must never hit the far right — if they touch the end the recording distorts, and that cannot be fixed afterwards. Move slightly further away. If the bars barely move, you are too quiet: move closer.
Sit about a hand span from the microphone, roughly 15 to 20 cm, and slightly off to one side rather than straight into it. That stops hard P and B sounds popping.
How to record
Press record once, and do not stop. Then, for each script:
- Add a marker. Press
Ctrl+Bon Windows, or⌘+.on Mac. A little box appears. - Type the script ID into it —
S-01— and press Enter. - Say the ID out loud. For example, "S dash oh one."
- Pause one second.
- Read the script.
- Pause two seconds, then move to the next one.
Why the marker and the spoken ID
The typed marker is what splits your recording. When you are finished, Audacity uses those markers to cut the one long file into a hundred separate files and name each one after the marker you typed. That is the whole trick, and it is why we ask you to use Audacity rather than whatever recorder you already have.
Saying the ID out loud on top of that is the safety net. If a marker goes missing or lands in the wrong place, we can still hear where each script starts and fix it ourselves. Belt and braces, and it costs you two seconds.
Say only the ID. Do not read the word "SLATE" out loud.
The rest of the rules
- If you stumble, do not stop the recording. Pause two seconds, say the same ID again, and read from the top. We keep the last clean take and throw the rest away. Stopping and restarting changes your tone, which is worse than any stumble.
- Take breaks whenever you need them. If you are away for more than about five minutes, stop the recording, save that file, and start a fresh one when you come back. Your voice shifts slightly after a break and we would rather keep those separate.
- Read it the way you would say it to a patient sitting across from you. Warm, direct, unhurried. Not a radio announcer, not a presentation voice. Whatever you sound like in this recording is what every future video sounds like, so sound like yourself.
- Say numbers as words. If the script says "three to four months", say "three to four months".
You will probably finish with two to four audio files, one per stretch between breaks. That is normal and correct.
Your scripts
Your scripts are not written yet
Tick each one as you record it, so you can stop for a coffee and know exactly where you were. This is just for you — nothing here is submitted.
Exporting — the important bit
Two exports out of the same recording. Do them in this order.
1. The continuous recording
- File → Export → Export as WAV
- Name it with your name and the date, for example
jane-smith-voice-2026-09-09-01.wav - If you recorded across several sittings, number them 01, 02, 03 in order.
2. The hundred separate files
- File → Export → Export Multiple
- Under Split files based on, choose Labels.
- Under Name files, choose Using Label/Track Name. This is what names each file after the script it holds.
- Leave Include audio before first label unticked.
- Choose an empty folder, pick WAV, and click Export.
You should end up with a folder of files named S-01.wav, S-02.wav
and so on. If the names look right, you are done — that folder is exactly what we need.
If WAV file sizes cause problems, MP3 at the highest quality setting is fine for the hundred clips. Keep the continuous recording as WAV if you can. You upload both on the next step — no Drive links, no email.
Before you finish
- Quiet room, no fan or air conditioning
- Proper microphone, not the laptop
- Audacity set to 48000 Hz, mono
- Thirty second test recorded and listened back, peaks near -12
- Microphone distance set, and nothing moved after that
- Worked through the whole script list, 30 minutes of speech minimum
- ID said out loud before every script
- Kept rolling through mistakes
- Continuous recording exported as WAV, named and numbered
- Export Multiple run, and the file count matches the script count
Upload your recordings
Both sets go here — the continuous recording and the folder of separate clips. They upload directly to our storage, so nothing goes through email and there is no file size limit you need to worry about.
1. The continuous recording
The one long file per sitting — usually two to four of them. This is also the file you upload to ElevenLabs on the next step, so keep it handy.
Drop your continuous recording here
or click to choose · WAV or MP3 · usually 2–4 files
2. The separate clips
Everything that came out of Export Multiple — S-01.wav, S-02.wav and
the rest. Select the whole folder's worth at once and drop them in together.
Drop all your script clips here
or click to choose · about 100 files · select them all at once
Saves automatically. If you skipped the markers and want us to split the recording from your spoken IDs, say so here.
- My continuous recording is uploaded
- My separate script clips are uploaded, or I have asked you to split them
Create your voice clone
This is the part only you can do. ElevenLabs verifies that the voice being cloned is the voice of the person signed in, so it has to happen in your account, from your microphone.
Use the same microphone, same room and same speaking style as your recording session. The verification compares the two, and a different setup is the usual reason it fails.
What to do
- Sign in at elevenlabs.io/app/voice-lab.
- Choose Add Voice → Professional Voice Clone.
- Upload your continuous recordings — the two to four long files, not the hundred separate clips. ElevenLabs learns better from long, unbroken speech, and a hundred short files gives it a hundred starts and stops to average out.
- Name the voice with your own name, so we can find it later.
- Record the verification passage when prompted — same mic, same room, same voice you used for the recording.
- Confirm the consent and training checkboxes and submit.
- Professional Voice Clone submitted with my recordings
- Verification passage recorded and accepted
- ElevenLabs emailed me that the voice is ready
Connect us to your account
So we can produce voiceovers without you, we need a key from your ElevenLabs account. You create it, you can see exactly what it is allowed to do, and you can revoke it in one click.
What this key can and cannot do
| Can | Turn a script into audio using your voice |
| Cannot | Change your password, your billing, or your plan |
| Cannot | Delete your voice or your account |
| You control | A monthly credit ceiling, and revoking it at any time |
Create the key
- Go to elevenlabs.io/app/settings/api-keys.
- Click Create API Key and name it
Health Hueso you can recognise it later. - If it offers you a scope or permissions choice, restrict the key to text to speech and reading voices. That is everything we need and nothing else. If you do not see that option, a default key is fine — it still cannot touch your billing or your password.
- Optionally set a credit limit. We suggest leaving headroom of about half your monthly allowance.
- Copy the key. ElevenLabs shows it once — if you lose it, delete that key and make another.
We check it against ElevenLabs before saving, so you will know immediately if it worked. It is encrypted before it is stored, and it is never shown again — not even to you on this page.
Film yourself
Talk to your phone for about two minutes at a time, 4 to 6 times, changing your outfit between them. That is the whole job.
We already have your voice. From here on we write each script, produce a clean voiceover in your voice, then lay that voiceover over this footage and match it to your mouth. So for this filming:
- We are not using the sound you record on the day — that audio gets thrown away.
- What you say does not matter. There is nothing to memorise.
- We only need your face, your expressions, and your mouth moving continuously. If you go quiet and still for a long pause, that part of the clip cannot be used.
Doing this again every 2 to 3 months in different clothes keeps a year of videos from looking like one afternoon.
The app to use
Download Blackmagic Camera. It is free, it is made by Blackmagic Design — the people who make professional cinema cameras — and it exists on both iPhone and Android. You record with it instead of your normal camera app because it lets you lock the settings, so the picture does not drift halfway through a take. Your built-in camera app is constantly re-deciding brightness and colour, and that flicker is very hard to fix afterwards.
| Phone | What you need |
|---|---|
| iPhone | iOS 16.6 or later · App Store |
| Android | Android 13 or later · Google Play. Supported on Samsung Galaxy S21 and newer, Pixel 6 to 9, OnePlus 11 and 12, Xiaomi 13 and 14, and Sony Xperia 1, 5 and Pro-I |
| iPad | iPadOS 18 or later, though a phone on a tripod is easier to place |
Set your phone up once
| Setting | Set it to |
|---|---|
| Resolution and frame rate | 4K (3840 × 2160) at 30 fps |
| Lens | 1×, the main lens. Never 0.5×, never digital zoom. To get closer, move the tripod |
| Colour | Standard Rec.709 / SDR. Never HDR, never a "log" profile |
| White balance | Manual, 5000 to 5600, then locked |
| Focus and exposure | Lock both on your face before you roll |
| Shutter | 1/60, and keep ISO as low as your light allows |
Set the room up once
- Phone on a tripod at eye level, about 5 feet (1.5 m) away, at 1×. Not looking down at you, not up at you.
- Shoot landscape (16:9) for the master. Keep yourself framed so a vertical 9:16 crop still works — roughly centred, with some headroom above you.
- Keep your eyes about a third of the way down the frame, not right in the middle.
- Main light in front of you, from the left or right at about 45 degrees, slightly above your eyes. A big window facing you works well. Soft light is best.
- Never sit with a window or bright lamp behind you — you will be a silhouette.
- Never use a light straight overhead — it throws shadows under your eyes.
- If one side of your face is too dark, put something white just out of frame on that side to bounce a little light back. A little shadow on one side is good; you are not trying to flatten your face out.
You will be talking direct to camera, so look at the lens. If you ever want an interview-style look instead, look 10 to 15 degrees off to the side — just pick one or the other for a batch and keep it.
Sound
We still want a clean sound track, so some clips can go out exactly as they are, and so we have something to match against.
- Use a lavalier (clip-on) or USB condenser mic, never the phone's built-in mic.
- Aim for peaks around -12 dB, no clipping. If the meter is hitting the top, turn it down.
- If the mic records into a separate device like your Mac, clap once on camera at the start of every take so we can line the two up.
Do a 30 second test
Record 30 seconds, play it back, and check four things:
- Your face is sharp, in focus.
- The brightness holds steady as you move — exposure is not pulsing.
- The audio is clean at around -12 dB, no clipping or distortion.
- There is no hum or buzz.
Fix anything that is off, then do not touch the setup again for the rest of the session.
What to do for each clip
- Talk for two minutes, one continuous take.
- Say anything at all. Explain a treatment the way you would to a patient, talk through your morning, or describe your weekend. Explaining a treatment usually looks best, because your face changes when you are making a point.
- Keep talking. Long silent pauses are the one thing that makes a clip unusable, because we need to see your mouth moving the whole time.
- Look at the lens, not at yourself on the screen.
- Move normally. Use your hands if it helps you talk — just keep them below your chin and do not let them cross your face.
- Change something, then go again. Change your outfit, or stand instead of sit, or move the tripod a little closer or further.
Aim for 4 to 6 clips in a session. More outfits are better if you have the energy, and changing your top is quick. Moving to another room is slower because the light has to be set again, so only switch walls if it is easy.
- Phone locked to 4K/30, 1×, Rec.709, manual white balance
- Tripod at eye level, 5 feet away, light in front at 45°
- 30 second test recorded and checked
- 4 to 6 clips, about two minutes each, different outfits
- Kept talking throughout — no long silent pauses
Upload your footage
Send each clip exactly as it came off the phone. Do not stitch them together, do not trim them, do not edit or compress them. These untouched files are the raw files we need.
Drop your video clips here
or click to choose · MOV or MP4 · usually 4–6 files, one per clip
- All my clips are uploaded and showing as done above
- I sent the original files, untrimmed and uncompressed
Send us your script ideas
This is the part that keeps going. Everything above happens once — this is how you feed the machine from here on.
You do not need to write scripts. You need to tell us what is worth saying. The best ones almost always come from your day: the question you answer for the fifth time this week, the thing patients get wrong, the treatment you wish people understood before they booked.
What makes a good one
- A question you actually get asked. "Will people be able to tell?" beats any topic we could invent.
- A misconception worth correcting. Something patients believe that costs them money or results.
- A judgement only you can make. What you would tell a friend, not what a brochure says.
- A rough note is enough. One line is enough. We write it properly.
Saves automatically. Come back and add to it whenever something comes up — you do not need to wait for us to ask.
Drop any reference material here
or click to choose · optional · video, audio, slides, documents
- I have sent through my first batch of ideas
That is everything
Once your voice clone finishes training and your footage is up, we take it from here. You will see the first videos for review before anything goes out. Anything at all, we are at support@healthhue.com.