A good voiceover takes a script, a voice that fits the channel, and a lot of small fixes: a name read wrong, a pause that is too short, a line you want to say differently after you hear it. Then you still need captions that match what was said. Most voice tools charge a separate monthly fee for this, and many stop at the audio file.
Voice Studio is a tab inside AM Jarvis YouTube Studio. You paste a script, pick a voice, fix what sounds wrong line by line, and download the voiceover with captions timed to the real audio. The voices run on the AM Jarvis server, so there is no ElevenLabs account to connect and no extra fee on top of your plan. Your plan includes a number of voiceover minutes each month.
What you get
- Library voices: 52 ready voices in 9 languages, from the open Kokoro voice model.
- My Voice: a copy of your own voice, made from 1 to 2 minutes of recordings, after a consent check.
- A line editor: preview any line, set the pause after it, and remake only the lines you change.
- A pronunciation list: tell the voice how to say brand names, people's names and short forms.
- Files: MP3 and WAV audio, plus SRT and VTT caption files timed from the real length of each line.
- Clean audio: upload a recording you made yourself to reduce steady background noise, trim silence and set the loudness YouTube plays at. It is free and does not use your minutes.
How it works
- Add your script. Paste it, or send one from the AI writer in YouTube Studio. Each line of the script becomes one part of the voiceover, and a blank line between paragraphs gives a longer pause.
- Pick a voice. Filter the library by language, accent and gender, play the sample, and choose one. Or pick your own voice if you made one in My Voice.
- Check the lines. Preview any line, change the pause after it, and add words to the pronunciation list if the voice says them wrong.
- Make the voiceover. Before you start, Voice Studio shows how long the audio will be, how many of your minutes it uses and roughly how long it will take. You can leave the page while it works.
- Download. Get the MP3, the WAV, and captions as SRT or VTT. If one line needs a fix, change it and apply. Only the changed lines are made again.
The voice library
The library voices come from Kokoro, an open voice model released under the Apache 2.0 license. AM Jarvis gives each voice a plain name and a short style note, such as "steady narrator", "warm" or "bright and lively", so you can pick by sound instead of by code. Every voice has a sample you can play before you choose it.
| Language | Voices | Notes |
|---|---|---|
| English (US) | 20 | Includes narrator, calm, upbeat and deep voices |
| English (UK) | 8 | Includes a classic narrator and a storyteller |
| Hindi | 4 | |
| Spanish | 2 | Spain accent |
| Italian | 2 | |
| Portuguese | 2 | Brazil accent |
| French | 1 | |
| Japanese | 5 | Beta: pronunciation is more basic |
| Mandarin | 8 | Beta: pronunciation is more basic |
A few voices are marked as top picks. They are the ones with the best grades in the model's own voice notes, and they are a good place to start. You can set the reading speed from 0.6 to 1.6 times normal for library voices.
The line editor
A script is split into lines, and each line is made on its own and then joined. That is what makes small fixes cheap. If the voice rushes one sentence, you raise the pause after it. If a line reads flat, you rewrite it. When you apply the change, only that line is made again, and pause changes only rejoin the audio without making anything new.
- Preview any line before you make the whole voiceover, so you hear the voice on your real words first.
- Pauses after each line can be set from none up to 2 seconds.
- Limits per voiceover: up to 300 lines, 500 characters per line and 30,000 characters in total, which is about 35 minutes of audio. Longer scripts can be split into two jobs.
The finished audio is set to -14 LUFS, the loudness level YouTube plays at, so your voiceover does not come out much quieter or louder than other videos.
A pronunciation list for names and brands
Every voice model trips over some words. Brand names, product codes, place names and people's names are the usual problems. In the pronunciation list you write a word the way it should sound, for example "AMJ" as "A M Jay" or "Nguyen" as "Win". The list holds up to 200 words and is used for every voiceover you make.
Your captions keep the original spelling. The voice says "A M Jay", but the caption still shows "AMJ", which is what viewers expect to read.
Captions that match the audio
Captions made from a script usually drift, because a tool guesses how long each sentence takes. Voice Studio times the captions from the real length of every line it made, plus the pauses you set. Each caption shows up to 42 characters a line and 2 lines at a time, which keeps them readable on a phone.
Upload the SRT file in YouTube Studio under Subtitles. Captions help viewers who watch without sound, and the Video analyzer in AM Jarvis YouTube Studio lists uploaded captions as one of its SEO checks. If you publish on other platforms, the VTT file works in most web video players.
My Voice: a copy of your own voice
If your channel is built on your own voice, a library voice can sound wrong next to your other videos. My Voice makes a copy of your voice, so you can make voiceovers from a script on days you cannot record. It works like this:
- Record samples. Read the lines shown on screen in your normal YouTube voice, 15 to 25 cm from the microphone, in a quiet room. One to two minutes in total works best, 40 seconds is the minimum, and you can record in up to 8 takes. Voice Studio warns you if the room is noisy or the level is too low or too loud.
- Read the consent sentence. It includes your name, the date and a one-time 4 digit code, for example "I, [your name], am creating a copy of my own voice on AM Jarvis today". The sentence works for 30 minutes.
- Confirm and create. Tick "This is my own voice". AM Jarvis transcribes the consent recording to check the words and the code, and compares it with your samples to check it is the same voice. A voice that fails the check 5 times has to be deleted and started again.
Once the voice is ready, you pick it in the voiceover editor like any library voice. Two sliders let you tune it: expressiveness (how much emotion goes into the voice) and pace (how closely it follows your recorded style, which also slows it down or speeds it up). My Voice reads English scripts.
Your own voice only
You can clone only your own voice. The consent sentence with a one-time code is there to stop anyone from cloning a voice from a video or a call. AM Jarvis keeps a record of each check: the consent recording, the transcript, the scores and the times. Every My Voice clip also carries an inaudible watermark (Perth, from Resemble AI) that marks it as AI generated. You can delete your voice and all its recordings at any time. The record of the consent check is kept after a delete, because it shows the voice was made with permission.
How long it takes
Voice Studio runs on the server's processor, not on rented graphics cards. That keeps it inside your plan with no extra fees, but it has an honest cost in speed:
- Library voices are quick. A short line is usually ready in seconds, and a full voiceover in minutes, depending on its length.
- My Voice is much slower. Expect several minutes of work for each minute of audio.
Before you start, Voice Studio shows an estimate based on recent jobs on the server, and while a job runs the Jobs list shows real progress and the time left. One job is made at a time, so a busy queue adds a wait, which the estimate includes. You can close the page and come back.
Clean audio for recordings you made
Sometimes you do not need a new voice. You recorded the voiceover yourself, and it has a fan humming in the background or long silences at the start. Clean audio fixes that:
- Upload MP3, WAV, M4A, OGG or WEBM, up to 30 minutes and 100 MB.
- It removes low rumble, reduces steady noise such as fans and hiss, trims silence at the start and end, and sets the loudness to -14 LUFS.
- An option also shortens long pauses (gaps longer than about a second) inside the recording.
- Download the clean MP3 or WAV. Cleaning is free and does not use your voiceover minutes.
It works best on steady noise. It cannot remove music, traffic that comes and goes, or other people talking.
Do you need to tell YouTube the voice is AI?
Sometimes. YouTube asks creators to disclose realistic content that was made or changed with AI, for example a voice that sounds like a real person saying something they did not say. You do that with the altered or synthetic content setting when you upload. A plain narration voice reading your own script for your own video usually does not need it. A cloned voice of a real person, presented as that person speaking, is the kind of case the rule is about. Check YouTube's current policy for your video, and when in doubt, disclose. This is general information, not legal advice.
Where your scripts come from
You can paste any script. Two other parts of AM Jarvis can also send one straight to Voice Studio:
- AI writer in YouTube Studio: the Script outline has a Send to Voice Studio button that turns the hook, sections and call to action into voiceover text.
- Ad Creative Studio in the Dropshipping Suite: the UGC video scripts it writes have a Make the voiceover button.
Both open Voice Studio with the script loaded, ready for you to pick a voice.
Plans and minutes
| Plan | Voiceover minutes a month | My Voice voices |
|---|---|---|
| Free | 3 | Not included |
| Starter | 30 | 1 |
| Pro | 120 | 3 |
| Agency | 500 | 10 |
Minutes are charged by the real length of the audio. My Voice minutes count double, because they take far more work to make. Jobs still in the queue hold their estimated minutes, so you cannot queue past your limit by accident. Clean audio is free. Voiceovers and cleaned files stay in your Jobs list for 30 days, so download what you want to keep. See the pricing page for everything else each plan includes.
What it does not do
- It does not clone anyone else's voice, and there is no way to skip the consent check.
- My Voice reads English scripts only. The library covers 9 languages, and Japanese and Mandarin are in beta.
- It does not edit video. You add the audio and captions in your video editor or in YouTube.
- Clean audio does not separate voices from music or remove other speakers.
- It is not instant for My Voice. Plan for several minutes of work per minute of audio.
Free tools that help
To hear a script read aloud in your browser before you spend any minutes, use the free text to speech tool, and time it with the word counter (about 150 words is one minute of speech). Once the voiceover is done, the YouTube title generator and the YouTube description generator help you package the upload, and the audiogram maker turns a short audio clip into a waveform video for Shorts and Reels.