A recording in, the same recording in another voice out. What Voice Changer actually does, why the target list is shorter than the voice catalogue, and what the three settings are for.
Lesson 1 · start · 5 min read
Open Voice ChangerVoice conversion keeps the performance and replaces the voice. Your words, your timing, your emphasis and your pauses all survive; what changes is who appears to be saying them. It is not an effects box and it does not rewrite what was said — one recording goes in, the same recording in a different voice comes back.
Careful
There are no effects. No robot, no alien, no monster, no preset rack, and nothing that runs live on your microphone — despite what you may have read on our own feature page, which is wrong and is being fixed. There are three settings on this tab and one of them is a pitch slider.
Note
There is no plan gate on this tool. Every account can run it, including a free one — the only limit is credits. Worth saying plainly, because much of the studio is not like that and readers assume a lock that is not there.
One file at a time, dragged in or picked from your machine. There is no recorder on this tab and no way to paste a link, so a phone memo or a call recording has to be a file before it can be converted. Video is accepted as well as audio — WAV, MP3, FLAC, OGG, OPUS, AAC, M4A, WEBM, MP4, MOV, AVI, WMV, MPG, MPEG, 3GP and MKV — and either way what comes back is a single audio file. How large and how long a source file may be depends on your plan, and the uploader enforces your account's real figures whatever the line under the dropzone says. That line is fixed text on this particular tab rather than your plan's ceiling, so treat a rejected upload as the authority and not the label.
Tip
Start with thirty seconds, not with the whole project. You are billed on the length of the source, so the cost of finding out whether a target suits your voice is entirely under your control — 495 credits for a half-minute test against 29,700 for a half-hour episode.
Not every voice you can speak text with can be a conversion target. Conversion needs a stored reference recording of the target, so the picker offers only finished voices that have one — your own cloned voices, and catalogue voices held as real recordings. Voices that exist as a preview rather than a stored clip are perfectly good for reading text and cannot be converted toward, so the list here is shorter than the catalogue and that is not a fault.
If nothing in the list is the voice you want, the answer is to create one: a cloned voice is a conversion target the moment it finishes. Cloning belongs to the Voice Studio track and is taught there rather than repeated here.
Tip
Cloning, directing a voice and reading typed text all live in the Voice Studio track at /guides/voice. This track starts where you already have a recording.
Conversion Settings is the whole of the control surface: Preserve Timing, Quality and Pitch Shift. The next lesson is about when to touch them; on a first run, do not.
Conversion Settings, with the tool's own ranges.
Note
The file type of the result — MP3, WAV or OGG, MP3 unless you change it — is set by the Format menu on the Text to Speech settings pane, not on this tab. It applies to conversions all the same. If you need a WAV, switch tabs, change Format, switch back, then run.
990 credits per minute of source audio, whatever you set the controls to — 990 credits for a minute, 9,900 for ten, 59,400 for an hour. Nothing you set changes that figure, which makes the arithmetic pleasantly simple and makes a failed conversion cost exactly as much as a good one.
Careful
The credit figure appears once the file is loaded, and your browser works it out from the length of the audio rather than asking the server. If it cannot read the length — an unusual container, a video whose audio track it will not decode — the estimate shows nothing and the job still runs at the same rate. That is the one case where you find out the cost afterwards.
The output of a finished job.
Conversions collect in the History tab, newest first, each with its player and its download. The delete control there appears on jobs that did not finish; a completed conversion stays in your history.
Not from memory. You need the two close together in your ears, because what you are judging is a difference and it is easy to talk yourself into one.
Check the obvious things first: every word present, nothing clipped off the front, the timing still yours. That part is usually right and quick to confirm.
Not whether it sounds like the target — whether it sounds like anybody at all. A conversion can be perfectly intelligible and still land in the uncanny middle, and that is the failure this tool actually has. Nothing on screen will mention it: the job says it succeeded either way.
If it held up, you are done: download it, or use function(){throw Error("Attempted to call OPEN_IN_STUDIO_LABEL() from the server but OPEN_IN_STUDIO_LABEL is on the client. It's not possible to invoke a client function from the server, it can only be rendered as a Component or passed to props of a Client Component.")} to take it straight into the editor. If it came back wobbly, smeared or oddly not-anyone, that is not a mistake you made with the settings — it is almost always something about the recording you put in, and the next lesson is about which something.
Questions
Free to start
Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.
1,000 credits every month. No card required.