5–20 second sample
Record in the browser or upload a clear, single-speaker audio clip.
SuraYomi is an online AI voice cloning and text-to-speech studio built around consent. Record your own voice in the browser or upload an authorized sample, add the spoken transcript, and save a reusable voice for future text-to-speech reads.
Record in the browser or upload a clear, single-speaker audio clip.
Select the saved clone, enter new text, and direct its delivery.
Manage each saved reference and delete it from your account.
Open the studio, choose Clone a voice, and either record directly from your microphone or upload an existing clip. Name the voice, choose its language, and enter the exact words spoken in the sample when possible. Save the voice, select it, type a new passage, then generate and download the result as MP3 or WAV.
The saved clone becomes a reusable voice in your account. For later reads, select the same voice and keep your language, pace, expression, performance direction, and seed together when you want a consistent production setup.
Use 5–20 seconds of natural speech with one speaker, little background noise, no music, and minimal echo. Keep a steady distance from the microphone and avoid whispering unless that is the voice you intend to preserve. A clean short sample is more useful than a longer recording with interruptions.
You can record on a supported desktop or mobile browser, or upload WAV, MP3, M4A, MP4, WebM, OGG, or FLAC audio up to 8 MB. Enter the exact transcript when possible so the system can relate the reference audio to the words that were spoken.
After saving the reference, SuraYomi uses the cloned voice as the speaker for new text. Choose Auto or a supported language, adjust pace and expression, describe the intended delivery, and use a seed as part of the reusable settings. Each generated read can be downloaded as MP3 for convenient sharing or WAV for editing.
The studio supports English, Japanese, Chinese, Korean, German, French, Spanish, Italian, Portuguese, and Russian language settings. Output quality can vary by language, sample, text, and direction, so review every generated file before publishing it.
Only clone your own voice or a voice whose speaker has given informed permission for this specific use. Permission should explain what will be generated, where the audio may be used, how long the voice will be retained, and how the speaker can withdraw consent.
Do not clone a public figure, colleague, family member, customer, child, or any other person without appropriate authorization. Voice cloning must never be used for impersonation, fraud, harassment, identity verification, political deception, or fabricated endorsements.
The normalized reference audio and its transcript are stored privately for your account until you delete the saved voice or request account deletion. Generated speech is returned to your browser and is not stored as a permanent audio library by SuraYomi.
Because voice data can be identifying, download and share generated files carefully. Remove a saved voice promptly when permission is withdrawn.
Voice cloning uses 1.5× the standard text-to-speech credit amount, rounded up. For example, a 250-character passage uses 38 cloning credits instead of 25 standard credits, and a 500-character request uses at most 75 cloning credits. The Free plan includes 50 monthly credits, so you can create an account and test the workflow without a payment card.
QUICK ANSWERS
Yes. In SuraYomi, record 5–20 seconds in the browser or upload a clear audio clip, add the transcript, and save the voice to your account.
Use a clear 5–20 second sample with one speaker, little background noise, no music, and minimal echo.
Yes. Select the saved voice, enter new text, adjust the delivery, and generate downloadable MP3 or WAV speech.
SuraYomi accepts WAV, MP3, M4A, MP4, WebM, OGG, and FLAC reference audio up to 8 MB.
Yes. Japanese is available alongside English and other supported language settings, and the full studio also has a Japanese interface.
Only your own voice or a voice for which you have clear, informed permission.
Yes. Deleting it removes the saved reference audio and voice record from the service.
Voice cloning uses 1.5× the standard credit amount, rounded up.
Begin with 50 credits. No card required.