Streaming TTS
CosyVoice2
Test
Configuration
Model
CosyVoice2
VoxCPM2 IPA
Inference Mode
EEOS Streaming
Bistream Streaming
Offline
Zero-Shot Clone
Cross-Lingual Clone
Pretrained Speaker
Instruct v1
Instruct v2
Chunked streaming with early exit-of-sequence. Requires a voice profile.
Language
English
Chinese
Mixed (CN+EN)
Chunk Size (k)
Words per chunk
Future Words
Lookahead context
Speed Normalization
Enabled
Disabled
Prevents speed drift
Crossfade (ms)
Overlap between chunks (0 = disabled)
Language
English
Chinese
Malay
Guidance Scale
Higher = more faithful to text
Inference Steps
More steps = better quality, slower
Max Length
Maximum token length
Input
Voice
New Voice
Speaker ID
Built-in speaker from the fine-tuned model
No voice profiles yet
Upload or record a reference audio clip to clone a voice
Create Voice Profile
×
Name
Reference Audio
Drop audio file here or
browse
WAV, MP3, FLAC, WEBM — max 30s recommended
or
Record
Transcription
Transcribe
Must match the reference audio content exactly
Reference Audio (optional, VoxCPM2 voice style)
Drop reference audio or
browse
Separate from prompt audio; used for voice cloning style in VoxCPM2
Text to Synthesize
Load example...
Insert Tag
Events
[breath]
[laughter]
[sigh]
[cough]
[noise]
[lipsmack]
[mn]
[hissing]
[quick_breath]
[clucking]
[vocalized-noise]
[accent]
Style Wraps
<strong>...</strong>
<laughter>...</laughter>
Emotion
future
<|HAPPY|>
<|SAD|>
<|ANGRY|>
<|NEUTRAL|>
IPA Phonemes
Select a word, type IPA, then insert. Uses VoxCPM2 boundary tokens.
<IPA-en>
<IPA-ms>
<PHO-zh>
Insert
ə
ɛ
ɔ
æ
ʃ
ʒ
ŋ
θ
ð
ˈ
ˌ
ɹ
ɑ
ʊ
ɐ
ː
Instruct Text
Guides the speaking style and delivery
Synthesize
Output
No audio generated yet
Download