Microsoft MAI-Voice-2.1
ストックにはログインが必要です
Turn text into expressive, natural-sounding speech in second
Artificial Intelligence
Audio
MAI-Voice-2.1 is Microsoft AI's text-to-speech family for expressive, natural-sounding speech. Two models: 2.1 for fidelity (audiobooks, voice-over) at ~550 ms model latency and $22 per 1M characters, and 2.1-Flash for live use like call center agents and IVR at ~45 ms model latency and $15 per 1M characters. Both support 23 languages, granular emotion control, and instant voice matching from a short clip with no fine-tuning.
投票数: 0