Alibaba cut the cost of building with voice AI by up to 95%
Alibaba shipped Qwen-Audio-3.1 on 24th Sep'26, a five-model voice stack covering speech recognition, text-to-speech and real-time conversation, plus two new models for audio creation and speaker identification. Alongside it, Alibaba slashed prices across the lineup. Its realtime voice API drops by roughly 85% and speech recognition (ASR) falls by up to 95%, with text-to-speech about 70% cheaper too. The new ASR-Next model can also identify individual speakers with timestamps and detect emotion and background noise.
What it actually means
If you've priced out a voice feature and shelved it because the per-minute cost didn't work, price it again. A 70 to 95% cut on the pieces that make up most voice-app bills, transcription and speech synthesis, changes which products are worth building, not just which are cheaper to run.
The catch is the usual one with Chinese model releases: the cheapest path to using this stack in the West is renting it through Alibaba Cloud or a reseller, which removes an old cost problem while adding a new vendor dependency. Test it against whatever you're paying today for ASR and TTS specifically, not the bundle, since that's where most voice spend actually sits.
The cut lands on exactly the two things a voice app actually pays for.
95%Alibaba's price cut on Qwen Audio speech recognition
Source: The Decoder · Thursday 24 September 2026