Alibaba's Qwen Rolls Out Audio 3.1 and Mobile Agents as Qwen 4 and Custom Silicon Signal Broader Ambitions
Key Takeaways
- •Qwen Audio 3.1 consists of five models: upgraded ASR, TTS, and Realtime systems plus new TTS-Next and ASR-Next models for audio creation and deeper comprehension.
- •The flagship TTS model ranks first on the independent Artificial Analysis Text-to-Speech Leaderboard and supports 16 languages, 20 Chinese dialect regions, and up to three minutes of continuous speech generation.
- •Alibaba reduced speech service prices at launch, cutting TTS costs roughly 70 percent, Realtime around 85 percent, and ASR up to 95 percent.
- •Qwen Intelligence launches with three agents, and the Mobile Use agent scored 82.1 on MobileWorld, 92.2 on MobileWorld Real, and 97.2 on AndroidDaily, with a reported 90 percent end-to-end task success rate.
- •At the Apsara Conference, Alibaba outlined Qwen 4 in four tiers without public specifications, announced plans to train models of five to ten trillion parameters, and unveiled the Zhenwu V900 accelerator with mass production scheduled for the first quarter of 2027.

Alibaba's Qwen team has rolled out two additions to its product lineup: Qwen Audio 3.1, a rebuilt speech stack, and Qwen Intelligence, a suite of mobile agents. Together, they show how the company is converting its model momentum into deployable products across modalities.
Qwen Audio 3.1: Five Models Across a Complete Speech Stack
Qwen Audio 3.1 consists of five models organized into two groups: an upgraded core lineup covering understanding, generation, and interaction, plus two new models aimed at creation and deeper comprehensionn
On the synthesis side, the flagship text-to-speech (TTS) model ranks first on the independent Artificial Analysis Text-to-Speech Leaderboard. It supports 16 languages and 20 Chinese dialect regions, generates up to three minutes of continuous speech in a single pass, and accepts free-form natural language instructions that control emotion, pace, timbre, and accent. It also offers 86 fine-grained inline tags for phrase-level effects such as pauses, breathing, laughter, and sighs.
According to the technical report, a 12.5 Hz low-frame-rate speech tokenizer reduces decoding cost, while a five-stage training pipeline coordinates the language and flow models for content consistency, voice similarity, and robustness. The system can produce usable speech even from noisy, reverberant, or degraded reference audio. A companion model, TTS-Next, pairs a language model with a diffusion framework to generate voice, sound effects, and background audio in a single pass, targeting audiobooks, podcasts, games, and advertising.
On the understanding side, the upgraded automatic speech recognition (ASR) model adds stronger multilingual and dialect recognition, along with native transcript polishing that automatically removes fillers and repetitions. ASR-Next extends this to multi-speaker recognition with speaker labels, timestamps, and aligned transcripts. It can also interpret emotions, ambient noise, and machine sounds, enabling sound captioning, event localization, and audio question answering.
The Realtime model supports simultaneous speaking and listening with interruption at any moment, and adjusts its pace and tone when it detects a low mood in the caller. Alibaba paired the launch with substantial price reductions: TTS costs fell roughly 70 percent, Realtime by around 85 percent, and ASR by up to 95 percent. The reductions cover the full core lineup — synthesis, realtime interaction, and recognition — arriving at launch alongside the new creation- and comprehension-focused models.
Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and ASR-Next for audio understanding. Five models, one complete audio stack: understanding, generation, interaction & creation. Plus big price cuts across the… pic.twitter.com/Qcbvcw0q9C
— Qwen (@Alibaba_Qwen) September 23, 2026 (on X)
Qwen Intelligence: Agents for Mobile Task Execution
Qwen Intelligence targets a different layer entirely: autonomous task execution on mobile devices. Its Planner agent decomposes and orchestrates complex tasks, ranking first on the company's MobilePA Bench, including its business and memory variants.
The Mobile Use agent executes actions primarily through APIs, falling back to graphical interface control when needed. It scored 82.1 on MobileWorld, 92.2 on the real-device MobileWorld Real benchmark, and 97.2 on AndroidDaily, with a reported 90 percent success rate on complete end-to-end tasks. The Creative agent generates a usable image from a single sentence in about three seconds — roughly twice as fast as leading alternatives, according to Alibaba.
The company also opened its benchmark suite, including MobilePA Bench, MobileWorld, MobileWorld Real, and MobileWorld Safety, to the research community — the yardsticks behind several of the reported agent scores — providing external reference points for evaluating mobile agents independently of Alibaba's own reporting.
Introducing Qwen Intelligence, bringing personal intelligence within everyone's reach. It launches with three SOTA agents: – Mobile Planner Agent: plans, decomposes & orchestrates complex tasks. #1 on MobilePA-Bench, MobilePA-Bench Business & Memory. – Mobile-Use Agent:… pic.twitter.com/OgCLpxDJgj
— Qwen (@Alibaba_Qwen) September 23, 2026 (on X)
Following the Conference Stage: Qwen 4 and the Infrastructure Announcements
The releases follow Alibaba's Apsara Conference in Hangzhou on September 22, where the Qwen 4 family was announced in four tiers: Max as the flagship positioned against top rival models, Flash for low-latency and high-volume workloads, Plus as a balanced multimodal tier, and a 27B open-weights variant for local use. None of the four has public specifications, pricing, or a launch date, making the announcement a statement of direction rather than a shipping product.
Other announcements at the conference clarified that direction. Chief Executive Eddie Wu said the Qwen team plans to train models of five to ten trillion parameters — a target Alibaba's own statements assign to Qwen 4.5 and Qwen 5 — to tackle more complex, longer-horizon tasks. To support models at that scale, Alibaba's T-Head unit unveiled the Zhenwu V900 accelerator, which promises three times the performance of its M890 predecessor, carries 216 GB of memory, and scales to clusters of up to 500,000 chips, with mass production scheduled for the first quarter of 2027.
Wu set a target of surpassing 20 gigawatts of global cloud capacity by 2032 and described a three-layer cloud architecture designed for agentic workloads. He also noted that demand for AI is outpacing supply, with AI supernodes coming online at commercial scale this quarter. Alibaba's Hong Kong shares rose 5.1 percent on the day.
The domestic chip push unfolds against tightening US export curbs. Analysts point to August's open-sourced Qwen3.8 Flash Next, with its sparse attention and n-gram embedding techniques, as the closest available preview of the next-generation architecture, though Alibaba has not confirmed the connection.
A Broader Trajectory: From Open Models to an Integrated Stack
The recent releases cap an accelerating trajectory. Since its debut in 2023, Qwen has become the most widely used open-weights model family, with more than 300 million cumulative downloads and over 100,000 derivative models on Hugging Face, according to Alibaba.
The current flagship, Qwen3.8 Max, released in August with open weights, contains 2.4 trillion parameters and handles one million tokens of context, while September's Qwen3.8 Omni Flash extends processing across text, image, audio, and video. The licensing structure has evolved in parallel: the flagship carries a clause requiring companies generating more than $50 million in annual revenue to share that revenue with Alibaba, while smaller distilled models ship under permissive Apache 2.0 licenses.
Several open questions will shape how this strategy unfolds: whether independently verified benchmarks confirm the vendor-reported results for the agents and speech models, whether chip production timelines align with the model roadmap, and how the revenue-sharing license is adopted in practice. With Qwen 4 specifications still undisclosed and the next generation already in training, Alibaba has laid out an unusually broad agenda spanning silicon, cloud infrastructure, models, and consumer agents. The coming quarters will show how these pieces converge in practice.
Source: Metaverse Post