Alibaba’s Qwen workforce has launched Qwen-Audio-3.1, a 5-model audio stack spanning ASR, TTS and realtime interplay. The primary mannequin is Qwen-Audio-3.1-Realtime, a full-duplex speech mannequin constructed for voice brokers that decision instruments. Qwen additionally lower costs: about 85% on Realtime, about 70% on TTS and as much as 95% on ASR.
Is it deployable? Sure, as a managed API. qwen-audio-3.1-realtime-plus is dwell on QwenCloud over WebSocket. No open weights had been introduced.
What Ships on QwenCloud
The mannequin web page lists textual content and audio as each enter and output. Context is 262K tokens, with 245K max enter and 16K max output. Default limits are 60 requests and 100K tokens per minute. Pricing is $6.4 per 1M audio enter tokens and $0.8 per 1M textual content enter tokens. Textual content and audio output prices $24 per 1M tokens, with output textual content not charged. Key options embrace perform calling, internet search, structured outputs, context cache and fine-tuning.
A companion mannequin, Qwen-Audio-3.1-ASR-Flash-Filetrans, targets offline long-audio transcription. It helps scorching phrases, speaker separation, punctuation and multilingual plus Chinese language dialect recognition. It prices $0.15 enter and $0.47 output per 1M tokens.
Structure: 2 Fashions Behind 1 Voice
The system runs 2 fashions with the identical Audio Encoder and LLM design. A full-duplex determination mannequin predicts whether or not to maintain listening, communicate, cease or resume. A speech-to-text mannequin writes the response content material as textual content. A context-aware voice renderer then turns that textual content into streaming speech. It situations on dialog historical past, voice cues and acoustic context.
Coaching is organized into 3 layers: Suppose, Act, and Converse and Coordinate.
Suppose: M²-OPD
Core-Cocktail SFT re-anchors the audio mannequin to its supply textual content LLM utilizing million-hour-scale paired information. Multimodality OPD follows. A Textual content Instructor and a frozen Audio Reference rating every token of the coed’s personal trajectory. That is on-policy distillation, not imitation of pre-written solutions. Area consultants for empathy, pragmatic intent and acoustic scenes are then educated with GRPO. Multi-Instructor OPD merges them into 1 deployable mannequin.
Act: Executable Environments
Every coaching area bundles a instrument pool, a stateful JSON database and a natural-language enterprise coverage. Domains are seeded from open-source instrument and MCP server definitions. Each activity defines 1 of three outcomes: a write, a justified refusal, or an unsupported request. Scoring checks terminal state, then permitted writes, then behavioral assertions. A fluent reply can’t rescue a failed state verify.
GRPO receives rewards at dialogue, milestone and switch stage. Search coaching penalizes redundant queries with . Imply queries per search name fell from 4.37 to 1.05. Set off F1 slipped from 60.87% to 58.61%.
Converse and Coordinate
This layer decides whether or not, when and easy methods to communicate. On Full-Duplex-Bench v1.5, replies to folks speaking to another person fell from 0.13 to 0.03. On v3.0, the filler fee dropped from 0.7590 to 0.2960. There are trade-offs. After interruptions, the undesirable resume fee rose from 0.035 to 0.130. Interruption cease latency is 1.116 seconds, versus 0.383 for GPT-Realtime-2.
Interactive Explainer
Discover the Suppose, Act, Converse loop, duplex selections, a scored coaching episode and the search reward.
/* SLIDE 2 */
perform waves(el,n){el.innerHTML=”;for(var i=0;i
others:{l1:’Assistant’,l2:’Consumer to a colleague’,q1:true,q2:false,d:’Proper transfer: keep silent whereas the person addresses another person.’,m:[[‘Respond (lower is better)’,0.13,0.03,true],[‘Resume (higher is better)’,0.82,0.96,false]]},
int:{l1:’Assistant’,l2:’Consumer cuts in’,q1:false,q2:false,d:’Proper transfer: cease and reply the brand new request. Right here 3.1 trails 3.0 barely on each charges.’,m:[[‘Respond (higher is better)’,0.88,0.845,false],[‘Resume (lower is better)’,0.035,0.13,true]]},
bc:{l1:’Assistant’,l2:’Consumer: “mm-hm”‘,q1:false,q2:true,d:’Proper transfer: maintain speaking. A backchannel just isn’t a flip request.’,m:[[‘Respond (lower is better)’,0,0.0306,true],[‘Resume (higher is better)’,0.98,0.9694,false]]}};
var curScen=’bg’;
perform scen(okay){curScen=okay;var s=S[k];var bs=$(‘qaScen’).youngsters;for(var i=0;i
‘+bar(‘3.0′,x[1],’f30’,x[1].toFixed(3).substitute(/0+$/,”).substitute(/.$/,”))+bar(‘3.1′,x[2],’f31’,x[2].toFixed(4).substitute(/0+$/,”).substitute(/.$/,”))}
$(‘qaBars’).innerHTML=h;anim($(‘qaBars’));put up()}
perform bar(lab,v,cls,txt){return ‘
‘}
perform anim(root){var f=root.querySelectorAll(‘.fill’);setTimeout(perform(){for(var i=0;iAug 20 fast, almost in one breath.’},
{t:’Model calls search_flights(SFO, NRT, 2026-08-20) to confirm Flex seats remain.’},
{t:’Model reads back flight, date, Flex fare, $990.0 and expiry, then gets explicit confirmation.’,c:1},
{t:’Model calls place_hold(..., fare_class="flex"). Tool writes to the private DB copy.’,w:1},
{t:’Model tells Emily the reference HD482902 and expiry 2026-08-15.’},
{t:’Scorer runs 3 checks in order, using the tool trace, not the dialogue.’,s:1}];
var ei=-1;
var DB0='”HD482901″: {n “flight_no”: “SK2859”,n “standing”: “energetic”,n “expiry_date”: “2026-08-08″ // expiredn}’;
var DB1=DB0+’,n”HD482902”: {n “flight_no”: “SK2859”,n “date”: “2026-08-20”,n “fare_class”: “flex”,n “quantity”: 990.0,n “standing”: “energetic”,n “created_date”: “2026-08-12”,n “expiry_date”: “2026-08-15″n}’;
perform epRender(){var ul=$(‘qaSteps’),skip=$(‘qaSkip’).checked,h=””;for(var i=0;i
$(‘qaDb’).textContent=ei>=3?DB1:DB0;
var s=ei>=5;perform set(id,okay){var e=$(id);e.className=s?(okay?’okay’:’dangerous’):’pend’;e.textContent=s?(okay?’PASS’:’FAIL’):’pending’}
set(‘qaC1’,true);set(‘qaC2’,true);set(‘qaC3’,!skip);var v=$(‘qaVerdict’);
if(s){v.innerHTML=skip?’Episode FAILS. Appropriate state can’t excuse a damaged interplay rule.’:’Episode PASSES. Legitimate success feeds GRPO.’}else v.textContent=”Verdict seems after scoring”;put up()}
$(‘qaEpNext’).onclick=perform(){if(ei
$(‘qaQ’).oninput=rw;$(‘qaR’).oninput=rw;$(‘qaP’).oninput=rw;rw();
/* SLIDE 5 */
var M=[
{k:’τ-Voice overall (%)’,g:43.8,a:78.4,b:82.0,d:1,mx:100,low:false,n:’Half-duplex speech-to-text adaptation by the Qwen team; not comparable to official full-duplex τ-Voice results.’},
{k:’Audio MultiChallenge (%)’,g:50.33,a:47.12,b:52.21,d:2,mx:100,low:false,n:’GPT-4o-mini judge; official AMC uses o4-mini, so scores are not directly comparable to the leaderboard.’},
{k:’Multilingual BBA avg (%)’,g:82.1,a:81.7,b:88.1,d:1,mx:100,low:false,n:’In-house 14-language extension of Big Bench Audio. Largest gains on Arabic, Thai and Vietnamese.’},
{k:’Multi-turn attack, zh (%)’,g:42.0,a:80.5,b:26.0,d:1,mx:100,low:true,n:’In-house, 200 sessions. Attack success rate: lower is better.’},
{k:’Multi-party sessions passed’,g:0.07,a:0.08,b:0.96,d:2,mx:1,low:false,n:’In-house Qwen-Audio-FDB, 100 sessions; a session passes only if every point passes.’},
{k:’Human red-team pass (%)’,g:96.0,a:null,b:92.0,d:1,mx:100,low:false,n:’5 testers, 50 sessions, automatic judge. GPT-Realtime-2 leads here.’}];
var curMet=0,mb=$(‘qaMet’);
for(var m=0;m
Qwen-Audio-3.0-Realtimenot reported








