Alibaba’s Qwen workforce has made Qwen3.8-Max broadly obtainable and confirmed that its open weights ship subsequent week. A second checkpoint, Qwen3.8-27B, can be going open-weights. Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts mannequin. It accepts textual content, picture and video as enter and returns textual content.
Is it deployable
Sure, however the deployable floor relies on which artifact you might be making use of.
The hosted API is deployable right this moment by any firm measurement. It’s OpenAI- and DashScope-compatible, so integration is a base-URL and model-ID change. The open weights are a unique matter. At 2.4T whole parameters, the checkpoint is a multi-node datacenter artifact. Alibaba has not disclosed the activated-parameter rely. Serving value subsequently can not but be modeled. Qwen3.8-27B is the checkpoint that matches unusual on-premise GPU {hardware}.
The printed function set maps cleanly onto 4 industries. These are software program engineering, authorized and monetary doc assessment, media and e-commerce operations, and design.
Purposes embrace repository-scale coding brokers and long-document data bases. Lengthy-video indexing, structured information extraction and multi-step analysis assistants additionally match.
Interactive explainer
“+row[0]+’‘+(vals[qi]!==null&&vals[fi]!==null?(vals[qi]>vals[fi]?’above’:vals[qi]
‘;
C.m.forEach(operate(mn,j));
el.innerHTML=h;L.appendChild(el)});
q(‘#bmw’).textContent=”Qwen3.8-Max above Fable5 on “+win+’ of ‘+cmp+’ comparable rows’;
setTimeout(operate(){qa(‘#bml .f’).forEach(operate(f){f.type.width=f.dataset.w+’%’})},40);
setTimeout(rs,140);}
/* RL curve */
var RL=[[0,.474],[500,.469],[1000,.487],[1500,.586],[2000,.606],[2500,.622],[3000,.616],[3500,.647],[4000,.725],[4500,.719],[5000,.689]];
var rlDone=false;
operate paintRl(){
if(rlDone)return;rlDone=true;
var g=q(‘#rlg’),X=operate(v){return 44+v/5000*570},Y=operate(v){return 210-(v-.44)/(.76-.44)*180};
operate mk(t,a){var e=d.createElementNS(‘http://www.w3.org/2000/svg’,t);
for(var okay in a)e.setAttribute(okay,a[k]);g.appendChild(e);return e}
[.45,.5,.55,.6,.65,.7,.75].forEach(operate(v){
mk(‘line’,{x1:44,y1:Y(v),x2:614,y2:Y(v),stroke:’#1d1830′,’stroke-width’:1});
var t=mk(‘textual content’,{x:6,y:Y(v)+3,class:’lbl’});t.textContent=v.toFixed(2)});
mk(‘line’,{x1:44,y1:Y(.474),x2:614,y2:Y(.474),stroke:’#4a4170′,’stroke-width’:1,’stroke-dasharray’:’4 4′});
var sb=mk(‘textual content’,{x:470,y:Y(.474)+14,class:’lbl’});sb.textContent=”SFT baseline 0.474″;
var pts=RL.map(operate(p){return X(p[0])+’,’+Y(p[1])}).be part of(‘ ‘);
var pl=mk(‘polyline’,{factors:pts,fill:’none’,stroke:’#7C5CFF’,’stroke-width’:2.6,’stroke-linejoin’:’spherical’,’stroke-linecap’:’spherical’});
var len=pl.getTotalLength?pl.getTotalLength():2000;
pl.setAttribute(‘stroke-dasharray’,len);pl.setAttribute(‘stroke-dashoffset’,len);
var an=d.createElementNS(‘http://www.w3.org/2000/svg’,’animate’);
an.setAttribute(‘attributeName’,’stroke-dashoffset’);an.setAttribute(‘from’,len);an.setAttribute(‘to’,0);
an.setAttribute(‘dur’,’1.5s’);an.setAttribute(‘fill’,’freeze’);pl.appendChild(an);
RL.forEach(operate(p,i){
var pk=p[1]===.725;
var c=mk(‘circle’,{cx:X(p[0]),cy:Y(p[1]),r:pk?6:3.6,fill:pk?’#B49CFF’:’#7C5CFF’,opacity:0});
var a2=d.createElementNS(‘http://www.w3.org/2000/svg’,’animate’);
a2.setAttribute(‘attributeName’,’opacity’);a2.setAttribute(‘from’,0);a2.setAttribute(‘to’,1);
a2.setAttribute(‘dur’,’.3s’);a2.setAttribute(‘start’,(i*0.13)+’s’);a2.setAttribute(‘fill’,’freeze’);c.appendChild(a2);
var t=mk(‘textual content’,{x:X(p[0]),y:Y(p[1])-12,’text-anchor’:’center’,class:’lbl’,fill:pk?’#B49CFF’:’#8b81ab’,opacity:0});
t.textContent=p[1].toFixed(3);
var a3=d.createElementNS(‘http://www.w3.org/2000/svg’,’animate’);
a3.setAttribute(‘attributeName’,’opacity’);a3.setAttribute(‘from’,0);a3.setAttribute(‘to’,1);
a3.setAttribute(‘dur’,’.3s’);a3.setAttribute(‘start’,(i*0.13)+’s’);a3.setAttribute(‘fill’,’freeze’);t.appendChild(a3);
if(ipercent2===0){var xl=mk(‘textual content’,{x:X(p[0]),y:228,’text-anchor’:’center’,class:’lbl’});xl.textContent=p[0]}});
var pk=mk(‘textual content’,{x:X(4000),y:Y(.725)-26,’text-anchor’:’center’,class:’lbl-b’,fill:’#B49CFF’});
pk.textContent=”peak”;
setTimeout(rs,200)}
/* cross-harness */
var HZ=[
{n:’CoWorkBench’,b:[[‘Fable5 (OpenClaw)’,75.9,0],[‘Opus4.8 (OpenClaw)’,72.3,0],[‘Qwen3.7-Max (OpenClaw)’,64.6,0],[‘QwenWork’,73.2,1],[‘Claude Code’,74.6,1],[‘Codex’,75.8,1],[‘OpenClaw’,74.8,1],[‘Hermes’,73.3,1]]},
{n:’WorkspaceBench’,b:[[‘Fable5 (OpenClaw)’,68.7,0],[‘Opus4.8 (OpenClaw)’,66.8,0],[‘Qwen3.7-Max (OpenClaw)’,61.4,0],[‘QwenWork’,67.0,1],[‘Claude Code’,67.3,1],[‘Codex’,67.6,1],[‘OpenClaw’,67.7,1],[‘Hermes’,71.2,1]]},
{n:’JobBench’,b:[[‘Fable5 (OpenCode)’,57.4,0],[‘Opus4.8 (OpenCode)’,48.4,0],[‘Qwen3.7-Max (OpenCode)’,31.3,0],[‘QwenWork’,58.0,1],[‘Claude Code’,59.0,1],[‘Codex’,57.5,1],[‘OpenClaw’,59.8,1],[‘Hermes’,59.6,1],[‘OpenCode’,53.4,1]]}];
HZ.forEach(operate(h){
var c=d.createElement(‘div’);c.className=”hc”;
var s=”
“+h.n+’
‘;
var mx=Math.max.apply(null,h.b.map(operate(x){return x[1]}));
h.b.forEach(operate(b){s+=’
‘+
‘‘+b[0]+’‘+
‘‘+b[1].toFixed(1)+’
‘});
c.innerHTML=s;q(‘#hzl’).appendChild(c)});
/* deploy */
var DEP=[
‘Path: hosted API only. Call model ID qwen3.8-max through the DashScope or OpenAI-compatible endpoint. Self-hosting 2.4T weights is out of reach. Best fits: coding agents, document and video ingestion, research assistants. Owner: the founding engineer or a solo AI engineer.’,
‘Path: API first, 27B for anything sensitive. Use the hosted flagship for long-horizon agent work and keep Qwen3.8-27B on your own GPUs for PII-bearing or high-volume routing. Best fits: support automation, analytics copilots, internal tooling. Owner: an ML platform team.’,
‘Path: both, with an evaluation gate. Multi-node clusters can host the open weights once activated parameters and license are published. Until then run the API behind your own harness. The cross-harness numbers suggest portability across Claude Code, Codex and OpenClaw. Owner: an applied AI group reporting to the CTO.’,
‘Path: wait, then take 27B. No license file has been published for Qwen3.8-Max, so procurement cannot clear it yet. Air-gapped deployment realistically means Qwen3.8-27B. Best fits: healthcare, defense, public sector. Owner: security architecture plus a data science lead.’];
qa(‘#who button’).forEach(operate(b){b.onclick=operate(){
qa(‘#who button’).forEach(operate(x){x.classList.take away(‘on’)});
b.classList.add(‘on’);q(‘#dep’).innerHTML=DEP[+b.dataset.w];setTimeout(rs,60)}});
q(‘#dep’).innerHTML=DEP[0];
/* proof */
var EV=[[
‘Full benchmark table across coding, agent, reasoning, multimodal, document, video’,
‘2.4 trillion total parameters, mixture-of-experts’,
‘1M context; 991K max input, 131K max output, 262K max reasoning’,
‘Text, image and video input; text output’,
‘$2.00 input, $6.00 output, $0.25 implicit cache read per 1M tokens’,
‘Built-in code_interpreter, web_search, web_extractor, t2i_search, i2i_search’,
‘RL scaling curve, including the decline past the 4,000-environment peak’,
‘Cross-harness results on QwenWork, Claude Code, Codex, OpenClaw, Hermes’,
‘Open weights for Qwen3.8-Max and Qwen3.8-27B announced for next week’
],[
‘Model card’,
‘License file’,
‘Activated parameters per token’,
‘Independent evaluation from Artificial Analysis or LMArena’,
‘Harness, attempt count and reasoning settings behind each score’,
‘Note: the multimodal table compares against Qwen3.7-Plus, not Qwen3.7-Max’
]];
var evMode=0;
operate paintEv(){var L=q(‘#evl’);L.innerHTML=”;
EV[evMode].forEach(operate(t,i){var e=d.createElement(‘div’);e.className=”q-ev”;
e.type.animationDelay=(i*55)+’ms’;
e.innerHTML=’‘+(evMode?’?’:’✓’)+’‘+t+’‘;
L.appendChild(e)});setTimeout(rs,120)}
qa(‘#evp button’).forEach(operate(b){b.onclick=operate(){
qa(‘#evp button’).forEach(operate(x){x.classList.take away(‘on’)});
b.classList.add(‘on’);evMode=+b.dataset.e;paintEv()}});
paintEv();paintBm();
window.addEventListener(‘load’,rs);setTimeout(rs,300);setTimeout(rs,900);
})();
” type=”width:100%;peak:600px;border:0;overflow:hidden;show:block” scrolling=”no” loading=”lazy” title=”Qwen3.8-Max Interactive Explainer”>
What’s Technically Out there
The mannequin web page lists a 1M-token context window. Most enter is 991K tokens, dropping to 983K when pondering is enabled. Most output is 131K tokens in each modes, and the utmost reasoning funds is 262K tokens. Charge limits are 2M tokens per minute and 15K requests per minute.
Pricing is $2.00 per 1M enter tokens and $6.00 per 1M output tokens. Implicit cache reads value $0.25 per 1M tokens. Specific cache creation is $2.50 and express cache reads are $0.17 per 1M tokens. Cached enter is eight occasions cheaper than contemporary enter. Prefix stability subsequently drives value greater than immediate size does.
Supported capabilities embrace operate calling, structured outputs, batches, prefix completion and fine-tuning. 5 built-in instruments ship on the Responses API: code_interpreter, web_search, web_extractor, t2i_search and i2i_search.


Efficiency
Alibaba printed a full benchmark desk with this launch. Qwen3.8-Max scores 86.6 on Terminal-Bench 2.1, forward of Claude Opus 4.8 and Claude Fable 5 at 84.6, behind GPT-5.6 Sol (max) at 88.8. It reviews 67.7 on SWE-bench Professional towards Fable 5’s 80.0, and 73.5 on FrontierSWE towards Fable 5’s 88.8. It leads PaperBench at 93.0 and IFBench at 82.8. GPQA Diamond lands at 92.6, up marginally from Qwen3.7-Max’s 92.4. The clearest features are multimodal and agentic, not reasoning. It tops most imaginative and prescient rows, together with OSWorld-Verified 86.1, Parametric CAD Bench 91.5, and OmniDocBench 1.5 at 92.1. In opposition to its personal predecessor the soar is giant: DeepSWE 1.1 strikes from 21.6 to 56.6, FrontierSWE from 40.7 to 73.5, JobBench from 31.3 to 53.4. Two caveats belong in any sincere learn. The multimodal desk benchmarks towards Qwen3.7-Plus, not Qwen3.7-Max, which flatters the generational delta. And Alibaba’s personal RL scaling curve peaks at 0.725 close to 4,000 coaching environments, then declines to 0.719 and 0.689.
Key Takeaways
- Qwen3.8-Max is a 2.4T-parameter MoE mannequin with 1M context, now usually obtainable.
- Pricing is $2 enter, $6 output and $0.25 cached enter per 1M tokens.
- Open weights for Qwen3.8-Max and Qwen3.8-27B are promised subsequent week.
- No benchmark desk, license, or activated-parameter rely has been printed.
- The 27B checkpoint, not the flagship, is the sensible on-premise deployment path.
Try the Technical particulars, API and Qwen Studio. Additionally, be at liberty to observe us on Twitter and don’t neglect to affix our 150k+ML SubReddit and Subscribe to our E-newsletter. Wait! are you on telegram? now you’ll be able to be part of us on telegram as properly.
Have to accomplice with us for selling your GitHub Repo OR Hugging Face Web page OR Product Launch OR Webinar and so on.? Join with us
Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is dedicated to harnessing the potential of Synthetic Intelligence for social good. His most up-to-date endeavor is the launch of an Synthetic Intelligence Media Platform, Marktechpost, which stands out for its in-depth protection of machine studying and deep studying information that’s each technically sound and simply comprehensible by a large viewers. The platform boasts of over 2 million month-to-month views, illustrating its recognition amongst audiences.









