← All stories
● Covered by 7 sources · 24 reportsMedium impact1 negative22 neutral1 positive

Alibaba Releases Qwen3.8-Max and Qwen3.8-27B AI Models, Including Open Weights

🔄 Updated 8d ago — new reporting from Tom's Hardware
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Alibaba released Qwen3.8-Max, a 2.4 trillion parameter multimodal AI model.
  • Qwen3.8-Max supports a 1 million token context window.
  • A 27-billion-parameter version, Qwen3.8-27B, was also released with open weights.
  • Both models are designed for complex tasks, coding, and research.
  • Alibaba claims Qwen3.8-Max performance rivals models from Anthropic and OpenAI.
  • Qwen 3.8 27B is an Apache 2 licensed vision-capable LLM.
  • Qwen 3.8 27B defaults to "xhigh" reasoning effort.
  • The "xhigh" setting leads to overthinking and extended generation times.
  • Qwen 3.8 27B was released on August 16, 2026.
  • Qwen3.8-27B achieved 50.44 tokens/second on an NVIDIA RTX PRO 4000 SFF.
  • The RTX PRO 4000 SFF has 24 GB VRAM.
  • A custom llama.cpp build increased throughput by 21.97% to 55.40 tok/s.
  • Embedded MTP increased throughput 2.81 times to 59.46 tok/s.
  • Throughput fell from 50.44 to 37.02 tokens/second with 69.2 MiB more precision.
  • Qwen3.8 27B was released on August 14, 2026.
  • Qwen3.8 27B scored 52 on the Artificial Analysis Intelligence Index.
  • The median score for comparable open-weight models is 9.
  • Qwen3.8 27B supports a 256k token context window.
  • Qwen3.8 27B generated 160M tokens during Intelligence Index evaluation.
  • The median token generation for Intelligence Index evaluation is 43M.
  • Qwen3.8 27B costs $0.00 per 1M input tokens.
  • Qwen3.8 27B costs $0.00 per 1M output tokens.
  • Qwen3.8-27B supports local execution on high-end consumer hardware.
  • Qwen3.8-27B includes native image and video understanding.
  • Qwen3.8-27B requires 56GB of GPU memory at full 16-bit precision.
  • An FP8 version of Qwen3.8-27B needs 28GB of GPU memory.
  • 4-bit quantization reduces Qwen3.8-27B to 17GB.
  • Unsloth released Dynamic v3.0 GGUFs.
  • Unsloth Dynamic v3.0 improves model quality preservation and accuracy.
  • Dynamic v3.0 delivers over 10% better top-1% accuracy for Qwen3.8-27B.
  • Dynamic v3.0 GGUFs work with llama.cpp and Unsloth Desktop.
  • Dynamic v3.0 preserves more model quality at the same size.
  • Dynamic v3.0 shows stronger results across Divergence-300 @32 and KL Divergence.
  • Unsloth Qwen3.8 had over 5.1 million downloads in 5 days.
  • Unsloth uses a higher-quality imatrix calibration dataset from diverse sources.
  • The imatrix dataset is refined for agentic coding, chat, and multilingual performance.
  • Unsloth improved layer selection and introduced more quantization techniques.
  • Unsloth does not train on the imatrix calibration dataset.
  • Unsloth does not use QAT or QAD, only post-training quantization.
  • Qwen 3.8 27B successfully reverse-engineered a commercial application's license check.
  • Qwen 3.8 27B completed the reverse-engineering task in 30 minutes.
  • Qwen 3.8 27B runs on a Lenovo ThinkStation PGX with an Nvidia GB10 Grace Blackwell chip.
  • The Lenovo ThinkStation PGX has 128 GB of unified memory and 273 GB/s bandwidth.
  • Qwen 3.8 27B achieves 15 to 30 tokens/second out of the box on the ThinkStation PGX.
  • With SGLang, NVFP4, and DFlash2, Qwen 3.8 27B reaches 50 tokens/second on code and reasoning.
  • Qwen 3.8 27B is the top open-weights model in the 4B to 40B size class out of 135 models.
  • Qwen 3.8 27B beats more expensive models on SWE-bench Pro.
  • Alibaba released Qwen3.8-Flash, a 125 billion parameter multimodal MoE model.
  • Qwen3.8-Flash is an early preview of the architecture for Qwen4.
  • Qwen3.8-Flash uses a hybrid Gated DeltaNet + Gated Attention design.
  • Qwen3.8-Flash has superior capabilities in coding and office tasks.
  • Qwen 3.8 27B is 17GB for four-bit quantized weights.
  • Qwen 3.8 27B grabbed attention from users with RTX 5090, RTX 4090, RTX 3090, Radeon RX 7900 XTX, Radeon AI Pro R9700, or Arc Pro B70.
  • Qwen3.8 27B 4-bit Q4_K_M version performs comparably to the full BF16 model.
  • 1-bit quantization significantly degrades Qwen3.8 27B performance.
  • Qwen3.8 27B 4-bit Q4_K_M fits on a 24 GB card, leaving room for 64k tokens of context.
  • Qwen3.8 27B 4-bit Q4_K_M matches the full model on Terminal-Bench 2.1.
  • Qwen3.8 27B 1-bit quantization performs around random chance on GPQA Diamond.
  • Qwen3.8-2.4T-A95B is the first Qwen-Max-class model with open weights.
  • Qwen3.8-2.4T-A95B has 95 billion activated parameters per token.
  • Qwen3.8-2.4T-A95B uses a hybrid linear-plus-full-attention architecture.
  • Qwen3.8-2.4T-A95B can be deployed on Amazon SageMaker HyperPod using vLLM.
  • Qwen3.8-2.4T-A95B can be deployed on a ml.p6-b300 instance (8x NVIDIA B300 Blackwell Ultra GPUs).
  • Tom's Hardware Premium published benchmarks for Qwen 3.8 27B.
  • Jeff Kampman benchmarked Qwen 3.8 27B on consumer devices.
  • Qwen 3.8 27B was benchmarked on Mac Mini, DGX Spark, and Strix Halo systems.
  • The PC market shifts to high-end AI PCs and ultra-light laptops.
  • Mid-range PCs are underserved due to AI data center demand.
  • Ajinomoto's ABF substrate supply chain is strained.
  • ABF substrate price increased by 30%.
  • ShapeLearn released optimized GGUF models for Qwen 3.8 27B.
  • ShapeLearn's GGUF models offer improved quality and speed.
  • ShapeLearn-Lite GGUFs were released on August 18, 2026.
  • The MTP draft head is bundled in every ShapeLearn GGUF.
  • DFlash2 uses a separate 1.1 GB draft model.
  • DFlash2 requires llama.cpp b10658 or newer.
  • Alibaba Cloud released Qwen Image 2.1, a 7 billion parameter image generation AI model.
  • Qwen Image 2.1 supports native transparency.
  • Qwen Image 2.1 supports improved image editing with up to 10 reference images.
  • Qwen Image 2.1 forbids commercial resale without a separate license.

Alibaba Unveils New Qwen3.8 AI Models

Alibaba has announced the release of its Qwen3.8-Max and Qwen3.8-27B AI models. Qwen3.8-Max is a multimodal model featuring 2.4 trillion parameters and a context window of up to 1 million tokens. The company states this makes it its most capable AI model to date. Qwen3.8-27B is a smaller, 27-billion-parameter version designed for local deployment.

Both models are built on the architectural foundation of Qwen3.5 and are intended for tasks such as coding, professional work, research, and long-horizon agentic tasks. Alibaba's own testing suggests Qwen3.8-Max performs comparably to, and in some cases surpasses, models like Anthropic's Fable 5 and GPT-5.6 Sol Max on specific benchmarks.

Open Weights and Accessibility

Alibaba has committed to releasing the open weights for Qwen3.8-Max and Qwen3.8-27B. The weights for Qwen3.8-Max are expected to be published on Hugging Face and ModelScope, marking the first time a Qwen-Max-class model will have downloadable weights. The Qwen3.8-27B model's open weights are available under the Apache 2.0 license, allowing it to run locally on consumer hardware.

The Qwen3.8-Max model is also available through QwenCloud and Alibaba Cloud Model Studio, priced at $2 per million input tokens and $6 per million output tokens. The release of open weights for these models brings Qwen-Max-class capabilities to the open-source community.

Architectural Details and Capabilities

Qwen3.8-Max utilizes a sparse mixture-of-experts (MoE) design with hybrid attention. While it contains 2.4 trillion parameters in total, it activates approximately 95 billion for each token, which allows the model to draw on specific parts of the model needed for each task. The model's capabilities include visual intelligence, coding, and handling complex, multi-step tasks.

The Qwen3.8-27B model is a dense model with native vision-language understanding, capable of processing images and videos. Both models are designed for improved reliability in completing complex tasks.

Performance Claims and Benchmarking

Alibaba's internal testing indicates Qwen3.8-Max's performance broadly matches or exceeds that of models from Anthropic and OpenAI. On the OSWorld-Verified benchmark, Qwen3.8-Max reportedly scored 86.1, surpassing GPT-5.6 Sol Max (83.2) and Fable 5 (85.0). It also posted the highest reported score on PaperBench.

However, independent benchmarks have shown varying results, with performance influenced by factors like time and token budgets. The Qwen3.8-27B model, despite its smaller size, is claimed to perform comparably to Anthropic's Opus 4.6 on coding and knowledge work tasks, and includes vision capabilities.

Updates

🕒 2026-09-23 · new reporting from Tom's Hardware
  • Alibaba Cloud released Qwen Image 2.1, a 7 billion parameter image generation AI model.
  • Qwen Image 2.1 supports native transparency.
  • Qwen Image 2.1 supports improved image editing with up to 10 reference images.
  • Qwen Image 2.1 forbids commercial resale without a separate license.
🕒 2026-09-18 · new reporting from Hacker News Front Page
  • ShapeLearn released optimized GGUF models for Qwen 3.8 27B.
  • ShapeLearn's GGUF models offer improved quality and speed.
  • ShapeLearn-Lite GGUFs were released on August 18, 2026.
  • The MTP draft head is bundled in every ShapeLearn GGUF.
  • DFlash2 uses a separate 1.1 GB draft model.
  • DFlash2 requires llama.cpp b10658 or newer.
🕒 2026-09-12 · new reporting from Tom's Hardware
  • Tom's Hardware Premium published benchmarks for Qwen 3.8 27B.
  • Jeff Kampman benchmarked Qwen 3.8 27B on consumer devices.
  • Qwen 3.8 27B was benchmarked on Mac Mini, DGX Spark, and Strix Halo systems.
  • The PC market shifts to high-end AI PCs and ultra-light laptops.
  • Mid-range PCs are underserved due to AI data center demand.
  • Ajinomoto's ABF substrate supply chain is strained.
  • ABF substrate price increased by 30%.
🕒 2026-09-10 · new reporting from AWS Machine Learning Blog
  • Qwen3.8-2.4T-A95B is the first Qwen-Max-class model with open weights.
  • Qwen3.8-2.4T-A95B has 95 billion activated parameters per token.
  • Qwen3.8-2.4T-A95B uses a hybrid linear-plus-full-attention architecture.
  • Qwen3.8-2.4T-A95B can be deployed on Amazon SageMaker HyperPod using vLLM.
  • Qwen3.8-2.4T-A95B can be deployed on a ml.p6-b300 instance (8x NVIDIA B300 Blackwell Ultra GPUs).
🕒 2026-09-08 · new reporting from Hacker News Front Page
  • Qwen3.8 27B 4-bit Q4_K_M version performs comparably to the full BF16 model.
  • 1-bit quantization significantly degrades Qwen3.8 27B performance.
  • Qwen3.8 27B 4-bit Q4_K_M fits on a 24 GB card, leaving room for 64k tokens of context.
  • Qwen3.8 27B 4-bit Q4_K_M matches the full model on Terminal-Bench 2.1.
  • Qwen3.8 27B 1-bit quantization performs around random chance on GPQA Diamond.
🕒 2026-09-08 · new reporting from Tom's Hardware
  • Qwen 3.8 27B is 17GB for four-bit quantized weights.
  • Qwen 3.8 27B grabbed attention from users with RTX 5090, RTX 4090, RTX 3090, Radeon RX 7900 XTX, Radeon AI Pro R9700, or Arc Pro B70.
🕒 2026-08-28 · new reporting from The New Stack
  • Alibaba released Qwen3.8-Flash, a 125 billion parameter multimodal MoE model.
  • Qwen3.8-Flash is an early preview of the architecture for Qwen4.
  • Qwen3.8-Flash uses a hybrid Gated DeltaNet + Gated Attention design.
  • Qwen3.8-Flash has superior capabilities in coding and office tasks.
🕒 2026-08-23 · new reporting from Hacker News Front Page
  • Qwen 3.8 27B successfully reverse-engineered a commercial application's license check.
  • Qwen 3.8 27B completed the reverse-engineering task in 30 minutes.
  • Qwen 3.8 27B runs on a Lenovo ThinkStation PGX with an Nvidia GB10 Grace Blackwell chip.
  • The Lenovo ThinkStation PGX has 128 GB of unified memory and 273 GB/s bandwidth.
  • Qwen 3.8 27B achieves 15 to 30 tokens/second out of the box on the ThinkStation PGX.
  • With SGLang, NVFP4, and DFlash2, Qwen 3.8 27B reaches 50 tokens/second on code and reasoning.
  • Qwen 3.8 27B is the top open-weights model in the 4B to 40B size class out of 135 models.
  • Qwen 3.8 27B beats more expensive models on SWE-bench Pro.
🕒 2026-08-19 · new reporting from Hacker News Front Page
  • Unsloth released Dynamic v3.0 GGUFs.
  • Unsloth Dynamic v3.0 improves model quality preservation and accuracy.
  • Dynamic v3.0 delivers over 10% better top-1% accuracy for Qwen3.8-27B.
  • Dynamic v3.0 GGUFs work with llama.cpp and Unsloth Desktop.
  • Dynamic v3.0 preserves more model quality at the same size.
  • Dynamic v3.0 shows stronger results across Divergence-300 @32 and KL Divergence.
  • Unsloth Qwen3.8 had over 5.1 million downloads in 5 days.
  • Unsloth uses a higher-quality imatrix calibration dataset from diverse sources.
  • The imatrix dataset is refined for agentic coding, chat, and multilingual performance.
  • Unsloth improved layer selection and introduced more quantization techniques.
  • Unsloth does not train on the imatrix calibration dataset.
  • Unsloth does not use QAT or QAD, only post-training quantization.
🕒 2026-08-18 · new reporting from VentureBeat
  • Qwen3.8-27B supports local execution on high-end consumer hardware.
  • Qwen3.8-27B includes native image and video understanding.
  • Qwen3.8-27B requires 56GB of GPU memory at full 16-bit precision.
  • An FP8 version of Qwen3.8-27B needs 28GB of GPU memory.
  • 4-bit quantization reduces Qwen3.8-27B to 17GB.
🕒 2026-08-17 · new reporting from Hacker News Front Page
  • Qwen3.8 27B was released on August 14, 2026.
  • Qwen3.8 27B scored 52 on the Artificial Analysis Intelligence Index.
  • The median score for comparable open-weight models is 9.
  • Qwen3.8 27B supports a 256k token context window.
  • Qwen3.8 27B generated 160M tokens during Intelligence Index evaluation.
  • The median token generation for Intelligence Index evaluation is 43M.
  • Qwen3.8 27B costs $0.00 per 1M input tokens.
  • Qwen3.8 27B costs $0.00 per 1M output tokens.
🕒 2026-08-17 · new reporting from Hacker News Front Page
  • Qwen3.8-27B achieved 50.44 tokens/second on an NVIDIA RTX PRO 4000 SFF.
  • The RTX PRO 4000 SFF has 24 GB VRAM.
  • A custom llama.cpp build increased throughput by 21.97% to 55.40 tok/s.
  • Embedded MTP increased throughput 2.81 times to 59.46 tok/s.
  • Throughput fell from 50.44 to 37.02 tokens/second with 69.2 MiB more precision.
🕒 2026-08-17 · new reporting from Hacker News Front Page
  • Qwen 3.8 27B is an Apache 2 licensed vision-capable LLM.
  • Qwen 3.8 27B defaults to "xhigh" reasoning effort.
  • The "xhigh" setting leads to overthinking and extended generation times.
  • Qwen 3.8 27B was released on August 16, 2026.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~34 min · 27 stories · Oct 02

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

How outlets covered it

Alibaba Cloud released Qwen Image 2.1, a new 7 billion parameter image generation AI model. The model supports native transparency and improved image editing with up to 10 reference images, and its developers claim it outperforms larger models like Google's Nano Banana 2.0 on internal benchmarks. The licensing agreement for Qwen Image 2.1 now explicitly forbids commercial resale without a separate license.

ShapeLearn has released its full set of optimized GGUF models for the Qwen 3.8 27B language model, offering improved quality and speed compared to their initial 'Lite' release and other quantizations. These new models enhance performance for users running Qwen 3.8 27B on various GPUs, with specific recommendations for different memory and throughput needs.

Tom's Hardware Premium published benchmarks for the Qwen 3.8 27B AI model across various consumer devices, highlighting the importance of configuration over raw performance metrics. The publication also analyzed the current state of the PC market, noting a shift towards high-end AI PCs and ultra-light laptops, leaving the mid-range underserved due to AI data center demand. Additionally, it examined the strained supply chain of Ajinomoto's ABF substrates, crucial for AI accelerators, which has led to a 30% price increase.

Alibaba's Qwen team released Qwen3.8-2.4T-A95B, their first Qwen-Max-class model with open weights, featuring 2.4 trillion parameters and a 262K token context. This release allows full control over data and inference behavior, and a guide demonstrates its deployment on Amazon SageMaker HyperPod using vLLM.

Benchmarking of Qwen3.8 27B quantizations reveals that the 4-bit Q4_K_M version (17GB) performs comparably to the full BF16 model on agentic coding and instruction-following benchmarks. However, 1-bit quantization significantly degrades performance, making the model operate at near-random chance levels on reasoning tasks. This indicates that 4-bit quantization offers a viable option for running large language models on consumer hardware without substantial quality loss, while 1-bit quantization is not practical.

Benchmarking of Alibaba's Qwen 3.8 27B open-weight AI model on various hardware, including the RTX 5090, indicates that VRAM capacity alone is insufficient for optimal performance due to significant software and inference engine bottlenecks. This analysis highlights that effective local AI inference requires more than just powerful GPUs, emphasizing the importance of a well-optimized software stack for practical application.

Alibaba has released Qwen3.8-Flash, an open-weight, multimodal Mixture-of-Experts (MoE) model with 125 billion parameters, serving as an early preview of the architectural changes planned for Qwen4. This release allows developers to examine and test new design forms, such as the hybrid Gated DeltaNet + Gated Attention, before their full integration into future Qwen models.

The Qwen 3.8 27B open-weights large language model demonstrated its advanced capabilities by successfully reverse-engineering a commercial application's license check on a local machine. This event highlights the increasing sophistication of local LLMs in complex, specialized tasks, even correcting its own errors during the process.

Unsloth released Dynamic v3.0 GGUFs, an update to its quantization method, which improves model quality preservation and accuracy for large language models. This iteration delivers over 10% better top-1% accuracy for Qwen3.8-27B at the same size compared to other providers, making quantized models more efficient without significant performance loss.

Alibaba released Qwen3.8-27B, a 27-billion-parameter multimodal AI model under an Apache 2.0 license, allowing local execution on high-end consumer hardware. This model offers capabilities like image/video understanding, a large context window, and support for coding and agentic workflows, making advanced AI accessible without cloud APIs.

Alibaba released its Qwen3.8 27B model on August 14, 2026, which scored 52 on the Artificial Analysis Intelligence Index. This places it significantly above the median score of 9 for comparable open-weight models of similar size, indicating its performance in intelligence benchmarks.

An experiment demonstrated that a carefully configured local inference setup for the Qwen3.8-27B model, utilizing a custom llama.cpp build and embedded MTP, achieved 50.44 tokens per second on an NVIDIA RTX PRO 4000 SFF with 24 GB VRAM. This result highlights that optimizing the interplay between components can yield better performance than simply using individually high-performing parts.

Alibaba's Qwen research lab released Qwen 3.8 27B, an Apache 2 licensed 27B parameter vision-capable LLM. The model's default "xhigh" reasoning effort setting leads to overthinking and extended generation times, despite producing high-quality outputs.

Alibaba has released an open-weights 27 billion parameter version of its Qwen3.8 model under the Apache 2.0 license, which can run locally on consumer hardware. Benchmarks indicate this model performs comparably to Anthropic's Opus 4.6, particularly in coding and knowledge work tasks, and includes vision capabilities.

Unsloth has released GGUF files for Qwen3.8-27B, a new version of the Qwen open-model family, featuring improved coding, agentic tasks, and native vision-language understanding. This update allows for better integration into development tools and offers more reliable task completion for complex, multi-step operations.

Qwen has released Qwen3.8-27B, an open-weight, 27-billion-parameter model with native vision-language understanding and improved agentic capabilities. This release provides a compact, deployment-friendly model for local inference, offering advancements in coding, research, and multi-step task completion.

Qwen has released Qwen3.8, a new large language model in its open-model family, featuring improved capabilities in coding, research, and agentic tasks. This release includes FP8-quantized model weights compatible with various inference frameworks, offering performance nearly identical to the original model.

Qwen has released Qwen3.8, the latest generation in its open-model family, featuring 2.4 trillion parameters and 95 billion activated parameters. This release brings Qwen-Max-class capabilities to the open-source community, improving performance in coding, professional tasks, research, and long-horizon agentic tasks.

Recent benchmarks for Alibaba's Qwen 3.8-Max model show significant performance differences depending on the time and token budgets allocated, revealing that raw benchmark scores and per-token pricing do not fully predict the actual cost of using AI models. This highlights the need for developers to consider 'cost per successful task' and explicitly define time and token budgets when evaluating and selecting models.

Alibaba announced Qwen3.8-Max, a new multimodal model with 2.4 trillion parameters, but faces criticism for promising open-source weights without immediately releasing them. This approach is seen by some as an API business model disguised as open source, similar to a previous move by Moonshot AI.

Alibaba's Qwen team unveiled Qwen3.8-Max, a 2.4-trillion-parameter multimodal LLM, which reportedly outperforms GPT-5.6 Sol Max and Fable 5 on agentic computing benchmarks like OSWorld-Verified. The company also announced plans to release open weights for Qwen3.8-Max next week, potentially reshaping enterprise AI adoption if released under a permissive license.

Alibaba has released Qwen3.8-Max, a multimodal AI model featuring 2.4 trillion parameters, available through QwenCloud and Alibaba Cloud Model Studio. This model is designed for complex, multi-day tasks and uses a sparse mixture-of-experts architecture, making it suitable for large organizations and inference providers due to its infrastructure requirements.

Alibaba introduced its Qwen3.8-Max AI model, featuring 2.4 trillion parameters and a 1 million token context window, positioning it as one of its most powerful models to date. This release contributes to the ongoing competition among Chinese tech companies to advance AI capabilities.

Alibaba has released its Qwen3.8-Max AI model, which the company states is its largest and most capable to date, with performance comparable to models from Anthropic and OpenAI. This release includes plans to make the model's weights available, marking a return to open-weight releases for Alibaba.