Back to news

Qwen3.8-2.4T-A95B Open source claims to beat Fable 5 and Opus 4.8!

Qwen3.8-2.4T-A95B Open source claims to be better than Fable 5 and Opus 4.8 in Swe and coding.

Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Beyond answering harder questions, Qwen3.8 is designed to carry complex, multi-step tasks through to completion with greater reliability.

Benchmark * (check at the bottom for more info):

qwen 3.8.png

Qwen3.8 Highlights Qwen3.8 features the following enhancements:

  • Core Capabilities: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks.
  • Agent Execution: Stronger autonomous planning and better handling of environment feedback, leading to more reliable end-to-end task completion.
  • Downstream Compatibility: Broader support for popular harnesses and development tools, making it easier to integrate into your existing stack.
  • Flexible Thinking Control: Reasoning depth can be tuned with reasoning_effort, and reasoning context from historical messages is retained via preserve_thinking.

Model Overview Type: Causal Language Model

  • Training Stage: Pre-training & Post-training Language Model
  • Number of Parameters: 2.4T in total and 95B activated Hidden Dimension: 8192 Token Embedding: 248,320 (Padded)
  • Number of Layers: 92 Hidden Layout: 23 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE))

Gated DeltaNet:

  • Number of Linear Attention Heads: 128 for V and 16 for QK
  • Head Dimension: 128

Gated Attention:

  • Number of Attention Heads: 64 for Q and 4 for KV Head Dimension: 256 Rotary Position Embedding Dimension: 64

Mixture of Experts:

  • Number of Experts: 512
  • Number of Activated Experts: 10 Routed + 1 Shared
  • Expert Intermediate Dimension: 2048
  • LM Output: 248,320 (Padded)
  • MTP (Multi-Token Prediction): trained with multiple steps
  • Context Length: 262,144 natively and extensible up to 1,010,000 tokens.

At this level of 2.4 trillion parameters every model is quite heavy and not consumer hardware friendly but still usable in case of necessity and most importantly open source. If the bench is true, it makes it the best model world wide opensource and nr 2 after chatgpt of OpenAI. However, Qwen has a path of integrity and responsibility to the users so few doubts about it´s performance and claims. I believe that this trend will impossible to keep soon so we at Hugston will consider/treat it as rare treasure, considering the packed knowledge beside the skills and abilities of this model. More updates follow in coming weeks and ofc as always you can run this model with HugstonOne.

Here we ask Qwen to code an svg of how it sees itself:

qwen-svg.png

The model is available for download here: https://www.modelscope.cn/models/Qwen/Qwen3.8-2.4T-A95B

With the benchmarks good to know *

  1. Fable5 results may involve fallbacks.
  2. Terminal Bench 2.1: Evaluated with Claude Code (avg@10), using a 5-hour timeout and max_tokens=131,072. For all other models, we report the best published score across harnesses: Claude Opus 4.8 and Claude Fable 5 with Terminus 2 from Artificial Analysis (https://artificialanalysis.ai/evaluations/terminalbench-v2-1); GPT-5.6 Sol with Codex (https://openai.com/index/previewing-gpt-5-6-sol/).
  3. SWE-bench Pro: Evaluated with the Claude Code harness, temp=1.0, top_p=0.95, and a 256K context window. Problematic tasks corrected and all baselines evaluated on the refined benchmark.
  4. DeepSWE 1.1: Evaluated with the Claude Code and mini-SWE-agent harnesses, temp=1.0, top_p=0.95, and a 256K context window. We report the highest score among both harnesses; notably, Qwen3.8-Max performs best on Claude Code.
  5. NL2Repo-Bench: Evaluated with the Claude Code harness. To prevent reward hacking, we disable Bash commands that attempt to access the specific repository, such as pip download, pip install, and git clone.
  6. FrontierSWE: Evaluated with the Claude Code harness. All other available MEAN@5 results are taken from the official FrontierSWE leaderboard (https://www.frontierswe.com) as of August 3, 2026. Dominance scores are recomputed from the raw scores using the official evaluation script. "--" indicates that no official MEAN@5 result was available as of that date.
  7. MLS-Bench-Lite: Evaluated with Claude Code using a 5-hour timeout and max_tokens=131,072. All other model scores are taken from the official leaderboard.
  8. PaperBench: Evaluated in the BasicAgent setting under Code-Dev mode, judged by Claude Opus 4.6, and averaged over 3 runs (max 12 hours per run).
  9. AndroidBench: Evaluated on the 95-task public subset, reporting avg@3 scores.
  10. QwenSWEBench: Inhouse coding benchmark to evaluate models' software engineering capabilities. Evaluated with the Claude Code harness. Reporting avg@3 with an 8-hour timeout, max_tokens=32,768, temperature=1.0, and a 256K-token context window.
  11. QwenQoderBench: Inhouse coding benchmark to evaluate user experience on Qoder. Evaluated with the Claude Code harness. Reporting avg@5 with a 6-hour timeout, max_tokens=32,768, temperature=1.0, and a 256K-token context window.
  12. QwenReactBench: Inhouse React project building benchmark using Claude Code as the harness, bilingual (EN/CN), 7 categories; auto-render + multimodal judge; BT/Elo rating.
  13. QwenSVGBench: Inhouse SVG code generation benchmark; bilingual (EN/CN), auto-render + multimodal judge; BT/Elo rating.
  14. CoWorkBench: Inhouse cowork benchmark for evaluating long-horizon tasks across computer science, finance, law, medical, and other productivity domains.
  15. SkillsBench: Evaluated on the public SkillsBench v1.1 benchmark across 87 tasks, reporting the average score over three runs per task. Opus 4.8 and Fable 5 are evaluated on Claude Code; GPT-5.6 Sol is evaluated on Codex; the Qwen-series are evaluated on OpenCode. All results are from our own testing.
  16. Automation-Bench: Evaluated on the 600-task public subset.
  17. WideSearch: Evaluated with the Claude Code harness for external models and the Qwen-Agent harness for ours, reporting the average item-F1 over four runs.
  18. $OneMillion-Bench: Evaluated using gemini-3.1-pro-preview.
  19. PLawBench: Evaluated using gemini-3.1-pro-preview.
  20. Empty cells (--): Scores are not yet available or are not applicable.

Comments

0 contributions

No comments yet.