The Three Numbers: Qwen 3.8 Fits in 24 GB, Runs 2.24x Faster, and Loses 2.1 Points When Uncensored

The Three Numbers: Qwen 3.8 Fits in 24 GB, Runs 2.24x Faster, and Loses 2.1 Points When Uncensored

0 View

Publish Date:
28 September, 2026
Category:
MSNBC
Video License
Standard License
Imported From:
Youtube

By Evan Vega

In five weeks, Qwen 3.8 27B moved onto a desk, got twice as fast and had its refusals removed. 24 GB: the native vision-language model runs a 262K-token window in laptop memory because 48 of its 64 layers never cache a token. 2.24x: MTPLX’s multi-token prediction on an M5 Max, with the output distribution unchanged. 2.1 points: the MMLU cost of abliteration, published by exactly one of four builds. Every figure is attributed to whoever measured it.


Read the full investigation →

Related: Frontier Watch

Read the full file →