Technology
Tests show four-bit Qwen3.8 27B model matches full version while one-bit collapses
A 17 GB version runs on consumer graphics cards while 1-bit formats fail, according to Hacker News.
The short version
- Benchmarking of the 55 GB Qwen3.8 27B model found its 17 GB Q4_K_M four-bit quantization performs on par with the full model on coding tests.[Hacker News]
- The 17 GB four-bit format fits on standard 24 GB consumer GPUs like the RTX 4090, avoiding high-end hardware requirements.[Hacker News]
- Tests revealed that one-bit quantization causes accuracy to plunge to random chance on complex benchmarks, with extended reasoning worsening outcomes.[Hacker News]
- The reproducibility of the tests is limited because file provider Unsloth replaced the tested v2 quantization files with v3 versions.[Hacker News]
Key facts
- The full BF16 Qwen3.8 27B model requires 55 GB of memory, exceeding the capacity of standard consumer hardware.[Hacker News]
- The 17 GB Q4_K_M quantization fits within 24 GB graphics cards and matches the full model on the Terminal-Bench 2.1 coding benchmark.[Hacker News]
- One-bit quantization causes performance on GPQA Diamond to degrade to random chance levels, deteriorating further with extended reasoning.[Hacker News]
- Unsloth claimed its UD-1bit quantizations retain roughly 72 percent top-1 percent accuracy despite an 89 percent size reduction.[Hacker News]
What remains uncertain
- Exact reproduction of most test results is complicated because Unsloth replaced its v2 quantization files with v3 files on August 19, 2026.[Hacker News]
Sources
Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.