← Latest briefing

Technology

Tests show four-bit Qwen3.8 27B model matches full version while one-bit collapses

A 17 GB version runs on consumer graphics cards while 1-bit formats fail, according to Hacker News.

The short version

  • Benchmarking of the 55 GB Qwen3.8 27B model found its 17 GB Q4_K_M four-bit quantization performs on par with the full model on coding tests.[Hacker News]
  • The 17 GB four-bit format fits on standard 24 GB consumer GPUs like the RTX 4090, avoiding high-end hardware requirements.[Hacker News]
  • Tests revealed that one-bit quantization causes accuracy to plunge to random chance on complex benchmarks, with extended reasoning worsening outcomes.[Hacker News]
  • The reproducibility of the tests is limited because file provider Unsloth replaced the tested v2 quantization files with v3 versions.[Hacker News]

Key facts

  • The full BF16 Qwen3.8 27B model requires 55 GB of memory, exceeding the capacity of standard consumer hardware.[Hacker News]
  • The 17 GB Q4_K_M quantization fits within 24 GB graphics cards and matches the full model on the Terminal-Bench 2.1 coding benchmark.[Hacker News]
  • One-bit quantization causes performance on GPQA Diamond to degrade to random chance levels, deteriorating further with extended reasoning.[Hacker News]
  • Unsloth claimed its UD-1bit quantizations retain roughly 72 percent top-1 percent accuracy despite an 89 percent size reduction.[Hacker News]

What remains uncertain

  • Exact reproduction of most test results is complicated because Unsloth replaced its v2 quantization files with v3 files on August 19, 2026.[Hacker News]

Sources

Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.