Technology
AI startup Magic claims major training efficiency gains over leading open models
Hacker News reports that Magic claims its pretraining process requires substantially less compute than rival base models.
The short version
- Magic announced that its new pretraining method exceeds the compute efficiency of leading open-weight base models by more than tenfold.[Hacker News]
- The startup claims it matched DeepSeek V4 Pro Base performance using about 50 times fewer floating-point operations.[Hacker News]
- Scaling the model further cost approximately $4 million and reportedly surpassed existing public base models on perplexity benchmarks.[Hacker News]
- Magic noted that achieving similar results under DeepSeek's approach would cost over $100 million based on scaling laws, though backend evaluation issues were reported.[Hacker News]
Key facts
- Magic stated that its pretraining recipe is more than 10 times more compute-efficient than leading open-weight base models.[Hacker News]
- The company reported matching DeepSeek V4 Pro Base performance using roughly 50 times fewer FLOPs.[Hacker News]
- Magic reported scaling training tenfold to roughly $4 million, outperforming available open base models on perplexity evaluations.[Hacker News]
- According to scaling laws cited by Magic, training a comparable model using DeepSeek V4 Pro's recipe would exceed $100 million.[Hacker News]
What remains uncertain
- Magic experienced backend issues during logprob evaluation across vLLM and SGLang setups, and the efficiency gains have not been independently verified.[Hacker News]
Sources
Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.
- >10x More Efficient PretrainingHacker News