Technology
Developer trains 3.8B parameter language model under 1000 dollars, report says
A developer completed training a 3.8B-parameter model for $998, according to Hacker News.
The short version
- A developer reported training a 3.8B-parameter model on 65 billion tokens in 43 hours for $998.[Hacker News]
- The project was debugged using a 5090 GPU before completing the final training run on rented B200 hardware.[Hacker News]
- Value embeddings constituted 19% of the overall parameters in the final 3.8B-parameter model.[Hacker News]
- An earlier test on an 858M model achieved a 60.45% score on the PIQA benchmark across 16.4 billion tokens.[Hacker News]
Key facts
- A reported 3.8B-parameter model achieved a 0.384 CORE score after 43 hours of training on 65 billion tokens costing $998.[Hacker News]
- The codebase was debugged on an individual 5090 GPU and concluded training using rented B200 GPUs.[Hacker News]
- Value embeddings comprised 19% of the total parameter count in the 3.8B model.[Hacker News]
- A preliminary 858M parameter run on FineWeb-Edu reached 16.4 billion tokens and scored 60.45% on PIQA.[Hacker News]
What remains uncertain
- Performance benchmarks and budget figures stem from a single self-reported project and have not been independently validated.[Hacker News]
Sources
Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.