← Latest briefing

Technology

Developer trains 3.8B parameter language model under 1000 dollars, report says

A developer completed training a 3.8B-parameter model for $998, according to Hacker News.

The short version

  • A developer reported training a 3.8B-parameter model on 65 billion tokens in 43 hours for $998.[Hacker News]
  • The project was debugged using a 5090 GPU before completing the final training run on rented B200 hardware.[Hacker News]
  • Value embeddings constituted 19% of the overall parameters in the final 3.8B-parameter model.[Hacker News]
  • An earlier test on an 858M model achieved a 60.45% score on the PIQA benchmark across 16.4 billion tokens.[Hacker News]

Key facts

  • A reported 3.8B-parameter model achieved a 0.384 CORE score after 43 hours of training on 65 billion tokens costing $998.[Hacker News]
  • The codebase was debugged on an individual 5090 GPU and concluded training using rented B200 GPUs.[Hacker News]
  • Value embeddings comprised 19% of the total parameter count in the 3.8B model.[Hacker News]
  • A preliminary 858M parameter run on FineWeb-Edu reached 16.4 billion tokens and scored 60.45% on PIQA.[Hacker News]

What remains uncertain

  • Performance benchmarks and budget figures stem from a single self-reported project and have not been independently validated.[Hacker News]

Sources

Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.