Moonfrost AI
A 777M-parameter Mixture-of-Experts model that I took from an empty folder all the way to a chat window. I trained the tokenizer, pretrained it twice, fine-tuned it to hold a conversation, scored it on six benchmarks and wrote the little app it talks through. The attention and expert layers follow DeepSeek-V2, rebuilt from the papers.
- 777M
- parameters, 161M active per token
- ~6B
- pretraining tokens (FineWeb-Edu)
- 2.7B
- chat fine-tuning tokens
- ~$55
- total training cost, ~12 H100-hours
- Python
- PyTorch
- MoE
- MLA
- FastAPI
- SSE
- SQLite
- Hugging Face