nanochat
nanochat is Andrej Karpathy's open-source experimental harness for training and using a small language model end to end on one GPU node. The repository brings tokenizer training, pretraining, supervised fine-tuning, evaluation, reinforcement-learning experiments and inference into a compact, readable codebase. A single depth setting controls much of the model scale. It is a project name, not a generic term for every small chatbot.
Origin and context
Karpathy introduced nanochat in October 2025 with a runnable `speedrun` workflow and positioned it as the broader successor to nanoGPT, which primarily covered pretraining. The nanochat README also credits modded-nanoGPT's measured speedrun and leaderboard approach. Later guides, model miniseries and active repository changes replaced parts of the original launch recipe, so the launch discussion is historical context rather than current operating documentation.
Why it matters
Many LLM stacks split data preparation, training, post-training, evaluation and serving across large frameworks. nanochat keeps those stages close enough to inspect and modify as one experiment. That makes it useful for teaching and controlled systems research. Independent reuse goes beyond commentary: Hugging Face added a NanoChat model implementation to Transformers, and researchers wrapped nanochat's training loop to compare DiLoCo with conventional distributed data parallel training.
Example
A researcher changes an optimizer or data-loading rule, trains several depth-controlled models with the repository's experiment scripts, compares validation bits per byte and CORE results, then runs the fine-tuning and chat stages on the selected checkpoint. For downstream inference, the team can use nanochat's own engine or a compatible NanoChat model through Hugging Face Transformers. The exact script, commit, hardware, data and evaluation bundle must be recorded for a reproducible comparison.
How it differs
Scaling laws
nanochat includes scripts for model miniseries and scaling-law experiments, but one repository's depth sweeps do not establish a universal scaling law or a frontier-compute limit.
Mid-training
Mid-training is a stage or methodology. nanochat is a concrete codebase whose historical and current pipelines may implement particular intermediate training steps.
Post-training
Post-training covers broad methods for adapting a pretrained model. nanochat provides specific supervised and reinforcement-learning scripts, not a definition or exhaustive framework for the field.
Maturity and evidence
Maturity is 3. The project has a dated release, continuing development, a documented full workflow, large public reuse, an independent Transformers integration and research adaptation. It remains intentionally experimental and single-node focused; project interfaces and recommended recipes change quickly, and independent evidence does not establish production reliability or competitive model quality.
Limits and open questions
nanochat favors clarity and a strong baseline over broad hardware, model and configuration coverage. The current README warns that CPU or Apple-silicon examples produce much weaker models, while primary experiments target costly multi-GPU nodes. Dollar and time estimates vary with hardware prices, code revisions and the selected depth; the repository's slogan is not a reproducibility guarantee. Its speedrun leaderboard is maintained by the project, and reported capability depends on chosen metrics. The inherited `~2000 lines` description is obsolete: an independent November 2025 paper described its snapshot as roughly 8K lines, and the repository has continued evolving.
Related terms
References
- nanochat READMEAndrej Karpathy · 2025 · class A
- Introducing nanochat: The best ChatGPT that $100 can buyAndrej Karpathy · 2025-10-13 · class A
- nanoGPT READMEAndrej Karpathy · 2025-11 · class A
- NanoChat model documentationHugging Face · 2025-11-27 · class A
- What happens when nanochat meets DiLoCo?arXiv · 2025-11-14 · class B
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as pytorch skill.
This term is also covered in the Skills Atlas as large language models skill.
This term is also covered in the Skills Atlas as model training skill.
This term is also covered in the Skills Atlas as hugging face skill.
This term is also covered in the Skills Atlas as model evaluation skill.
This term is also covered in the Skills Atlas as distributed training skill.
This term is also covered in the Skills Atlas as reinforcement learning skill.