Glossary · term

nanochat

nanochat is Andrej Karpathy's open-source experimental harness for training and using a small language model end to end on one GPU node. The repository brings tokenizer training, pretraining, supervised fine-tuning, evaluation, reinforcement-learning experiments and inference into a compact, readable codebase. A single depth setting controls much of the model scale. It is a project name, not a generic term for every small chatbot.

Karpathy2025-10-13ExternalMaturity: 3/5

Origin and context

Karpathy introduced nanochat in October 2025 with a runnable `speedrun` workflow and positioned it as the broader successor to nanoGPT, which primarily covered pretraining. The nanochat README also credits modded-nanoGPT's measured speedrun and leaderboard approach. Later guides, model miniseries and active repository changes replaced parts of the original launch recipe, so the launch discussion is historical context rather than current operating documentation.

Sources: s1, s2, s3

Why it matters

Many LLM stacks split data preparation, training, post-training, evaluation and serving across large frameworks. nanochat keeps those stages close enough to inspect and modify as one experiment. That makes it useful for teaching and controlled systems research. Independent reuse goes beyond commentary: Hugging Face added a NanoChat model implementation to Transformers, and researchers wrapped nanochat's training loop to compare DiLoCo with conventional distributed data parallel training.

Sources: s1, s4, s5

Example

A researcher changes an optimizer or data-loading rule, trains several depth-controlled models with the repository's experiment scripts, compares validation bits per byte and CORE results, then runs the fine-tuning and chat stages on the selected checkpoint. For downstream inference, the team can use nanochat's own engine or a compatible NanoChat model through Hugging Face Transformers. The exact script, commit, hardware, data and evaluation bundle must be recorded for a reproducible comparison.

Sources: s1, s4

How it differs

Scaling laws

nanochat includes scripts for model miniseries and scaling-law experiments, but one repository's depth sweeps do not establish a universal scaling law or a frontier-compute limit.

Mid-training

Mid-training is a stage or methodology. nanochat is a concrete codebase whose historical and current pipelines may implement particular intermediate training steps.

Post-training

Post-training covers broad methods for adapting a pretrained model. nanochat provides specific supervised and reinforcement-learning scripts, not a definition or exhaustive framework for the field.

Maturity and evidence

Maturity is 3. The project has a dated release, continuing development, a documented full workflow, large public reuse, an independent Transformers integration and research adaptation. It remains intentionally experimental and single-node focused; project interfaces and recommended recipes change quickly, and independent evidence does not establish production reliability or competitive model quality.

Sources: s1, s2, s4, s5

Limits and open questions

nanochat favors clarity and a strong baseline over broad hardware, model and configuration coverage. The current README warns that CPU or Apple-silicon examples produce much weaker models, while primary experiments target costly multi-GPU nodes. Dollar and time estimates vary with hardware prices, code revisions and the selected depth; the repository's slogan is not a reproducibility guarantee. Its speedrun leaderboard is maintained by the project, and reported capability depends on chosen metrics. The inherited `~2000 lines` description is obsolete: an independent November 2025 paper described its snapshot as roughly 8K lines, and the repository has continued evolving.

Sources: s1, s2, s5

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as pytorch skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as large language models skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model training skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as hugging face skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model evaluation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as distributed training skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as reinforcement learning skill.