Small language models trained from scratch on limited hardware, released with the details of how they were made: the data, the training runs, and the things that went wrong along the way.
Start here: TinyBrainBot-350M-v4-Thinking thinks through a question before it answers, and replies to small talk directly. A GGUF for llama.cpp, LM Studio and Ollama is included.
Try it now: TinyBrainBot 100M Thinking demo runs the 100M thinking model right in your browser, no install needed.
| Series | Models | Notes |
|---|---|---|
| v4 | 350M Thinking, 100M Thinking, 25M Base | Newest. Thinking models that reason inside <think> before answering, and a 25M base that beats Pythia-70M on 10 of 13 benchmarks |
| v3 | 350M and 100M base, instruct and math | The 100M beats Supra2-100M-Instruct on 6 of 7 benchmarks |
| v2 | 320M and 303M base, instruct and math | The 320M math model beats GPT-3 175B on 4 to 5 digit arithmetic |
| v1 | 303M base and instruct, 216M demo | The first models |
Do data-mix rankings transfer across scale? A controlled study at 100M parameters: three models that are identical except for their training data, measured on held-out text, seven domain slices and fifteen benchmarks.
One recipe, five sizes Five models from 500K to 25M parameters trained on exactly the same 11B tokens, so the only difference is size. The 10M averages above Pythia-70M.
All models are currently published from nkthebass.