let's build an llm · learning in public
From 1,330 verses to
a machine that answers
entry № 000
3 × DGX Spark
status: preparing
3 × DGX Spark
status: preparing
Part two of a from-scratch journey. Part one built an 815,744-parameter model on the Thirukkural and learned an honest lesson: tiny data teaches form, never meaning. This time: real pretraining, an assistant that gets smarter phase by phase, and — at the very end — kurals that mean something.
← part one: mykuralgpt.com — where this story startedNOT YET TAKEN
the 20-question exam · administration № 0
One fixed exam — simple facts, small tasks, one nonsense question — taken by every model this project trains. The scores go up as the data grows. Every transcript published, failures included.
PHASE 0
The smallest possible assistant — 135M parameters, 5B tokens, one exam. Up next.
PHASE 1
The modernization ladder — RoPE, RMSNorm, SwiGLU, GQA: every upgrade earned by ablation.
PHASE 2
The two-week run — 360M parameters, 40B tokens, two Sparks in harness.
PHASE 3
Base → assistant — SFT, then DPO. The caveat section of part one, made real.
PHASE 4
The capstone — a Tamil model, and the unfinished ending: kurals with meaning.