let's build an llm · learning in public

From 1,330 verses to
a machine that answers

entry № 000
3 × DGX Spark
status: preparing

Part two of a from-scratch journey. Part one built an 815,744-parameter model on the Thirukkural and learned an honest lesson: tiny data teaches form, never meaning. This time: real pretraining, an assistant that gets smarter phase by phase, and — at the very end — kurals that mean something.

← part one: mykuralgpt.com — where this story started
NOT YET TAKEN

the 20-question exam · administration № 0

One fixed exam — simple facts, small tasks, one nonsense question — taken by every model this project trains. The scores go up as the data grows. Every transcript published, failures included.

PHASE 1 The modernization ladder — RoPE, RMSNorm, SwiGLU, GQA: every upgrade earned by ablation.
PHASE 2 The two-week run — 360M parameters, 40B tokens, two Sparks in harness.
PHASE 3 Base → assistant — SFT, then DPO. The caveat section of part one, made real.
PHASE 4 The capstone — a Tamil model, and the unfinished ending: kurals with meaning.
3.9 1.9 training steps → the curves start soon