nanochat-omni

Author	SHA1	Message	Date
Fam Zheng	3c1cc3302f	omni: W1 audio align smoke — synthetic dataset + 50-step script End-to-end smoke proving the audio path: wav -> WhisperEncoder (frozen) -> Projector -> prepend to text embeddings -> tiny d6 GPT (random init) -> CE loss on text only Pass criterion is a plain "loss drops by at least 0.5". On a 4090 the run finishes in ~1 s and goes 5.55 -> 0.17 over 50 steps, so the threshold has plenty of headroom against false positives. Two design calls worth keeping in mind: 1. Synthetic sine clips, not LibriSpeech. W1 is forward-path proof, not alignment quality, and a deterministic offline dataset means no network on the smoke path. data/audio_smoke/manifest.jsonl is the only thing committed; wavs are regenerated by audio_smoke_data.py and gitignored. W2 swaps in real LibriSpeech. 2. Standalone byte-level tokenizer (UTF-8 bytes + a single BOS, vocab=257). Avoids depending on a trained nanochat BPE — the d6 GPT is random anyway, so vocab choice doesn't matter for "does the gradient flow" smoke. W2 onwards uses the real BPE on a real base. Caveat documented in doc/todo.md: because the LM is also random and being trained, the loss-down here mostly reflects the LM memorising 5 short strings, not Whisper-Projector alignment. That's fine for proving plumbing; W2 freezes the LM so projector-only gradient is the only path to lower loss.	2026-05-05 22:39:20 +01:00
Fam Zheng	b585e07dc2	omni: CI smoke + docs + README preamble smoke / nanochat-smoke (push) Successful in 2m30s Details - .gitea/workflows/smoke.yml: gitea CI on ailab gpu runner (manual git clone since actions/checkout@v4 mis-resolves subpath gitea); injects WANDB_API_KEY + CI_RUN_TAG=smoke-$run_number - scripts/smoke.sh: in-place smoke (uv sync + 1 shard + tokenizer + d=6 50-step base_train); idempotent cache at /data/nanochat-smoke/ - doc/research_feasibility.md: voice-first multimodal feasibility study (mochi) - doc/todo.md: phase-by-phase roadmap (W1 Whisper smoke → W4 MVP) - README.md: omni preamble pointing at upstream nanochat README - .gitignore: exclude .claude/ runtime files	2026-05-05 22:21:31 +01:00
Andrej Karpathy	64a651a63c	include .claude is ok	2026-01-29 00:35:02 +00:00
Andrej Karpathy	f5a0ea4d3f	take out these gitignore dirs	2026-01-08 18:18:42 +00:00
Andrej Karpathy	ccf4b7f9bf	nudge hyperparameters of the base script with the results of the sweeps and miniseries. vocab size down to 32K. D:N ratio from 20 to 8. add miniseries script	2026-01-07 22:11:59 +00:00
Andrej Karpathy	ed2082fbc4	sane secrets management	2026-01-04 19:29:22 +00:00
Andrej Karpathy	aa42f40e66	delete the inline rustbpe project. it was ugly to have a project within project and rustbpe is now nicely a separate repo on my github karpathy/rustbpe and it's on pypi etc., so we just add it as a depedency to uv. i think it is appropriate that this is a separate repo because 1) it doesn't have too many knobs, other than the ones that are exposed - the regex pattern and vocab size and 2) all of its complexity is not algorithmic (it's equivalent to minbpe), instead it is efficiency-related, so it is ok to hide relatively speaking	2026-01-03 23:55:28 +00:00
Luke Stanley	760af62e11	Git ignore eval_bundle	2025-10-21 23:14:34 +00:00
Andrej Karpathy	fe5aed940b	add personality to nanochat. breaks previous code on git pull and requires download of a new file from s3, but there is a helpful error message so hopefully its ok	2025-10-21 15:04:58 +00:00
karpathy	3a5e0bc50b	initial commit	2025-10-13 06:49:24 -07:00

10 Commits