Distributed Training on Heterogeneous Macs
964 tok/sData-parallel and tensor-parallel QLoRA across an M3 Max and an M4 joined by Thunderbolt, including sharded LoRA for column-parallel layers.
A 238M-parameter model trained only on medical literature published before 1982, then tested on whether it proposes the bacterial cause of peptic ulcers before Marshall and Warren did.
A dependency-free programming paradigm. Describe a function in a YAML spec, let a model on your own machine write it, and verify by syntax tree that the result has no imports at all.
Six national AI programmes compared. Investment scale turns out to predict almost nothing about whether a country produces a frontier model.
A single direction in wav2vec2 embedding space that separates Parkinson's from other motor speech disorders, not merely from healthy speech.
Data-parallel and tensor-parallel QLoRA across an M3 Max and an M4 joined by Thunderbolt, including sharded LoRA for column-parallel layers.
A playable world model learned entirely from screen recordings of a phone game, with actions recovered from the video itself.
A head-to-head test of energy-based reasoning against a conventional language model on 718 Lean 4 theorems, and what happens when you combine them.
Data-driven per-layer bit allocation when converting PyTorch models to MLX. Four bits where it is safe, eight where it matters.
How small can a transformer get and still do exact arithmetic? Small enough to write out by hand, as it turns out.
Sponsors get the private research repository: full source for every study, trained checkpoints, datasets, evaluation harnesses, and the run logs and negative results that never make it into an article.
| $10 once | An hour of compute, and a shoutout on Twitter. |
|---|---|
| $25 / month | Full access to the private research repository, and new work as it lands. |
| $100 / month | Repository access, plus your name or logo on this site and in every article. |