Latent Node

Independent AI research
  • Model capability
  • Training systems
  • Code and verification
  • Applied research

How close can a frozen 149M encoder get to a model built for typed decisions?

On our 400-case benchmark, TypeSafe's Jev scores 0.727, close to the teacher's self-agreement ceiling of 0.735, while a frozen 149M ModernBERT baseline scores 0.646. The encoder's probabilities are far closer to the gold distribution, with a KL divergence of 0.223 against Jev's 1.442.

Typed Decisions Sep 2026 Read the study

What if most of a model never read the prompt at all?

Grafting Qwen3.5-0.8B at layer 7 cuts time-to-first-token by 4.0x at 128K, decode unchanged, validation loss slightly better than the original. A tool-indexing probe then fell from 0.97 to 0.27 while loss, benchmarks and passkey retrieval all looked fine. A 15% code mix, frozen upper layers and a deeper cut recover full parity at 1.5x.

Model Grafting Sep 2026 Read the study

Can a model trained only on literature from before a discovery arrive at the discovery itself?

Trained on medicine published up to 1981, a 238M-parameter model proposes bacterial causation of peptic ulcers at a rate of 1.7%, p = 0.0006 against controls. Marshall and Warren published the hypothesis two years later.

Scientific Backtest Apr 2026 Read the study

Can you delete your dependencies and have a model write them instead?

A 9B model running on the machine solves 88% of a hundred-spec benchmark, and every function it writes is checked import-free by syntax tree. Across five real applications the auditable code surface falls by 13 times.

Conjure Jun 2026 Read the study

Why do some national AI programmes produce frontier models while others do not?

Across six countries, how much was invested predicts almost nothing. Coordination, differentiation and patient capital separate the three that succeeded from the three that did not.

Sovereign AI Feb 2026 Read the study

Is there a direction in an audio model's embedding space specific to Parkinson's, rather than to disordered speech in general?

One steering vector reaches 91.5% AUC and still rejects 53.5% of dysarthria from other causes, so it is not merely detecting abnormal speech.

PASS Jan 2026 Read the study

Can two ordinary Macs train a model that fits on neither of them?

Tensor-parallel QLoRA over a Thunderbolt bridge runs a 30B model at 964 tokens per second across an M3 Max and an M4, with LoRA sharded across column-parallel layers.

Dist-MLX Mar 2026 Read the study

Can a playable world model be learned from nothing but screen recordings?

79.2M parameters trained on phone gameplay video, with the actions recovered from the frames themselves, reconstruct the game at 27.9 dB PSNR and 0.83 SSIM.

World Model Mar 2026 Read the study

Do energy-based reasoning models out-reason language models?

On 718 Lean 4 theorems a 39M energy-based model proves 26.5% against 28.4% for a 494M language model, so the claim does not hold. Run together they reach 31.2%, which neither does alone.

EBRM Prover Mar 2026 Read the study

Which layers of a model actually need the extra bits?

Sensitivity to quantization varies 56-fold between layers. Allocating precision by measurement rather than uniformly recovers 17 points on GSM8K while using 57% less memory.

OptiQ Mar 2026 Read the study

How small can a transformer be and still do exact ten-digit addition?

392 parameters, at 99.4% exact match, against a previous record of 491.

Addition Transformer Feb 2026 Read the study

Sponsors get the code, data and checkpoints behind every study.

That includes the evaluation harnesses, the run logs, and the negative results that do not appear in the write-ups.

$10 once
An hour of compute, and a shoutout on Twitter.
$25 a month
Access to the private research repository, and new studies as they are published.
$100 a month
Repository access, plus your name or logo on this site and in every article.
Become a sponsor