Proprioceptive AI shows targeted edits can improve LLM predictions
A startup claims lightweight neural probes can detect and suppress hallucinations in real time without retraining the underlying model
What if you could fix an AI model’s bad habits without ever touching the model itself? Proprioceptive AI, a startup founded by Logan Matthew Napolitano, is betting its entire thesis on that premise, developing tiny neural probes that monitor large language models from the inside and intervene before problematic outputs ever reach the user.
The company’s approach treats LLMs less like black boxes and more like patients on a heart monitor. Small diagnostic probes tap into the hidden states of a model, reading its internal dynamics in real time and making targeted corrections during inference. The result, according to the company’s reported benchmarks: an 85.8% reduction in confident-wrong outputs, the kind of hallucinations where a model states something incorrect with full conviction.
How the probes actually work
The core insight behind Proprioceptive AI’s technology is geometric. The company describes LLM internal representations as operating along two channels: a rank-1 primary channel that handles predictions, and a lower-dimensional “behavioral” channel that carries something like the model’s self-knowledge about its own potential outputs.
These probes are remarkably small. Each one contains roughly 1 million parameters and adds approximately 0.003% overhead to the host model’s parameter count. Latency sits in the sub-millisecond range, meaning the monitoring happens fast enough that users would never notice the extra computation.
The probes reportedly achieve separation ratios between 125x and 1,376x across different model families, including LLaMA and Mistral. That separation ratio measures how cleanly the probes can distinguish between safe and problematic internal states. The company also claims cross-model transfer capability, meaning probes trained on one architecture can generalize to others without retraining.
Why this matters for alignment
Proprioceptive AI’s framework leaves the base model frozen. No weights are changed during training or deployment. Interventions happen only at inference time, which means the underlying model retains its full capability set while the probes handle behavioral correction.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
The company reports internal benchmark scores between 0.96 and 0.999 AUC for detection efficacy. On the intervention side, per-token steering reportedly produced a 70 percentage point improvement on GSM8K tasks in gated experimental setups, a standard benchmark for mathematical reasoning.
The commercial play
Proprioceptive AI is building toward a commercial product called “The Cradle,” designed to provide behavioral adaptation and monitoring for models in the 3B to 32B parameter range.
The company has filed over 100 provisional patents, including methods built around Koopman operators, a mathematical framework borrowed from dynamical systems theory that allows nonlinear systems to be analyzed using linear techniques.
What to watch
The reported numbers are impressive, but they come with a significant caveat: all benchmarks disclosed so far are internal. Independent validation will be critical for establishing credibility with both the research community and potential enterprise customers. The company has indicated engagement with leading researchers, but peer-reviewed results would carry considerably more weight than company-reported metrics.
There is also the question of how these probes perform at scale beyond the 32B parameter ceiling currently targeted by The Cradle. The largest frontier models from OpenAI, Google, and Anthropic operate at significantly higher parameter counts. Whether the behavioral channels Proprioceptive AI has identified remain readable and actionable at that scale is an open question.