I am a fourth-year computer science PhD candidate at Harvard University, advised by Professor Demba Ba. I am supported by the Kempner Institute Graduate Fellowship. In summer 2026, I worked as a Research Fellow at .

I'm interested in the phenomenology of neural representations: their content and geometry, and the mechanisms and principles that give rise to them. Ultimately, I believe understanding model representations is essential for AI safety and alignment.

Previously, I worked at the AI Institute in Dynamic Systems with Nathan Kutz and Ryan Raut. I earned my B.S. in computer science from the Allen School at the University of Washington, where I worked with Rajesh Rao and William Noble.

Mozes Jacobs

Research

* equal contribution
Under review
Distance-to-target rulers
Goodfire A Phenomenology of Neural Geometry in Vision-Language-Action Models
Mozes Jacobs*, Aiden Swann*, Mathilde Papillon*, Siddharth Boppana, Fenil R. Doshi, Usha Bhalla, Thomas Icard, Matt Feiszli, Thomas McGrath, Vasudev Shyam, Demba Ba, Leon Bergen, Nina Miolane, Owen Lewis, Jack Merullo, Ekdeep Singh Lubana, Atticus Geiger, Matthew Kowal, Thomas Fel
We study the phenomenology of neural representations in VLA policies, building a taxonomy of how seven task variables are encoded across simulated and real-world robots, and find that physical constraints shape their geometry.
abstract

Robotic manipulation is inherently geometric, but what shape does the world take inside the policy that performs it? Prior work on vision-language-action (VLA) policies establishes that task-relevant variables are linearly readable from activations, yet readability certifies only that a variable is present and says nothing about the shape its values trace, which may be curved and multi-dimensional. Using controlled sweeps, natural rollouts, linear probes, and activation interventions, we build a taxonomy of seven task-relevant concepts in VLA models. We find that realized geometry is predominantly multi-dimensional and is shaped by the constraints under which the policy acts. Physical joint limits, for example, leave wrist orientation an open arc rather than a circle, and distance to the target, a single scalar, is realized in two orthogonal subspaces, a coarse ruler for the approach and a fine ruler for the grasp. The same geometries reappear in two further VLA models on real-world robot data, indicating that they are neither model-specific nor artifacts of simulation. We demonstrate how the findings are actionable by steering along the geometry a concept realizes. Overall, this work lays out an initial phenomenology of realized geometry in VLA policies and demonstrates that a concept's shape carries as much importance as its readability.

Under review
Block-sparse featurizers
Goodfire Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds
Thomas Fel*, Matthew Kowal*, Mozes Jacobs*, Dron Hazra*, Usha Bhalla*, Lee Sharkey, Lucius Bushnaq, Satchel Grant, Tal Haklay, Thomas Icard, Can Rager, Michael Pearce, Daniel Wurgaft, Aiden Swann, Fenil Doshi, Siddharth Boppana, Curt Tigges, Nick Cammarata, Thomas Serre, Vasudev Shyam, Owen Lewis, Thomas McGrath, Jack Merullo, Ekdeep Singh Lubana, Atticus Geiger
Block-sparse featurizers model concepts as low-dimensional manifolds rather than single directions, describing activations more compactly and revealing concept manifolds in InceptionV1, DINOv3, and SDXL.
arXiv  / 
abstract

What is the geometry of a visual percept? The most widely used protocols for decomposing neural network representations into interpretable parts treat concepts as isolated directions, yet recent work shows that concepts are often realized as geometric structures in low dimensional regions of activation space. We turn to the literature of structured sparsity to close this gap, and show that block sparsity, which groups directions into blocks, is the prior matched to a generative model in which a representation is a sparse sum of low-dimensional manifolds. We implement three variants of block-sparse featurizers (BSFs) and, through a minimum-description-length analysis, show that all three describe activations more compactly than direction-based featurizers, with the recovered concepts typically two- to four-dimensional. We then use BSFs to (i) recontextualize prior work, showing that curve detectors in InceptionV1 read from a single continuous curve manifold, (ii) discover novel manifolds including shadows and lighting in DINOv3, and (iii) support interpretable control of image generation in diffusion models (SDXL) via manifold steering.

Under review
Attention sinks
A Unifying View of Attention Sinks: From Mechanisms to Architectural Interventions
Lukas Fesser*, Mozes Jacobs*, Thomas Fel*, Andy Keller, Sham Kakade
Attention sinks hide two distinct algorithms, nop and broadcast; gating and registers each fix only one, and combining them gives complementary gains.
arXiv  / 
abstract

Attention sinks share a visual signature but hide two distinct algorithms: nop, where a head suppresses its update by routing to a null token, and broadcast, where a sink aggregates and redistributes global information. Each mechanism leaves distinct traces — nop sinks have negligible value norms; broadcast sinks induce low-rank outputs — which we use to derive practical diagnostics. Applied to pretrained vision transformers, we find both mechanisms coexist at scale. Gating and registers, the two dominant interventions, each implicitly target only one mechanism; combining them yields complementary gains. Training LeJepa with both gating and registers improves downstream semantic segmentation performance beyond either alone.

ICLR 2026
Raptor
Block-Recurrent Dynamics in ViTs
Mozes Jacobs*, Thomas Fel*, Richard Hakim*, Alessandra Brondetta, Demba Ba, T. Andy Keller
Trained ViTs are approximately block-recurrent: a 2-block recurrent surrogate recovers 96% of DINOv2 probe accuracy and enables dynamical interpretability.
arXiv  / 
abstract

We introduce the Block-Recurrent Hypothesis (BRH), arguing that trained ViTs admit a block-recurrent depth structure. To validate this, we train recurrent surrogates called Raptor. We demonstrate that a Raptor model can recover 96% of DINOv2 ImageNet-1k linear probe accuracy in only 2 blocks while maintaining equivalent runtime. We leverage our hypothesis to perform dynamical interpretability, revealing directional convergence into class-dependent basins, token-specific trajectory dynamics, and low-rank attractor structure in late layers.

CCN 2025 Oral
Traveling waves
Traveling Waves Integrate Spatial Information Through Time
Mozes Jacobs, Robert C. Budzinski, Lyle Muller, Demba Ba, T. Andy Keller
Recurrent networks that learn traveling waves integrate global spatial context, matching non-local U-Nets on segmentation with far fewer parameters.
pdf  /  blog  /  talk  / 
abstract

We investigate how traveling waves of neural activity enable spatial information integration in convolutional recurrent networks. Our models learn to generate traveling waves in response to visual stimuli, effectively expanding receptive fields of locally connected neurons. This mechanism significantly outperforms local feed-forward networks on semantic segmentation tasks requiring global spatial context, achieving comparable performance to non-local U-Nets while using significantly fewer parameters.

Preprint
Lorenz attractor
HyperSINDy: Deep Generative Modeling of Nonlinear Stochastic Governing Equations
Mozes Jacobs, Bingni W. Brunton, Steven L. Brunton, J. Nathan Kutz, Ryan V. Raut
A variational encoder and hypernetwork discover sparse stochastic governing equations from data, with uncertainty quantification.
arXiv  / 
abstract

HyperSINDy is a deep generative framework for discovering stochastic governing equations from data. A variational encoder and hypernetwork produce sparse differential equations — learned via a trainable binary mask — whose coefficients are driven by Gaussian white noise. HyperSINDy accurately recovers ground-truth stochastic dynamics and provides uncertainty quantification that scales to high-dimensional systems.

T-REX: Tied Recurrence Extraction
Mozes Jacobs, T. Andy Keller, Thomas Fel, Bingbin Liu, Richard Hakim, Yilun Du, Demba Ba
ICML 2026 Weight-Space Symmetries Workshop  /  OpenReview
Traveling Waves Integrate Spatial Information Into Spectral Representations
Mozes Jacobs, Robert C. Budzinski, Lyle Muller, Demba Ba, T. Andy Keller
ICLR 2025 Re-Align Workshop
Gradient Origin Predictive Coding
Mozes Jacobs, Linxing Preston Jiang, Rajesh N.P. Rao
Undergraduate senior thesis, 2022

Website template from Jon Barron