I am a fourth-year PhD candidate at Harvard University, advised by Professor Demba Ba. I am supported by the Kempner Institute Graduate Fellowship. In summer 2026, I worked as a Research Fellow at .

I work on mechanistic interpretability and representation learning. I'm interested in the phenomenology of neural representations: their content and geometry, and the mechanisms and principles by which they arise. Ultimately, I believe understanding model representations is essential for AI safety and alignment.

Previously: AI Institute in Dynamic Systems (Nathan Kutz, Ryan Raut)  ·  B.S. Computer Science, Allen School, University of Washington (Rajesh Rao, William Noble)

Mozes Jacobs

Research

* equal contribution  ·  ‡ core contributor
Under review
Distance-to-target rulers
Goodfire A Phenomenology of Neural Geometry in Vision-Language-Action Models
Mozes Jacobs*, Aiden Swann*, Mathilde Papillon*, Siddharth Boppana, Fenil R. Doshi, Usha Bhalla, Thomas Icard, Matt Feiszli, Thomas McGrath, Vasudev Shyam, Demba Ba, Leon Bergen, Nina Miolane, Owen Lewis, Jack Merullo, Ekdeep Singh Lubana, Atticus Geiger, Matthew Kowal, Thomas Fel
We study the phenomenology of neural representations in VLA policies, building a taxonomy of how seven task variables are encoded across simulated and real-world robots, and find that physical constraints shape their geometry.
abstract

Robotic manipulation is inherently geometric, but what shape does the world take inside the policy that performs it? Prior work on vision-language-action (VLA) policies establishes that task-relevant variables are linearly readable from activations, yet readability certifies only that a variable is present and says nothing about the shape its values trace, which may be curved and multi-dimensional. Using controlled sweeps, natural rollouts, linear probes, and activation interventions, we build a taxonomy of seven task-relevant concepts in VLA models. We find that realized geometry is predominantly multi-dimensional and is shaped by the constraints under which the policy acts. Physical joint limits, for example, leave wrist orientation an open arc rather than a circle, and distance to the target, a single scalar, is realized in two orthogonal subspaces, a coarse ruler for the approach and a fine ruler for the grasp. The same geometries reappear in two further VLA models on real-world robot data, indicating that they are neither model-specific nor artifacts of simulation. We demonstrate how the findings are actionable by steering along the geometry a concept realizes. Overall, this work lays out an initial phenomenology of realized geometry in VLA policies and demonstrates that a concept's shape carries as much importance as its readability.

Under review
Block-sparse featurizers
Goodfire Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds
Thomas Fel*, Matthew Kowal*, Mozes Jacobs*, Dron Hazra*, Usha Bhalla*, Lee Sharkey, Lucius Bushnaq, Satchel Grant, Tal Haklay, Thomas Icard, Can Rager, Michael Pearce, Daniel Wurgaft, Aiden Swann, Fenil Doshi, Siddharth Boppana, Curt Tigges, Nick Cammarata, Thomas Serre, Vasudev Shyam, Owen Lewis, Thomas McGrath, Jack Merullo, Ekdeep Singh Lubana, Atticus Geiger
Block-sparse featurizers model concepts as low-dimensional manifolds rather than single directions, describing activations more compactly and revealing concept manifolds in InceptionV1, DINOv3, and SDXL.
arXiv  / 
abstract

What is the geometry of a visual percept? The most widely used protocols for decomposing neural network representations into interpretable parts treat concepts as isolated directions, yet recent work shows that concepts are often realized as geometric structures in low dimensional regions of activation space. We turn to the literature of structured sparsity to close this gap, and show that block sparsity, which groups directions into blocks, is the prior matched to a generative model in which a representation is a sparse sum of low-dimensional manifolds. We implement three variants of block-sparse featurizers (BSFs) and, through a minimum-description-length analysis, show that all three describe activations more compactly than direction-based featurizers, with the recovered concepts typically two- to four-dimensional. We then use BSFs to (i) recontextualize prior work, showing that curve detectors in InceptionV1 read from a single continuous curve manifold, (ii) discover novel manifolds including shadows and lighting in DINOv3, and (iii) support interpretable control of image generation in diffusion models (SDXL) via manifold steering.

Under review
Proprioceptive-state feature manifold
Goodfire Multi-Dimensional Featurizers Yield Fine-Grained Steering and Data Shaping in Robot Foundation Models
Aiden Swann‡, Mozes Jacobs‡, Mathilde Papillon‡, Siddharth Boppana, David Chanin, Dron Hazra, Usha Bhalla, Curt Tigges, Mac Schwager, Owen Lewis, Jack Merullo, Ekdeep Singh Lubana, Atticus Geiger, Thomas Fel, Matthew Kowal
Block-sparse featurizers recover multi-dimensional features in VLA models, enabling fine-grained steering and isolating the <1.4% of training tokens behind a specific behavior.
abstract

Can better feature discovery improve control over robot behavior through internal interventions and training data? Recent work in vision and language models reveals multidimensional features, motivating block-sparse featurizers (BSFs), which represent each feature in a low-dimensional subspace. We investigate whether recovering this structure provides an interpretable interface for steering vision-language-action models and shaping the data from which they learn. We show that BSFs recover several multi-dimensional task-relevant features underlying the model's computation (e.g., perception, language, proprioceptive state, episode phase, and control variables). We causally validate these features through fine-grained steering: targeted changes in visual perception, semantic representation, and spatial localization produce predictable changes in downstream action prediction. Finally, we use feature-based data attribution to isolate where a specific behavior (failure recovery) arises in the training dataset, identifying fewer than 1.4% of training tokens whose masking is sufficient to remove the targeted behavior. Overall, our results show how matching featurizers to representation geometry can provide more effective tools for understanding and changing robot behavior.

Under review
Robot arm redirected between a banana and an apple
Goodfire Causal Analysis of Robot Foundation Models Enables Efficiency Gains
Aiden Swann‡, Mathilde Papillon‡, Mozes Jacobs‡, Siddharth Boppana, Tal Haklay, Eric Bigelow, Andrew Lee, Conor Watts, Thomas Icard, Mac Schwager, Owen Lewis, Jack Merullo, Ekdeep Singh Lubana, Atticus Geiger, Thomas Fel, Matthew Kowal
Activation patching in three VLA models reveals a narrow band of layers through which perception and language condition action; keeping only that interface cuts compute by 33% at matched success.
abstract

Vision-language-action (VLA) models are a promising approach to general purpose robot manipulation. While the internal computation of vision-language models, which compose the backbone of VLAs, has been studied extensively, less is known about how their perceptual and semantic representations are transformed into action in VLA models. We study this question mechanistically in three frontier VLA models (π0.5, MolmoAct2, and MolmoBot) by analyzing the causal pathways through which the input image and the instruction shape actions. Across the three models, activation patching shows that causally relevant information influences the action expert (AE) through a narrow band of mid-to-late cross-attention layers, which we term the action-conditioning interface. The content of this information changes with token and layer within the model. Early-to-mid layer representations of the target noun token determine which object is selected, whereas mid-to-late layers carry the spatial information in text tokens that redirects the movement when patched. Where spatial information is stored in prompt tokens varies based on the architecture. In MolmoAct2 the model forms task-conditioned representations at the text positions that follow the instruction, whereas π0.5 additionally stores this information at image token positions. We show that these findings can be translated directly into improving model efficiency. We find that restricting MolmoBot's action expert conditioning to 3/36 layers reduces inference FLOPs by up to 33% while maintaining comparable task success with the default model across more than 3,000 evaluation episodes (46.7% vs. 46.9%), and the same principle carries over to real-robot deployment of π0.5.

Under review
Attention sinks
A Unifying View of Attention Sinks: From Mechanisms to Architectural Interventions
Lukas Fesser*, Mozes Jacobs*, Thomas Fel*, Andy Keller, Sham Kakade
Attention sinks hide two distinct algorithms, nop and broadcast; gating and registers each fix only one, and combining them gives complementary gains.
arXiv  / 
abstract

Attention sinks share a visual signature but hide two distinct algorithms: nop, where a head suppresses its update by routing to a null token, and broadcast, where a sink aggregates and redistributes global information. Each mechanism leaves distinct traces — nop sinks have negligible value norms; broadcast sinks induce low-rank outputs — which we use to derive practical diagnostics. Applied to pretrained vision transformers, we find both mechanisms coexist at scale. Gating and registers, the two dominant interventions, each implicitly target only one mechanism; combining them yields complementary gains. Training LeJepa with both gating and registers improves downstream semantic segmentation performance beyond either alone.

ICLR 2026
Raptor
Block-Recurrent Dynamics in ViTs
Mozes Jacobs*, Thomas Fel*, Richard Hakim*, Alessandra Brondetta, Demba Ba, T. Andy Keller
Trained ViTs are approximately block-recurrent: a 2-block recurrent surrogate recovers 96% of DINOv2 probe accuracy and enables dynamical interpretability.
arXiv  / 
abstract

We introduce the Block-Recurrent Hypothesis (BRH), arguing that trained ViTs admit a block-recurrent depth structure. To validate this, we train recurrent surrogates called Raptor. We demonstrate that a Raptor model can recover 96% of DINOv2 ImageNet-1k linear probe accuracy in only 2 blocks while maintaining equivalent runtime. We leverage our hypothesis to perform dynamical interpretability, revealing directional convergence into class-dependent basins, token-specific trajectory dynamics, and low-rank attractor structure in late layers.

CCN 2025 Oral
Traveling waves
Traveling Waves Integrate Spatial Information Through Time
Mozes Jacobs, Robert C. Budzinski, Lyle Muller, Demba Ba, T. Andy Keller
Recurrent networks that learn traveling waves integrate global spatial context, matching non-local U-Nets on segmentation with far fewer parameters.
pdf  /  blog  /  talk  / 
abstract

We investigate how traveling waves of neural activity enable spatial information integration in convolutional recurrent networks. Our models learn to generate traveling waves in response to visual stimuli, effectively expanding receptive fields of locally connected neurons. This mechanism significantly outperforms local feed-forward networks on semantic segmentation tasks requiring global spatial context, achieving comparable performance to non-local U-Nets while using significantly fewer parameters.

Preprint
Lorenz attractor
HyperSINDy: Deep Generative Modeling of Nonlinear Stochastic Governing Equations
Mozes Jacobs, Bingni W. Brunton, Steven L. Brunton, J. Nathan Kutz, Ryan V. Raut
A variational encoder and hypernetwork discover sparse stochastic governing equations from data, with uncertainty quantification.
arXiv  / 
abstract

HyperSINDy is a deep generative framework for discovering stochastic governing equations from data. A variational encoder and hypernetwork produce sparse differential equations — learned via a trainable binary mask — whose coefficients are driven by Gaussian white noise. HyperSINDy accurately recovers ground-truth stochastic dynamics and provides uncertainty quantification that scales to high-dimensional systems.

T-REX: Tied Recurrence Extraction
Mozes Jacobs, T. Andy Keller, Thomas Fel, Bingbin Liu, Richard Hakim, Yilun Du, Demba Ba
ICML 2026 Weight-Space Symmetries Workshop  /  OpenReview
Traveling Waves Integrate Spatial Information Into Spectral Representations
Mozes Jacobs, Robert C. Budzinski, Lyle Muller, Demba Ba, T. Andy Keller
ICLR 2025 Re-Align Workshop
Gradient Origin Predictive Coding
Mozes Jacobs, Linxing Preston Jiang, Rajesh N.P. Rao
Undergraduate senior thesis, 2022

Website template from Jon Barron