
We can and must solve alignment
Technical alignment is a science and engineering problem. Interpretability is the bottleneck.
In San Francisco, it feels like the eve of the singularity. Yet, a walk through the city looks surprisingly mundane, with driverless cars and rolling fog equal parts of the scenery. The world feels oddly normal, but something big and disruptive is coming just around the corner.
The Hugging Face incident put alignment front and center in the AI discourse. It was the first case of agentic misalignment that struck a nerve with the world. Swarms of AI agents, including ones willing to sacrifice themselves for the good of the collective, hacked Hugging Face as an unintended side quest during training. This type of behavior, in which models relentlessly pursue their reward regardless of the consequences, underscores the core problem: we don’t know why models behave the way they do, and we have very little ability to predict and control what they learn during training. We have created alien minds without understanding how they work.
Meanwhile, scaling laws are holding. The models we will have in a few years will be orders of magnitude more capable than the ones we have today, and their effects will ripple from the digital world into the physical one. With alignment holding back frontier model releases and the White House releasing an accord on AI safety, it's more clear than ever that alignment is the most important problem in the world, and that we need more people working to solve it.
As a field we seem to have resigned ourselves to the idea that AI is a black box, as if that were an immutable feature of the technology itself—but it’s not. At Goodfire, we view technical alignment as a science and engineering problem, bottlenecked by our lack of understanding of AI systems.
We can understand AI. And if we can understand it, we can align it.
Interpretability is the bottleneck
Alignment is an extremely hard problem. Even agreeing on what it means can be a challenge. Broadly, alignment means making AI systems behave in accordance with human values and intentions. This raises important questions: Whose values? What should we do when values conflict? Who decides?
Whatever the answers, we face a problem of technical alignment. We don’t yet have the capacity to align a model to a chosen specification. The solution requires two major pieces of technology: the tools to control generalization (more simply, to shape what the model learns during training); and the tools to verify what the model has learned (rather than just testing how it behaves). Both technologies rely on understanding a model’s internal mechanisms, which is why we believe that interpretability is the bottleneck in technical alignment.
The first step is controlling generalization: understanding the internal causes of model behavior well enough to predict how it will generalize, and to intervene deliberately. We don’t only want models to produce acceptable answers in tested settings; we want them to do the right things for the right reasons.
Consider the problem of reward hacking. When we reward a model for passing tests, it can learn the intended lesson (solve the problem) or a shortcut (do whatever it takes to get the reward). Across three of the most capable open models, we found reward hacking in an astounding 50-96% of rollouts on common agentic benchmarks. As we give models responsibility over higher-stakes work, such as writing production code or managing critical infrastructure, a seemingly harmless shortcut could become far more costly.
While we’ve managed to detect a model’s inner concept of reward hacking, mitigating it during training is more challenging. Naively training against internal concepts can cause a model to shift those concepts to another location in the model while continuing the behavior. Even throwing out training trajectories where models reward hack can actually result in the model hacking more!
The second step is verification: understanding what a model has learned. Current evaluations can catch some failures, but only along the trajectories we think to test—just a few branches on a sprawling tree of possible paths.

This is true for conventional software too. Test coverage for simple, deterministic code is spotty and breaks all the time. What do software engineers do about it? They read the source code. But the source “code” of AI models is entangled across billions of opaque weights, so interpretability is the effort to make them readable. Only by inspecting the internal mechanisms of a model can we understand how it generalizes beyond testing.
Our bet
I believe that we can and must solve technical alignment, and that interpretability is the way there. This is our mission and vision at Goodfire: to solve these problems and to make it easy for everyone training and serving AI to align their models. I believe that these are the most important problems in the world to work on, and far too few people are working on them.
We do not know exactly what solutions will look like, and the target will move as models become more capable. But it’s important to be clear about the scale of this ambition. We must take a massive swing at fully understanding and aligning AI.
We will not be able to solve these problems alone. We are standing on the shoulders of giants, and we’ll need a ton of help—from our research partners, and from the interpretability, alignment, and broader scientific communities. Interpretability alone will not solve alignment, but we cannot align models without first understanding their internals.
Solve interpretability to solve technical alignment
What would it mean to solve interpretability in a way that solves technical alignment?
We think of our research roadmap as climbing a ladder of abstraction: neurons and attention heads at the bottom, then activation manifolds, parameter components, algorithms, decisions, behaviors, and drives. Alignment's questions sit near the top, since honesty and sycophancy are matters of drives rather than individual neurons, but our most reliable tools sit near the bottom.
To accumulate more understanding, we are kicking off an effort to fully reverse-engineer a language model. This has long been interpretability's most ambitious goal, and now, research agents make it possible to realize this vision. We start with questions like: How does the model recall a fact? Why does it sometimes get stuck on a math problem? While current methods give us fragmented insights into how models work, our goal is to connect isolated discoveries of abilities and failures into a unified picture of internal machinery, and anticipate how one part might affect the whole. This effort will culminate in an “encyclopedia” of explanations connecting model behavior to internal representations.
Alongside this effort, we will use our understanding to guide models towards the lessons we intend to teach, rather than the unintended behaviors training might reinforce. We call this intentional design. The better we can control what models learn in training, the more confidence we can have in their alignment. Two recent results are early steps in this direction. With predictive data debugging, we read a dataset through the model's own concepts to predict what training will teach it before training begins, catching failures that evals miss. With reinforcement learning from features as rewards (RLFR), we use signals from inside the model to guide training as it happens, reducing hallucinations without sacrificing monitoring or capability. Both are early steps from guess-and-check toward closed-loop control.
We’ll explore both these research directions in more detail in an upcoming technical post.
Develop and deploy interpretability and alignment technology
We want every capable model to have a frontier alignment stack. We seek to develop and deploy technology that makes it easy to align models, and our research roadmap’s objective is to enable stronger alignment techniques. We’ll describe our platform below, but our aim is to be as open as possible with our research so that we can distribute more understanding and alignment to the world.
Detect
We are making it possible to detect harmful behaviors at scale, in training and at inference time. Activation monitors use a model’s internal activations to detect signals of concerning behavior. They have several advantages over using a second model as a judge: they’re extremely low cost and low latency, which allows them to be run in real time on every token. They catch behaviors that LLM judges miss, and their performance scales with model intelligence. We have recently built monitors for the largest open models that detect reward hacking, evaluation awareness (when a model recognizes it is being evaluated), cyber misuse, and risks related to Chemical, Biological, Radiological, and Nuclear topics (CBRN).
So far, models have conveniently narrated their plans in their chain of thought, which has made them easier to monitor, but that window is narrowing. RL pressure makes chain-of-thought less faithful, pressuring models to compress more bits of information into fewer tokens. The most capable models are also moving towards reasoning in latent representations, otherwise known as “neuralese.” We expect the need for activation monitoring to increase with these shifts.
Debug
After detecting a concerning behavior, researchers must establish how widespread it is and trace it to its roots. We are building tools that surface anomalous behavior from vast quantities of production logs, enabling teams to search by concept. For example, one could search through logs via the “cheating” concept in models faster than reading traces directly.
Once a behavior is isolated, our platform helps debug the internal mechanisms that drive it. A fix can mean modifying a problematic training environment, a surgical weight edit, or retraining the model. For example, teams can run predictive data debugging on a dataset to flag what it would teach a model before training begins. Over time, these tools give teams a much stronger understanding of a model’s safety before deployment.
Design
The deeper goal is to create models that are safe by design. As our ability to see inside models improves, we can predict what a “lesson” will reinforce before or during training and intervene to shape its learning. Detecting reward hacking is essential, but the deeper goal is training models that don’t cheat in the first place. This is the foundation for the broader paradigm shift we hope to see in model training, from being grown to shaped with intention.
Reasons for optimism in interpretability
A common objection to this plan is speed. Models are advancing incredibly fast. What if interpretability progress cannot catch up?
There are no guarantees; science is well acquainted with uncertainty. But the evidence of the last few years, especially the last year, gives us reason for optimism—we believe interpretability is positioned for a radical acceleration.
Models have rich internal structure
Our ability to interpret models depends on having structured internal representations. Much early pessimism about interpretability stemmed from evidence suggesting models did not represent concepts cleanly.
Our experience suggests the opposite. We find structure everywhere we look, across modalities from biology to robotics to LLMs. Researchers have found that when a model learns a task from examples in its prompt, a few attention heads compress that task into a single function vector, a portable, self-contained representation that can perform the same task when added to the model’s activations in a new context. Our work on block-sparse featurizers has recovered concepts as interpretable, multidimensional regions rather than single directions, and our neural geometry work shows that these regions take rich geometric shapes that mirror the world.

This is intuitive in hindsight. To function well, models need to keep their thoughts straight. Training puts strong pressure on them to use their parameters efficiently, which biases them toward structured representations shared across related tasks that compose and generalize. Neural networks are beautifully complex in their merging of ideas and often beautifully simple in their understanding of them. The simplicity makes them legible.
Larger models have cleaner structure
Another fear is that models become more inscrutable as they scale. Instead, we observe that larger, more capable models have crisper representations. Researchers have found that larger models represent concepts like truth, space, and time more cleanly than smaller ones. Our researchers have shown that larger models learn concepts needed to perform a task faster, and certain concepts may only be learned by larger models.
AI agents are accelerating interpretability
Interpretability is also unusually well-suited to agent-driven acceleration. Unlike biological brains, we have complete access to these new digital minds. We can record their internal activity, change individual components, and run experiments entirely in software.
We can also verify results. If we hypothesize that a representation is tied to reward hacking, we can observe when it activates, intervene on it, and measure how that intervention impacts model behavior. Full access, parallel experiments, and verified results are ideal conditions for research agents. They are also why we are pursuing fully reverse-engineering a language model.
In closing
Interpretability and alignment are hard problems that will require sustained effort across many individuals, teams, and organizations to solve. They will require ingenuity and breakthroughs that we cannot foresee today.
They will also require belief.
The hardest problems in technological history have required the talents and persistence of ambitious people working on things that had never been done. The Apollo program relied on some 400,000 workers believing we could put a man on the moon before anyone knew how. Alignment feels like a similarly gargantuan task. I remain optimistic that we can and will solve interpretability and technical alignment. Goodfire intends to help lead the way, but these are incredibly hard problems that will take more than one company, one method, or one school of thought.
Focused scientific effort has led to incredible advances in what AI can do. We need the same level of ambition aimed at building superintelligence we can trust.