Machine learning · AI safety

On Problems That Will Outlive Our Models

Which difficulties in machine learning safety would survive a complete change of methods?

Originally published December 12, 2024. Rewritten .

Essay / 001
In this essay Contents

Early in my PhD I worried about how quickly ML safety research seemed to go stale. Something that looked like a core problem when I started could look very different three years later. A new architecture or training method would come along, an old limitation would stop mattering, and new subfields kept appearing. Now and then a problem actually got solved, but more often it had just been renamed. For a student this is a practical worry. If the papers around your topic might be unrecognizable in five years, how are you supposed to choose one?

The first time I tried to answer this, I wrote down nine areas of ML safety that I thought would last, including distribution shift, adversarial robustness, uncertainty estimation, reward hacking, interpretability, and scalable oversight. Looking back, I cut things too finely. A lot of those nine are local versions of problems that are much older than deep learning. People in statistical decision theory, formal verification, control theory, software testing, and systems safety were working on them decades ago. I’d now ask a different question. If we threw away every current algorithm and architecture tomorrow, which difficulties would we still have? I think there are four.

How do we specify what we actually want?

A system can do exactly what we told it to do and still not give us what we wanted. ML people call versions of this reward hacking, specification gaming, proxy optimization, or objective misspecification. Computer scientists were worrying about it long before any of those terms existed.

Hoare’s 1969 paper on axiomatic program correctness [1] is a good place to start. You reason about a program in terms of preconditions and postconditions. Writing {P}  C  {Q}\{P\}\;C\;\{Q\} means, roughly, that if program CC starts in a state where PP holds, it ends in a state where QQ holds, given the usual correctness conditions. It’s a powerful method, and it also makes a gap easy to see. You can prove a program meets its specification and still have a specification that misses something you care about. The proof says nothing about that.

The gap gets bigger once a system runs over time. Pnueli’s temporal logic (1977) [2] gave us a way to state properties of whole sequences of states. Alpern and Schneider (1985) [3] then split such properties cleanly into safety (nothing bad ever happens) and liveness (something good eventually happens). That split is useful because getting a task done and staying out of trouble are two separate demands. Picture a system that has to finish a job eventually (liveness) and must never leak private data while doing it (safety). You can make it finish more jobs, or finish them faster, and be no closer to guaranteeing the second property.

Machine learning makes this very visible, because we usually train against a number. Let a system’s behavior produce a trajectory τ\tau. We can compute and optimize a reward r(τ)r(\tau). What we actually care about is some value U(τ)U(\tau) that we can only partly observe or only partly specify. Nothing forces the trajectory that maximizes r(τ)r(\tau) to also maximize U(τ)U(\tau). Worse, the harder you optimize the proxy, the more room there is for the two to come apart in ways that matter.

Social scientists hit the same wall with measurement. Donald Campbell, writing about social indicators in 1979 [4], described how an indicator gets corrupted once institutions start making decisions based on it. In ML, Inverse Reward Design (Hadfield-Menell et al., 2017) [5] suggested a change of stance: read the reward a designer wrote as evidence about what they wanted, instead of as a perfect statement of it. Gao et al. (2022) [6] later measured reward-model overoptimization directly, showing that pushing hard on a learned reward eventually makes things worse when you check against a reference objective. The two papers work in very different settings, and both run into the same limit: if something important is missing from the specification, optimizing harder won’t put it back.

Things get harder still when the objective is uncertain, depends on context, or is something people actively disagree about, as opposed to merely being hard to write down. It may be that no single numerical utility function captures everything relevant. Reasonable people weigh tradeoffs differently. And some constraints are much more natural to write as “never do this” than as a penalty term. So the problem I expect to last is broader than reward hacking. It includes representing an objective we’re unsure about, telling hard constraints apart from preferences, catching mistakes in a specification, and building systems that still accept corrections after we’ve changed our minds about what we want.

I’d like papers in this area to say which of these they are working on. Training a better reward model, proving a property of a program, and finding the right set of safety constraints are related, and a paper that does one of them hasn’t done the others.

What can we infer about situations we have not observed?

Everything we know about a system was learned under some particular conditions, and the system will eventually run under different ones. Most of ML is built on generalizing from data. We gather examples, fit a model, and trust that what it learned will keep working. Unless we can say precisely what assumptions make that trust reasonable, it is just hope.

Say we build a predictor ff on data drawn from PP and then deploy it where the data come from QQ. The number we care about is the expected loss under QQ, and the loss we measured under PP is only a stand-in for it: RQ(f)=E(x,y)∼Q[ℓ(f(x),y)]R_Q(f) = \mathbb{E}_{(x,y)\sim Q}[\ell(f(x),y)](1) If nothing connects PP to QQ, good performance on one says nothing about the other.

Deep learning didn’t create this problem; statisticians have been working on it for a long time. Huber’s robust estimation work (1964) [7] studied estimators that keep working when a nominal distribution is contaminated in specified ways. Shimodaira (2000) [8] worked out predictive inference under covariate shift, including when reweighting the training data helps.

The main thing I took from that literature is that “robust” always means robust to some class of changes. If we write down a set Q\mathcal{Q} of distributions we think we might be deployed under, we can aim to bound the worst case: sup⁡Q∈QRQ(f)\sup_{Q\in\mathcal{Q}} R_Q(f)(2) Choosing Q\mathcal{Q} is where it gets difficult. A narrow class gives comforting guarantees that may miss the failures you actually run into, while an unrestricted one usually leaves you unable to guarantee anything useful about the worst case unless you add more assumptions.

This framing ties together several areas that are normally studied apart. Adversarial robustness asks how a model performs when someone deliberately picks the input or the environment. Out-of-distribution detection tries to notice when an input is different, in some relevant way, from what the model has seen. Uncertainty estimation tries to put a number on what the model doesn’t know about a given prediction. Calibration checks whether predicted probabilities match how often things actually happen. These are different practical questions, but each one is about what we’re entitled to believe beyond the data we have. They are still different problems. A strange-looking input may be one the model handles fine, and a completely normal-looking one may produce a bad error. A model can be well calibrated on its test population and badly off on one subgroup, or after conditions shift.

That’s why I think uncertainty should be connected to decisions. The idea that a classifier should sometimes refuse to answer goes back a long way. Chow’s 1957 paper on optimal character recognition [9] made rejection part of the decision problem itself, so the cost of a wrong answer could be weighed against the cost of not answering. I like this better than asking a model to report its uncertainty and leaving it there. What I actually want to know is whether answering, abstaining, gathering more information, or handing the case to someone else has the lowest expected cost.

Newer methods handle parts of this much more carefully. Ovadia et al. (2019) [10] tested how uncertainty estimates hold up under dataset shift. Conformal Risk Control (Angelopoulos et al., 2024) [11] provides ways to keep specified losses under control given certain statistical assumptions. They don’t remove the uncertainty, but they make it clear which guarantees you can get and under which conditions. The gap I find most interesting is between guarantees that hold on average over a population and how much you can trust one particular decision. An average-case guarantee doesn’t have to cover every subgroup, and it certainly doesn’t cover every situation a deployed system might find itself in.

The questions I expect to stick around are which assumptions about the environment are reasonable, which kinds of uncertainty we can estimate reliably, how to notice when our assumptions have stopped holding, and what a system should do once its guarantees no longer apply. Every system will run into things it hasn’t seen before, and I’d settle for one that still makes decisions it can defend when that happens.

How can we know whether a system is behaving correctly?

Say we’ve pinned down the behavior we want and made our peace with some assumptions about the environment. We still need evidence, good enough to convince a skeptic, that a complicated system will behave the way we intend. The obvious first step is testing: give it inputs, look at what comes out, and count how often it’s right. That only works if we can tell what a right answer looks like. Software engineering has a name for this, the test-oracle problem. Elaine Weyuker’s 1982 paper On Testing Non-Testable Programs [12] dealt with programs where the correct output is unknown, or too expensive to work out. Reading it now, it is very close to what we call scalable oversight.

Checking arithmetic is easy. Checking an unfamiliar proof, a complicated scientific hypothesis, or a long chain of decisions that depend on each other may not be. A system can give an answer that looks fine, where spotting the mistake takes as much expertise as producing the answer did, or more. Producing a result and checking it turn into separate skills. Weak-to-strong supervision is one way people study this today. Burns et al. (2024) [13] tested whether a weaker supervisor can draw out the abilities of a stronger model. They found some encouraging results and some clear limits.

Even when we can check outputs, rare failures are expensive to measure. A few hundred clean test runs tell you very little about a failure that shows up once in ten thousand. If trials are independent and identically distributed, you need about 3/ε3/\varepsilon failure-free tests before you can put an approximate 95% upper confidence bound of ε\varepsilon on the failure rate. Real deployments often break those assumptions. Failures can be correlated, can cluster in particular situations, or can be triggered by an adversary hunting for exactly the conditions your tests missed. Testing alone won’t get us there.

Interpretability tries to look inside the computation instead of judging only from inputs and outputs. Mechanistic interpretability in particular tries to explain how internal components cause the behavior we see. Circuit Tracing (2025) [14] shows how far this has come, and its authors are also candid about how incomplete and imperfectly faithful the resulting explanations still are. Knowing how part of the computation works is a long way from knowing how the whole system works. An explanation of one output also doesn’t tell you what the system will do on other inputs. Formal verification is a third option. Instead of sampling inputs and watching the system succeed, you prove it has some property under stated assumptions. That proof is only as good as your model of the system and its environment, though, so it can’t stand in for empirical checks when that model might be missing something.

Each of these covers something the others don’t. Tests tell you about behavior you’ve observed. Interpretability may tell you about internal mechanisms. Formal methods establish properties of a mathematical model, conditional on its assumptions. Monitoring tells you about behavior after deployment. I find it much harder to see how to assemble all of it into a safety claim you could defend. That’s roughly the idea of an assurance case or safety case, a structured argument that connects a specific safety claim to the evidence, assumptions, and reasoning behind it. Clymer et al. (2024) [15] and others have explored how this could work for advanced AI.

As research questions, I find these more interesting than building one more benchmark or one more interpretable representation. What evidence is enough for which claim? How do you combine several pieces of evidence that rest on some of the same assumptions, without counting those assumptions twice? And how do you evaluate a system that can tell it’s being evaluated, or can affect the evaluation? That last one worries me most. If a system acts differently when it believes it’s being tested, then its test results stop being a fair estimate of how it will behave once deployed. In the end this is an epistemic problem. We will only ever have partial evidence, and we need to know how far it should move our confidence in a complex system.

How can we keep a system safe while it interacts with the world?

The first three problems were about saying what we want, predicting what will happen, and collecting evidence. A deployed system also does things. It acts, gets feedback, changes the world around it, and sometimes changes itself. That’s the fourth problem. Safety has to hold for as long as the system and its environment keep influencing each other. Checking it once, at launch, doesn’t cover that. This is old territory for control theory and cybernetics. Wiener’s Cybernetics (1948) [16] was built around feedback: systems whose behavior depends on information coming back from the environment they are acting on.

The standard abstraction is a dynamical system with state sts_t, action ata_t, and outside disturbance wtw_t: st+1=F(st,at,wt)s_{t+1} = F(s_t,a_t,w_t)(3) A controller picks actions from what it can observe. Each action changes the next state, which changes what the controller sees and decides next. Safety is now a property of the whole trajectory. A typical requirement is that the state stay inside some safe set SS for the entire run, for any disturbance within a range we specify up front. Being right on average is a much weaker requirement. A controller can make one sensible decision after another and still end up somewhere unsafe, through small errors piling up or through feedback.

Classical control already has useful tools here. Page’s 1954 paper on continuous inspection [17] helped found the field of sequential change detection. Ramadge and Wonham (1987) [18] studied supervisory control of discrete-event systems, where a supervisor limits what the system may do so that only acceptable executions are possible. Supervisory control fits modern ML systems well, since they are built from parts we can’t trust to always decide correctly. You don’t have to make the learned component perfectly reliable. You can limit which actions it is allowed to take.

Alshiekh et al. (2018) [19], in Safe Reinforcement Learning via Shielding, did this for RL agents using temporal-logic specifications. Greenblatt et al.’s AI Control work (2024) [20] looked at protocols that prevent harmful outcomes even when a powerful, untrusted model is actively trying to get around them. The two settings have little in common, but both put the safety in constraints and oversight around an imperfect component, instead of relying only on fixing the component.

Interventions have failure modes of their own. A monitor might not have access to the signal that reveals danger. A fallback policy might be unsafe in some conditions. A human overseer might be too slow, or not know enough, to step in usefully. Even switching the system off isn’t automatically safe: in some physical systems, abruptly stopping control is dangerous in itself.

It gets harder again if the system keeps learning after deployment. Learning tasks one after another can make a model better at the new one and worse at the old ones, which McCloskey and Cohen documented as catastrophic interference in 1989 [21]. A deployed model can also shift the data it will see later through its own decisions. Perdomo et al. (2020) [22] formalized one version of this as performative prediction. The system then partly creates the distribution it will later have to adapt to. So I see continual learning, post-deployment monitoring, fail-safe design, and adaptive control as parts of one problem: keeping behavior acceptable while the system and its environment both change. We still don’t know how to build monitors we can trust, when to intervene, how to keep safety properties intact across updates, or whether safe behavior is achievable at all given realistic limits on what we can observe and control.

What does it mean for a problem to be timeless?

When I call these problems timeless, I don’t mean nobody will ever solve them. Many of them have clean mathematical answers once you fix the assumptions. We can verify some classes of programs, make predictions with statistical coverage guarantees, build estimators that tolerate limited contamination, and keep invariants for some dynamical systems. What’s hard is applying those results to systems whose environments, goals, and internals we only partly understand.

When I’m sizing up a proposed ML safety project, I run it through four questions. What do we actually want the system to do? Under what conditions do we expect it to do that? What evidence would justify believing it will? And what will keep its behavior acceptable when something changes? Projects that seem to have nothing in common often turn out, under those questions, to be blocked by the same thing.

It’s also why I think the older literature is worth reading. People in statistics, formal methods, software engineering, control theory, and systems safety ran into problems with the same structure long before anyone was training today’s models. Their answers won’t always carry over. Their definitions, counterexamples, impossibility results, and proof techniques can still save us a lot of time spent rediscovering old questions.

If I were advising someone just starting out, I’d tell them to spend less effort looking for a promising architecture or benchmark and more effort finding a limitation they can state without mentioning either. Then read carefully what earlier work proved about it, and work out which of its assumptions you’d have to loosen for the result to apply to a more realistic setting. The names we use for these problems will probably change in the next ten years. I’d rather work on the questions that will still make sense after they do.

References

1.
C. A. R. Hoare (1969). An axiomatic basis for computer programming. Communications of the ACM. Read source.
Back to text
2.
A. Pnueli (1977). The temporal logic of programs. 18th Annual Symposium on Foundations of Computer Science (sfcs 1977). Read source.
Back to text
3.
B. Alpern, F. B. Schneider (1985). Defining liveness. Information Processing Letters. Read source.
Back to text
4.
D. T. Campbell (1979). Assessing the impact of planned social change. Evaluation and Program Planning. Read source.
Back to text
5.
D. Hadfield-Menell et al. (2017). Inverse Reward Design. Read source.
Back to text
6.
L. Gao, J. Schulman, J. Hilton (2022). Scaling Laws for Reward Model Overoptimization. Read source.
Back to text
7.
P. J. Huber (1964). Robust Estimation of a Location Parameter. The Annals of Mathematical Statistics. Read source.
Back to text
8.
H. Shimodaira (2000). Improving predictive inference under covariate shift by weighting the log-likelihood function. Journal of Statistical Planning and Inference. Read source.
Back to text
9.
C. K. Chow (1957). An optimum character recognition system using decision functions. IRE Transactions on Electronic Computers. Read source.
Back to text
10.
Y. Ovadia et al. (2019). Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift. Advances in Neural Information Processing Systems. Read source.
Back to text
11.
A. Angelopoulos et al. (2024). Conformal Risk Control. International Conference on Learning Representations. Read source.
Back to text
12.
E. J. Weyuker (1982). On Testing Non-Testable Programs. The Computer Journal. Read source.
Back to text
13.
C. Burns et al. (2024). Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. International Conference on Machine Learning. Read source.
Back to text
14.
Back to text
15.
Back to text
16.
Back to text
17.
E. S. Page (1954). Continuous Inspection Schemes. Biometrika. Read source.
Back to text
18.
P. J. Ramadge, W. M. Wonham (1987). Supervisory Control of a Class of Discrete Event Processes. SIAM Journal on Control and Optimization. Read source.
Back to text
19.
M. Alshiekh et al. (2018). Safe Reinforcement Learning via Shielding. Proceedings of the AAAI Conference on Artificial Intelligence. Read source.
Back to text
20.
R. Greenblatt et al. (2024). AI Control: Improving Safety Despite Intentional Subversion. International Conference on Machine Learning. Read source.
Back to text
21.
M. McCloskey, N. J. Cohen (1989). Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem. Psychology of Learning and Motivation. Read source.
Back to text
22.
J. Perdomo et al. (2020). Performative Prediction. International Conference on Machine Learning. Read source.
Back to text

Cite this essay

Debargha Ganguly. “On Problems That Will Outlive Our Models.” Research notes, December 12, 2024; revised october 9, 2026. https://blog.debargha.com/posts/ml-safety-research-landscape/

@misc{ganguly2024problemsoutlive,
  author       = {Ganguly, Debargha},
  title        = {On Problems That Will Outlive Our Models},
  howpublished = {Research notes},
  year         = {2024},
  month        = dec,
  url          = {https://blog.debargha.com/posts/ml-safety-research-landscape/},
  note         = {Revised October 9, 2026},
}
Download .bib
∎

Thank you for reading.

Download bibliography