Blog

Nobody will know what the system is

For a bit over a year now coding agents have been superhuman at programming, and we're writing a lot more code than we used to. Our development processes haven't caught up though. It feels to me like the way we get changes into a codebase is up for debate right now, so I want to go through the options as I see them. I think of them as the four levels of self-driving codebases, like the levels of autonomy for self-driving cars.

  1. Humans read PRs and merge them as they see fit. Comfortable, and somebody actually knows what's happening to the canonical code.
  2. Agents review PRs and humans merge them. Probably how a lot of work gets done right now. Nobody really knows what's happening, but we pretend they do.
  3. Adversarial agents review PRs and they get merged automatically. Humans decide what should be built.
  4. Same as 3, except agents also decide what to build. The endgame, we just feed power in.

By adversarial I mean agents whose job is to find what's wrong with a change rather than say LGTM. What I mostly have in mind here is supply chain attacks in open source. If PRs to a popular library get merged without anyone reading them, the reviewers have to assume the person sending the PR might be trying to sneak something in. Think of the xz backdoor, except the attacker wouldn't need to spend two years earning a maintainer's trust first. Closed source has the same problem, it's just usually not malicious. Some junior doesn't pay enough attention in an important part of the system, breaks something, and the mistake gets past everything.

What I'm curious about is how we get to level 3 from wherever we are now, whether that's level 1 or 2. I think it depends a lot on what the system does.

For less critical systems it seems very doable already. I wouldn't lean on tests for it though. I don't trust AI-written tests at all, and good tests that actually stop things from breaking are rare and might need the whole codebase rewritten to even be possible. What has always been much stronger is version control. If main works, and the small change coming into main is something you can reason about, then by induction you can reasonably believe main still works. So keep the changes small enough for the adversarial reviewers to reason about, merge them automatically, and keep humans in the loop for releases.

Saying that nobody will know anymore what the system actually is sounds crazy. But if you think about it in terms of the path of least resistance, that's where this is going. Every time you skip reading something and nothing breaks, it gets easier to skip it next time. Not that long ago it sounded crazy that nobody would read code anymore, and now plenty of us don't.

I'm not sure it's as bad as it sounds though. Nobody understands all of Linux either, or everything in their node_modules, and we manage. What worries me more is losing track of what the system is supposed to do, as opposed to how it does it. That feels like the part humans should hold on to, and the moment we don't need to anymore, I'd say that's AGI.

In that world I'm interested in what programming language AI will use. It no longer needs to be optimized for humans, which doesn't necessarily mean it'll be unreadable to us, since it was trained on human text. My first guess was that it would be about efficiency of expression and care less about things like good variable names.

I suspect names don't help models as much as we think, especially for local variables. Function names and parameters absolutely matter, but only in a weakly typed language. In a strongly typed one the types speak for themselves. That's just a hypothesis though, and it should be quite easy to settle with some experiments. And if agents are checking each other's work, what matters most might be how easy the code is to verify.

From my experience models are really good at Rust because of the strong type system, where the compiler gives them a lot of feedback. It works especially well when the model is forced to encode the system into the type system, with Results, error enums and a more functional style, instead of Options everywhere, default types, and strings thrown around JavaScript or Python style like it's a try/catch language. Then a lot of mistakes just don't compile. That doesn't mean AI will end up writing Rust though. I think it's going to be something else, or an existing language written in a very specific way that humans won't like, for example local variables named x0, x1 and so on.