far.in.net


~Time to write

Pour tomber du ciel, il faut y être monté, ne fût-ce qu’un instant…
—Théophile Gautier, 1874, Histoire du romantisme, p. 153

I’m writing from the liminal space between London and San Francisco. Time works differently here (roughly speaking, three hours become ten). There’s also no internet, and that means no agents or students to prompt. Seems like the perfect opportunity to prime myself for my visit to Berkeley. Let’s see what I can make of it.

§United Kingdom

I’m thinking back to 2020. It was a whole year of liminal time. Civilisation was going somewhere, then it suddenly stopped. When would we get back to what we were doing before? Would we ever?

I don’t claim I had the worst of it, but I did go from an intellectually transformative experience in wintry Zürich to living on European time out of a caravan in the Yarra Valley.

Begrudgingly, civilisation and I carried on. In my case, I slept by day and studied the foundations of learning and computation by night. I discussed papers on AI risk and careers in AI safety with students from EA Zürich, the blue glow on my face betraying my timezone to my webcam.

I also attended a virtual EA student summit. This was later in the year, so I’d finished my moonlight semester, but the schedule was mostly optimised for European and American students, so I had to pull another all-nighter to attend the talks.

I don’t remember much from the summit now. But I do remember a piece of advice from fellow Australian Buck Shlegeris on how students should orient towards AI safety:

We really need people who’ve tried to think through the whole thing.

Soon after Buck’s talk, the caravan power failed, and I was left in the dark. When would the light come back? Would it ever? I went to sleep.

§Iceland

Buck’s challenge struck a chord. A deep one, that is still quietly resonating with me to this day.

I’ve spent a lot of the last six years working really hard. I studied the foundations of learning and computation. I learned how to be a researcher, and found my way to the frontier of the science of deep learning. I watched closely as language models went from sleep-talking about unicorns to solving open problems in mathematics.

In six years, the one thing I haven’t done is spend a lot of time “thinking through the whole thing.” Trying to resolve big-picture questions about the nature of the problem we are up against. Updating my models of the problem in light of new developments in the field and in technological/governance progress. Thinking about how the work I am doing eventually connects to making a positive impact on the chance that the future goes well.

I know thinking clearly about these issues is extremely important. I know I need to spend more time on it. The problem is, there has always been an object-level project to work on instead.

When I was a junior research student, I didn’t have a lot of control over the projects I had the opportunity to work on, so I didn’t have a lot of opportunities to set aside time to dig all the way to the foundational questions and spend time there.

I thought that when I was more senior, I would have the freedom to buck some of the publication pressure and I would naturally be able to get back to basics. I was right about the freedom—as a doctoral student at Oxford, I have complete control over what I spend most of my time on. I was wrong about time to think coming naturally.

These days, I have so many appealing project ideas and collaboration opportunities that I struggle to even keep track of them all, let alone do them justice, and the natural state is for me to ricochet between projects and periods of mild burnout.

Besides, thinking is hard. I’m not always in the mood for something so challenging, compared to the excitement of running parallel Claude sessions working on experiments or software projects I’ve had on my TODO list for years. Not to mention teaching, the strongest of my addictions.

Over the last few months, I have ghosted colleagues on two projects of this kind. Recently, one of my colleagues called me out for this:

Spit it out!

It tilted me a little, but maybe it was just what I needed to remind me of what is important.

§Greenland

I haven’t made zero progress thinking through the whole thing. It comes in bursts. Shallow flashes of intuition when I read something dissatisfying or ideas popping into my head when I walk through Oxford (without headphones).

Sometimes, I make non-trivial progress in a deeper burst of effort. I can point to two public examples: when I replied to Zuckerberg’s particularly maddening takes on AI safety and when I wrote a lecture on superintelligence for UniMelb’s Ethics of AI class. There have been a number of other instances where I made some progress, but didn’t take it all the way to publishing something.

So inspiration to think carefully does sometimes come naturally. The problem is that I could probably count these moments on two hands. It’s not enough progress—there is too much left to think through. I can’t keep waiting for the stars to align.

The only way I know how to deliberately spend time thinking is by writing. Writing is this magical activity where I can take thoughts from inside my head, where they seem to all hang together into a coherent structure, to before my eyes, where the misshapen and cracked structure of the thoughts becomes visible. I then work on the structure, replacing the missing links and rebalancing the different elements, often changing the conclusion dramatically by the end.

Given my style of thinking, you can see why I ended up a computer scientist:

Maybe by rigorously vetting my beliefs, I can become less wrong, more quickly. Maybe a new conceptual distinction, or a re-framing of the problem, leads to a new path forward. Ultimately, the more progress I make, the more I might be able to contribute to a solution, either by sharing my insights or executing on their implications for action myself.

§Canada

Is this the best way for me to contribute to the goals of the field of AI safety? This is a tough question that I don’t yet have an answer for. But the we’re only half way through liminal time. Let me break it down into three component questions, and see if I can make progress on any of them.

Question 1: If we want to make intellectual progress as a field, is it useful to have more people attempting to contribute to making intellectual progress?

The obvious way to make intellectual progress is to spend more person-hours thinking about it. OK, there are lots of failure modes to watch out for here. You want the hours to be spent thinking carefully, engaging productively rather than destructively, and so on. And you could get a lot of progress by sheer luck or genius. But to increase your exposure to luck and geniuses, the more the merrier.

In particular, the way to sustain intellectual progress over a long timescale seems to be to have each new generation critique its civilisation’s core truths to the best of its abilities. In the course of doing so, the new generation either comes to appreciate these truths more than they would if the truths had been dogmatically imposed, or it reveals genuine problems with the ‘truths’ and replace them with better ones.

In a rationalist society, intellectual leaders should justify their beliefs and practices, and be responsive to valid critiques of their arguments. In a classically liberal society, this invitation should extend to everyone who is willing to participate in this process in good faith. This is the political philosophy underpinning modern Western civilisation, if not in implementation, at least in aspiration. The same dynamic also plays out in the sciences.

It seems to me to be an open question how well this approach applies to the issue of AI. I don’t have the political philosophy background to know the details, but presumably the arguments in favour of this approach have their own assumptions.

For example, it seems you have to assume that the intellectual ‘progress’ that comes out of this process naturally moves civilisation in the direction of truth and goodness, rather than degenerating. This may or may not be true in general, depending on how reliable the participants are at evaluating such developments, how many people are involved, and how fast they try to make progress.

The issue of AI is rapidly developing, and we might not have the time required to properly implement the fully liberal approach. This doesn’t assume particularly short timelines to crucial AI developments (at least not “particularly short” by today’s standards): even given a tranquil decade, there’s a lot of ground to cover. Maybe we can’t afford to broaden the conversation and let it run its course by the time key decisions are forced. Likewise, if we try to rush the process, maybe we lose the reliability properties of discourse the political philosophers took for granted.

The above is just to play devil’s advocate. On priors, the liberal approach is more robust than strategies that are more dogmatic and authoritarian, since, speaking from dogma and without authority, dogma and authority are bad. (It’s consistent with the liberal view to entertain dogma and authority, at least, since we need to see for ourselves if the liberal approach holds up.) In the end, I just come out pretty uncertain on this question, since I haven’t thought a lot about political philosophy.

Question 2: Does the field need to make intellectual progress at all?

Buck Shlegeris called for this in his talk in 2020, but the field has matured somewhat since then. And in any case, I should see if I can decide for myself.

Thinking takes time, which imposes an opportunity cost. Instead, marginal time could be devoted to one of a number of existing directions, for example:

  1. refining specific safety techniques from existing proposals;
  2. building more capable AI systems (this makes sense, to a point, if you think more capable AI systems are more helpful for safety research automation, or if you think they’ll just be safe anyway); or
  3. pursuing global coordination, such as via regulation, such that people don’t recklessly build substantially more capable AI systems.

You might not need to have a very clever strategy for allocating field resources between these directions (and others)—it may be the case that just letting people choose between them based on personal fit/inclination and serendipitous opportunities leads to an approximately optimal allocation.

Moreover, there’s only so much you can do with a fixed amount of data before you need to collect more. So it’s possible that spending more time thinking about the problem is not the right move for the field right now, and instead the field should work as hard as possible towards gathering more data by developing more intelligent systems or empirically investigating risk models (also, towards pursuing other existing directions such as those enumerated above).

In these worlds, the field isn’t blocked on intellectual progress. By contrast, we want more intellectual progress if we live in one or more of the following kinds of worlds:

It’s unclear to me exactly which kind of world we are in, which leaves me pretty uncertain about this question too. But, in particular, I intuitively think there’s a decent chance we’re in one of the worlds where there is fruitful progress still to be made on the big-picture questions.

Question 3: Am I, personally, well-placed to contribute to intellectual progress?

It’s hard to say. It’s a challenging task. For all my academic background, field knowledge, and aptitude for writing, I might not have what it takes to contribute to conversations about civilisation and technology at the highest levels.

One piece of evidence that suggests I might have something to contribute is that I often feel intuitively sceptical of the kinds of approaches and arguments that the field currently pursues. (Unfortunately, I can’t give examples, since I haven’t thought through them, but doing so is a near-term goal.)

Maybe I need to lean into this intuitive dissatisfaction. Maybe if I follow it where it leads, I’ll find something that wouldn’t otherwise have been found. I’ll write about it, share it, and the field will move a little closer to the true and the good.

On the other hand, maybe my intuitions are systematically wrong, and I’m kidding myself that I’ll find something important that the many people who are already working at the intellectual frontiers of AI and AI safety have missed. Many of these people have presumably come a lot closer to “thinking through the whole thing” than me, so it’s too early to trust my vague intuitions.

But that’s the thing. The only way to find out is to think things through. Maybe my thinking will amount to nothing, and in hindsight, it will not have been worth the effort compared to spending the extra time working on research, teaching, advocacy—whatever else. Well, that doesn’t seem so bad. If this were the gamble, I’d feel somewhat compelled to try my luck.

§United States

Is that the time? We’re landing in a couple of hours.

I’ve come out pretty uncertain about whether the field needs more people to think through the whole thing, and whether I personally would have anything positive to contribute to this effort were I to try. However, part of that uncertainty includes the possibility that this is exactly what I should be doing, since it could help the field and help me do better research.

The above is an inside view account. Stepping out of the field for a moment, there’s a big outside view reason to think that we’ll need a lot of intellectual progress. Advances in technology are inviting profound and novel challenges to our civilisation’s core truths. In the liberal view, now is the time of greatest need for those who can help re-examine and refine the foundations of our civilisation, that they may hold against the disruptions to come.

I think this licenses me to spend some time trying to think through the whole thing, and see how far I get. That brings me to the purpose of my visit to Berkeley. Nominally (per my travel grant), the high-level goal of my trip is to spend two weeks refining my strategic model of AI existential risk and technical research, by writing essays and discussing these ideas with colleagues. As part of the trip, I’ll also attend the ILIAD conference at Lighthaven, where I hope to talk with even more people about these kinds of topics.

This summer really seems like a good time to prioritise writing. It is an exciting time for the field and for myself.

For the field, an exciting recent development is the launch of Geoffrey Irving’s new org, Resolution. Since this new org is subsuming Timaeus as its new Learning Theory division, I think I’m now technically a Resolution Research Affiliate.

I have extremely high hopes for this new organisation. It’s shaping up to become a large and well-resourced institution dedicated to tackling the hard parts of the issue of AI. It could be comparable in ambition to the scaling labs, but with a federated structure and relatively isolated from lab-like incentive structures so as to preserve research integrity and intellectual diversity. This seems like a solid foundation for making real progress on the real problem. I want Resolution to achieve its greatest potential, and I am excited that by thinking carefully and talking with the founders, I might be able to contribute to that.

As for myself, I’ve reached the halfway point of my doctoral studies. Looking back, I have made some intellectual progress, but not as much as I was hoping. I want to kick-start that side of my degree with an intense period of writing. I hope to attack several of the pieces I’ve had kicking around on my TODO list for months/years, waiting for the right moment. As discussed, it’s something I need to be deliberate about and set aside time for. There’s nothing like a change of scenery to shake up your routine.

In framing this trip as a writing retreat, I’m taking inspiration from the Inkhaven initiative. I might not be able to match that pace (500+ word post per day), partly because I am going to insist on writing-as-thinking, which can often enough lead me to dead ends and unpublishable results. But I do intend to make intellectual progress my main focus during the trip, and I’m excited to see what I can do in two weeks if I put my mind to it. I’m publishing this post as a pre-commitment to help ensure I put the effort in while I’m here, and don’t cave to new distractions.

What exactly will I write? I don’t know yet—I haven’t thought about it. Maybe I’ll start by writing-as-thinking about a list of topics on my mind. For now, the sun is still high in the sky, but it feels like the middle of the night. I’m going to try to use what remains of my liminal time to sleep.