Bricks and exponentials: A note on how I evaluate projects

This is an essay that I wrote to a colleague at Palisade, articulating why I feel unsatisfied with goals and projects that others on the team (on average) feel more enthusiastic about. It describes one aspect of how I, personally, am doing strategic analysis and choosing which projects to invest in.

Related: Compounding Resource X

Bricks for a wall

Say you need 70 million bricks to build a wall. You also need architects and builders, and 50,000 tonnes of mortar (all which you also don’t have right now), but you’ll eventually need 70 million bricks,). You can maybe get away with using only 50 million bricks, if you rely on clever architectural tricks, but less than that is not going to cut it.

You ran a labor-intensive 6 month project to make 100,000 bricks.

Now, one of three things could happen:

  1. You, or someone, is using the bricks that you made to build kilns. You are contributing to a self-reinforcing industrial process that is producing an order of magnitude more bricks each year. [You’re upstream of an exponential]
  2. Someone starts a rapidly-growing brick-making school, which will churn out another 1000 brick maker teams. [You’re “downstream” 1of an exponential]
  3. Neither of the above.

In case 1, your brick-making project was leveraged and (potentially) hugely important. There were several critical constraints and you removed one of them. Probably that was an amazingly good use of 6 months of labor. This project is at least in the running for the best thing you ever did. Hell yeah. 

That impact-story critically depends on the bricks you made being an input to a compounding process. Your win wasn’t making bricks that would be used for the wall. Your win was setting up a process that would lead to there being more than enough bricks for the wall.

In case 2, you did a noble thing, bearing some of the weight of the total task of making bricks. But also, the part you did wasn’t very neglected or counterfactual.

It’s very improbable if the brick-making school is able to scale up brick-making by almost 3 orders of magnitude, but runs out of steam just short of the critical 70 million mark. Either it got a good process going, and way overshot the 70 million target and bricks are cheap, or it couldn’t get a good process going, and they didn’t get anywhere near the necessary 70 million target. It’s very unlikely that we get by, by the skin of our teeth, with barely enough bricks.

In either of these subcases, I would still celebrate someone who contributed the 100,000 bricks that they could produce. But strategically, in case 2, I would say that they misallocated their 6 months. However, basically the only thing that mattered for whether enough bricks were gotten was whether the brick-making school worked, or not.

In case 3, your 100,000 bricks were not nearly enough to build the wall. No wall, or only a very inadequate wall, is built. While a noble attempt, it didn’t amount to anything important. History would have gone the same way, regardless of whether or not you had been born or bothered to try.

Little shifts in worldview for changing society

Palisade’s first youtube video got 850,000 (so far) views. Probably a lot of those people said “huh” and changed their worldviews a bit, and then went back to what they were doing. 

People who have had their worldview changed bit are like the raw material for possible changes in the world. Those people now have the prerequisites for another argument step. And if 70 million US citizens knew that the AIs are grown and not crafted, and that we can’t always control them, that would have a notable impact on the landscape of voter sentiment about various AI proposals (presuming that others do the work of designing and introducing and advocating for those proposals).

So how should I evaluate the impact of this video?

If this video contributes to building a machine that compounds, so that exponentially more people have their view changed, over the next three years, that video was amazing— maybe the best thing Palisade has ever done. We’ve changed (or helped to change) history.

If someone sets up some process (maybe a media-fellowship) that churns out orders of magnitude more content like that Palisade video, then we should be proud of doing our part and helping with the effort of informing the world. But if that process can 1000x the amount of content, that will have made such content cheap and no longer neglected and not a bottleneck. And on the other hand, if that process is only 10x’s the amount of similar content, we’re still in a range that is too small to matter.

Either way, we didn’t change history. We did do our part in a wave of change that would have succeeded or failed without us (which is still noble and good and admirable).

If nothing like either of those happens, then 100,000 people who say “huh”, change their worldview a little, and go back to their lives, is couch-change. Basically irrelevant. We get some dignity for trying at all, but we do not change history, and we do not win.

The actual situation

I don’t claim to understand the dynamics of the memetics around topics like this. I’m totally open to stories whereby there is a self-reinforcing or compounding process that results from videos like Palisade’s. I

I could totally believe, for instance, that one creator with an existing large audience watched that video, and is now inspired to make a video on a similar topic, which will be seen by even more people, and that this will be part of a compounding wave that eventually leads to the New York times writing about RSI  the  way they write about the Iraq war—not as an interesting speculative theory, but as mundanely real thing that is actually happening. I don’t know that’s how it works, but I wouldn’t be shocked to find out that that’s how it works..2

I’m not trying to make an object level claim about whether that video is or isn’t upstream of a compounding exponential process.

But the standard that I’m using to evaluate that video is “does it look like this will contribute to an exponential process that can get all the way to the region of relevance”, not “did we succeed in impacting 100,000 people, who can be fodder for later steps down the line.”

I don’t want to discourage people from producing the 100,000 helpful units that they can produce. I salute everyone who’s contributing what they can see to contribute to the effort securing our future.

But, for myself, at least I want to have my eye on the ball of whether 100,000 helpful units is several orders of magnitude too few, and searching for plans that can bridge gaps of that size. Which means that I generally believe in proposals that are upstream of an exponential (or plausibly, but not definitely, upstream of an exponential), and not plans that aren’t.

  1. Technically you’re lateral to the exponential, unless the brick-making school also trained you. But this is analogous to being downstream of the exponential, insofar you are producing value in the same units as the exponential process is. ↩︎
  2. Palisade doesn’t have to singlehandedly cause the exponential process. If we contribute to, and get some reasonable shapely value of an exponential that might or might not succeed without our help, that’s still a massive win worth the commitment of my life.

    But if the exponential process in question looks overdetermined—it will succeed with high reliability with or without our help, then I consider that part of the problem under control and turn my attention to the parts that are more critically on fire. ↩︎

Investing in wayfinding, over speed

A vibe of acceleration

A lot of the vibe of early CFAR (say 2013 to 2015) was that of pushing our limits to become better, stronger, faster. How to get more done in a day, how to become superhumanly effective.

We were trying to save the world, and we were in a race against Unfriendly AI. If CFAR made some of the people in this small community that focused on the important problems 10% more effective and more productive, then we would be that much closer to winning. [ 1 ]

(This isn’t actually what CFAR was doing if you blur your eyes and look at the effects, instead of following the vibe or specific people’s narratives. What CFAR was actually doing was mostly community building and culture propagation. But this is what the vibe was.)

There was sort of a background assumption that augmenting the EA team, or the MIRI team, increasing their magnitude, was good and important and worthwhile.

A notable example that sticks out in my mind: I had a meeting with Val, in which I said that I wanted to test his Turbocharging Training methodology, because if it worked “we should teach it to all the EAs.” (My exact words, I think.)

This vibe wasn’t unique to CFAR. A lot of it came from LessWrong. And early EA as a whole had a lot of this.

I think that partly this was tied up with a relative optimism that was pervasive in that time period. There was a sense that the stakes were dire, but we were going to meet it with grim determination. And there was a kind of energy in the air, if not an endorsed belief, that we would become strong enough, we would solve the problems, and eventually we would win, leading into transhuman utopia.

Like, people talked about x-risk, and how we might all die, but the emotional narrative-feel of the social milieu was more optimistic: that we would rise to the occasion, and things would be awesome forever.

That shifted in 2016, with AlphaZero and some other stuff, when a MIRI leadership’s timelines shortened considerably. There was a bit of “timelines fever”, and a sense of pessimism that has been growing since. [ 2 ]

My reservations

I still have a lot of that vibe myself. I’m very interested in getting Stronger, and faster, and more effective. I certainly have an excitement about interventions to increase magnitude.

But, personally, I’m also much more wary of the appeal of that kind of thing and much less inclined to invest in magnitude-increasing interventions.

That sort of orientation makes sense for the narrative of running a race: “we need to get to Friendly AI before Unfriendly AI arrives.” But given the world, it seems to me that that sort of narrative frame is mostly a bad fit for the actual shape of the problem.

Our situation is that…

1) No one knows what to do, really. There are some research avenues that individual people find promising, but there’s no solution-machine that’s clearly working: no approach that has a complete map of the problem to be solved.

2) There’s much less of a clean and clear distinction between “team FAI” and “team AGI”. It’s less the case that “the world saving team” is distinct from the forces driving us towards doom.

A large fraction of the people motivated by concerns of existential safety work for the leading AGI labs, sometimes directly on capabilities, sometimes on approaches that are ambiguously safety or capabilities, depending on who you ask.

And some of the people who seemed most centrally in the “alignment progress” cluster, the people whom I would have been most unreservedly enthusiastic to boost, have produced results that seem to have been counterfactual to major hype-inducing capability advances. I don’t currently know that to be true, or (conditioning on it being true) know that it was net-harmful. But it definitely undercuts my unreserved enthusiasm for providing support for Paul. (My best guess is that it is still net-positive, and I still plan to seize opertunities I see to help him, if they arise, but less confidently than I would have 2 years ago.)

Going faster and finding ways to go faster is an exploit move. It makes sense when there are some systems (“solution machines“) that are working well, that are making progress, and we want them to work better, to make more progress. But there’s nothing like that currently making systematic progress on .

We’re in an exploration phase, not an execution phase. The thing that the world needs is people who are stepping back and making sense of things, trying to understand the problem well enough to generate ideas that have any hope of working. [ 3 ] Helping the existing systems, heading in the direction that they’re heading, to go faster…is less obviously helpful.

The world has much much more traction on developing AGI than it does on developing FAI. There’s something like a machine that can just turn the crank on making progress towards AGI. There’s no equivalent machine that can take in resources and make progress on safety.

Because of that, it seems plausible that interventions that make people faster, that increase their magnitude instead refining their direction, disproportionately benefit capabilities.

I’m not sure that that’s true. It could be that capabilities progress marches to the drumbeat of hardware progress, and everyone including the outright capabilities researchers moving faster relative to growth in compute is a net gain. It effectively gives humanity more OODA loops on the problems. Maybe increasing everyone’s productivity is good.

I’m not confident in either direction. I’m ambivalent about the sign of those sorts of interventions. And that uncertainly is enough reason for me to think that investing tools to increase people’s magnitude is not a good bet.

Reorienting

Does this mean that I’m giving up on personal growth or helping people around me become better? Emphatically not.

But it does change what kinds of interventions I’m focusing on.

I’m conscious of deferentially promoting the kinds of tech and the cultural memes that seem like they provide us more capacity for orienting, more spaciousness, more wisdom, more carefulness of thought. Methods that help us refine our direction, instead of increase our magnitude.

A heuristic that I use for assessing practices and techniques that I’m considering investing in or spreading: “Would I feel good if this was adopted wholesale by DeepMind or OpenAI?”

Sometimes the answer is “yes”. DeepMind employees having better emotional processing skills, or having a habit of building lines of retreat, seems positive for the world. That would give the individuals and the culture more capacity to reflect, to notice subtle notes of discord, to have flexibility instead from a the tunnel vision of defensiveness or fear.

These days, I’m aiming to develop and promote tools, practices, and memes, that seem good by that heuristic.

I’m more interested in finding ways to give people space to think, than I am in helping them be more productive. Space to think seems more robustly beneficial.

To others

I’m writing this up in large part because it seems like many younger EAs are still acting in accordance with the operational assumption that “making EAs faster and more effective is obviously good.” Indeed, it seems so straightforward, that they don’t seriously question it. “EA is good, so EAs being more effective is good.”

If, you, dear reader, are one of them, you might want to consider these questions over the coming weeks, and ask how you could distinguish between the world where your efforts are helping and the world where they’re making things worse.

I used to think that way. But I don’t anymore. It seems like “effectiveness” in the way that people typically mean it is of ambiguous sign, and actually what we’re bottleneck on is wayfinding.


[ 1 ] – As a number of people noted at the time, the early CFAR workshop was non-trivially a productivity skills program. Certainly epistemology, calibration, and getting maps to reflect the territory were core to the techniques, and ethos. But also a lot of the content was geared towards being more effective, not being blocked, setting habits, and getting stuff done, and only indirectly about figuring out what’s true. (notable examples: TAPs, CoZE as exposure therapy, Aversion Factoring, Propagating Urges, GTD) To a large extent, CFAR was about making participants go faster and hit harder. And there was a sense of enthusiasm

[ 2 ] – The high point of optimism was probably early 2015, when Elon Musk donated 10 million to the future of life institute (“to the community” as Anna put it, at my CFAR workshop of that year). At that point I think people expected him to join the fight.

And then Elon founded OpenAI instead.

I think that this was the emotional turning point for some of the core leaders of the AI-risk cause, and that shift in emotional tenor leaked out into community culture.

[ 3 ] – To be clear, I’m not necessarily recommending stepping back from engagement with the world. Getting orientation usually depends on close, active, contact with the territory. But it does mean that our goal should be less to affect the world, and more to just improve our own understanding enough that we can take action that reliably produces good results.