7 Comments
User's avatar
Sophie Kim's avatar

I think critique without solutions/alternatives is often fairly unproductive, so I want to be clear about my intent here: I’m sharing this piece to stress-test my theory on MAIM's instability before I move on to the implications of this dynamic. I’m opening this up for commentary specifically to see where my logic might be flawed, so I can pivot toward more practical, stable policy implications in future piece(s).

Thanks for reading! :)

Oscar Delaney's avatar

This analysis seems right to me. I wonder if the original authors were mainly trying to convince China to MAIM the US, given fears around the US failing to solve alignment. That is, the main reason to advocate for MAIM seems to be to buy more time to work on AI safety, which we may not get without strong pressure from China threatening to or actually MAIMing. But I agree that if one doesn't buy misaligned AI takeover worries then MAIM seems very bad from the AI leader's perspective.

Sophie Kim's avatar

> I wonder if the original authors were mainly trying to convince China to MAIM the US, given fears around the US failing to solve alignment.

I hadn't considered that angle re: original author intent, but that seems plausible given it was a CAIS project!

> But I agree that if one doesn't buy misaligned AI takeover worries then MAIM seems very bad from the AI leader's perspective.

I think even given fears re: loss of control, MAIM still doesn't appear to be a good deal for the U.S.-- would be curious to hear your thoughts re the following:

---

If the U.S. is sufficiently x-risk-pilled, the logical choice would be to either:

(1) Pause development, or

(2) Focus on making American ASI safer

In the former case: MAIM becomes unnecessary

In the latter case: The U.S. will still not want to accept a regime of mutual sabotage

Essentially, I think something like: in any world where the U.S. continues development, it stands to reason the country wouldn't want to accept sabotage.

---

Thanks for your comment! I'm a huge fan of your Substack by the way!! :)

David Krueger's avatar

> if China doesn’t know that the U.S. is nearing superintelligence, the chance of a MAIMing strike is virtually zero, and the U.S. can continue development without high risk of sabotage.

This seems wrong. I think this would be correct if China were confident that the US was not nearing superintelligence, or did not understand the strategic implications. But if China is highly uncertain, then it seems like "precautionary" MAIMing might be justified.

Sophie Kim's avatar

Hi David! This is a fair point, thank you for bringing up-- I address this more thoroughly in my follow-up post on hardening:

> This uncertainty deters action even when stakes are high, because the costs of being wrong (triggering escalatory retaliation) remain constant while confidence in the necessity of action drops.

I have a few more thoughts, but it's a lot to put in a comment, so I'll be writing a separate post on precautionary MAIMing later this week. Thank you for the feedback!

*edit to add: am currently sick so precautionary MAIM research has been postponed lol

David Krueger's avatar

Makes sense but I don’t think this gets to near-0 probability. It’s pretty clear that there’s a good chance US companies are rapidly approaching superintelligence.

Sophie Kim's avatar

This is fair! I do think I may have overstated the "near-0 probability" bit in my original post-- it would be more accurate to say it lowers the probability considerably, but perhaps not to such a low extent.

I explored the idea of "MAIMing under uncertainty" a bit more in this follow-up: https://thecounterfactual.substack.com/p/extended-discourse-on-maim-part-2

More thinking re: retaliation dynamics and how they affect the likelihood of MAIMing will also be posted shortly :)

Thank you for your feedback!