Policy Memo: Why the U.S. Should Not Pursue MAIM
A Rational Case Against AI Sabotage Frameworks
Note: This is a simplified version of a more detailed analysis on positional incentives; you can read the original here.
Policy Memo:
The United States should not pursue a strategy of sabotage against states that could develop advanced AI because U.S. government-sponsored MAIMing normalizes the practice as acceptable and invites other countries to MAIM in retaliation– both of which make a MAIM equilibrium more likely to emerge.
The MAIM framework is fundamentally based upon the idea that as rival countries get progressively closer and closer to developing superintelligence and/or advanced AI that confers decisive strategic advantage, other countries should MAIM to prevent this reality from materializing– MAIM is a self-preservation mechanism that ostensibly represents what scenario would emerge if all states involved can be considered “rational actors.” This is the descriptive part of the argument by Hendrycks, Schmidt, and Wang– that MAIM is likely to emerge by default.
As the current leader in the ASI race, the U.S. would be uniquely disadvantaged by an AI deterrence regime like MAIM—becoming the primary target for sabotage while gaining comparatively little from the ability to sabotage others. Subscribing to MAIM essentially means inviting sabotage on the country’s own AI development, which is why the country is incentivized to do everything in its power to ensure MAIM fails– whether by preventing the doctrine from developing in the first place or actively undermining a MAIM equilibrium if one were to occur.
Rational U.S. Strategy:
Strongly deter sabotage from other countries by making it clear that any sabotage of U.S. AI projects will trigger escalatory counterstrikes.
Deterrence doesn’t work if one country’s “deterrent” is interpreted as another country’s “first strike” – if another country thinks a MAIM could trigger a kinetic counterstrike, for example, a MAIM becomes much less likely to occur.
China doesn’t want the U.S. to develop a DSA, but are they willing to risk war to MAIM based on the (uncertain) belief the U.S. may develop one soon?
A counterargument to this is that cyberattacks might still work given their low salience and the difficulty of credibly committing to, say, a kinetic counterstrike in response to a cyber incident, but I find this counterargument problematic for a few reasons:
It seems to contradict Hendrycks and Khoja’s own argument in their follow-up article that “norms will adapt to strategic realities,” which is to say that policymakers will “grasp the alarming implications of ASI” in the future– if it’s true that policymakers will understand the implications of ASI, then the U.S. could likely credibly deter MAIMing, even if it’s done through cyberattacks, by establishing that MAIM incidents will be treated like acts of war given the stakes of ASI.
One might argue that cyberattacks are difficult to counter because they can be difficult to attribute, or sometimes even difficult to detect, but the authors of MAIM themselves recognize that there is a trade-off here: they write in the normative section of Superintelligence Strategy that “[w]hen subtlety proves too constraining, competitors may escalate to overt cyberattacks… in a way that directly—if visibly—disrupts development” – not if, but when.
It still doesn’t change the fact that the U.S. should employ this strategy to do their best to deter MAIMing– also, the issue of cyberattacks will be discussed further in #2:
Maintain maximum opacity on AI progress– without high confidence in the inevitability of a breakthrough that confers DSA, other countries likely won’t MAIM, especially if the risks of executing a MAIM strike are raised by #1.
This also includes increasing cybersecurity, especially around sensitive projects like the Genesis Mission.
Rather than preserving mutual vulnerability, the U.S. should harden datacenters, distribute training runs across geographically dispersed facilities, and otherwise make it extremely difficult to sabotage the country’s AI infrastructure.
Initiating a MAIM attack would normalize AI sabotage as an acceptable practice, which would make it impossible for the U.S. to credibly deter MAIMing from other countries and unreasonable to escalate in the event of another country’s MAIM.
Making MAIM equilibrium more likely to emerge is not in the country’s interest.
Key Uncertainties & Important Considerations:
In large part, my analysis is focused on what the U.S. should rationally do given that MAIM confers a disproportionate disadvantage on the leader of the ASI race. If the US’s status changes, and China becomes the frontrunner, the basis of this entire memo is flipped on its head. I would strongly consider whether the US’s stance should change based on the risk of another country becoming leader– if China has a >x% chance of overtaking the US, then making MAIM the default might be worth accepting sabotage because it ensures China can’t employ any of the strategies detailed, if MAIM is already normalized by the time they take the lead.
Key questions to research include:
Is China good at making chips? How long until their domestic supply chains catch up with the US’s?
How many months behind are they currently?
What do their espionage capabilities look like?
What events might enable the country to gain a lead?
Given that the U.S. is undermining its own talent pool with immigration restrictions, what effect could this have on the pace of AI development relative to China?
A major consideration that I didn’t include in the memo itself is the issue of superintelligence itself and the danger developing it presents. MAIM is uniquely disadvantageous for the US, but if you assume that the U.S. cannot be trusted to develop aligned superintelligence – which is a reasonable assumption – then it is in the country’s best interest to undermine its own development, buying time for safety researchers to develop better technological methods and governance frameworks to catch up to the breakneck pace of AI so that the U.S. doesn’t commit collective suicide on behalf of the entire world by creating ASI it cannot control.
However, in a world where the U.S. is sufficiently superintelligence-pilled to believe this, the rational choice would be to instead commit to a pause on developing ASI altogether, which I unfortunately don’t see as likely for several reasons– including lack of trusted verification mechanisms, difficulty establishing concrete red lines for superintelligence development, and the need for international coordination.
Consider that we don’t even have an agreed-upon definition of ASI! Hendrycks and the CAIS team wrote one for AGI, but I don’t think a reputable ASI counterpart exists (or at least, I haven’t yet seen or heard of it).
MAIM itself assumes development continues, but with sabotage.
For the cyberattacks section, it’s possible that even if the U.S. commits to escalatory retaliation, political realities will make a kinetic counterstrike for a cyberattack impossible— public opposition to such escalation could be too great even if the national security establishment is willing to act on their word because they recognize how much of an existential threat another country developing ASI first is.
However, I think this strengthens the argument that MAIM is a bad deal for America, because it means the U.S. is largely defenseless against a MAIM attack (cannot credibly threaten escalatory retaliation); I wrote about this more in detail on the longer post I linked.
Important Note on Author Perspective
Quite honestly, I support the idea that given the current state of alignment research, “if anyone builds it, everyone dies.” I wrote this considering what actions would make sense from a U.S. perspective given game theory, competitive dynamics, and national security priorities; I do not, in fact, think that these strategically rational actions are normatively desirable.
This memo analyzes what U.S. policymakers are incentivized to do under current conditions, not what would maximize human survival. Descriptive analysis of incentives ≠ normative endorsement of outcomes.
The piece attempts to describe the world as it is, not as it should be.

