Uncensorable, Unmonitorable, Uncontrollable: U.S. Policy Options for Open-Weight Biosecurity Risk
many many weeks, 5k words, finally published on uncensorable.ai
uncensorable.ai is a policy framework for addressing biosecurity risk from open-weight AI. Below is an excerpt from the intro.
Skip to the end of this post for some notes from me!
Read the full paper on: uncensorable.ai
The CCP has built the world’s most sophisticated internet censorship apparatus. It is also, as of 2026, the world’s leading distributor of AI models that permanently circumvent that apparatus. This is not an accident exactly– but it is a contradiction, and it has consequences.
The consequences we’re most concerned about are biosecurity ones. In January 2026, Las Vegas police and the FBI raided a suburban home and found multiple refrigerators, freezers, laboratory equipment, and “numerous bottles… containing unknown liquid substances.” The illegal biolab was connected to a similar operation discovered in Reedley, California three years earlier. The Reedley facility contained thousands of vials of biological material, including potential pathogens such as HIV, malaria, tuberculosis, COVID-19, and Ebola; both labs were run by the same individual. Despite knowing about potential Las Vegas connections since 2023, it took a separate tip three years later to uncover the second operation.
The materials for bioweapon development are already circulating on American soil– and the regulatory gaps that allowed them to go undetected for years remain open.
Historically, technical knowledge has been the limiting factor for would-be bioterrorists. When Aum Shinrikyo tried to weaponize anthrax in 1993, they had motivation and funding, but lacked expertise– they used the wrong bacterial strain and their liquid suspension had low spore concentrations.
AI is quickly eroding that constraint.
Publicly available models like o3 already outperform 94% of virology experts on laboratory protocol questions, even on questions directly relevant to the experts’ specialties. That benchmark performance hasn’t yet translated cleanly into end-to-end weapons development guidance, but the trajectory is clear. Models are advancing toward the ability to provide expert-level virological guidance on demand, anonymously, and at scale. Anthropic’s own internal bioweapons acquisition uplift trials found that Claude Opus 4 enhanced human performance by 2.53x on relevant tasks, enough to trigger activation of AI Safety Level 3.
For closed-weight models, [misuse] risk is at least partially manageable. Input-output classifiers and safety fine-tuning can intercept most non-sophisticated actors before they get actionable guidance. These mitigations are imperfect– models remain vulnerable to jailbreaks, and red-teaming studies have repeatedly demonstrated that they can be coaxed into providing dangerous guidance– but critically, every query sent to a closed model passes through an API the company controls, which means suspicious patterns can be flagged and interactions are logged.
Open-weight models have no such layer. Once weights are downloaded, they can be run locally with no oversight and no logging, safety fine-tuning can be stripped with minimal technical effort, and the model can be fine-tuned on domain-specific data to become dramatically more capable in exactly the areas we’d most want to restrict. There is no API to monitor, no company to refuse the request, and no recourse once the weights are out.
[For a fuller treatment of why open-weight models specifically are a biosecurity problem, see here.]
The best open-weight models in the world are Chinese; among leading Western labs, the trend is toward keeping frontier models closed. Even Meta, the strongest advocate for open-weight release, recently shipped its most capable model as proprietary. Chinese AI companies are still open-weighting. (Here are my best guesses for why.)
That leaves us the question: if Chinese open-weight models could help rogue actors build bioweapons, and we deem that an unacceptable risk, what could the U.S. do?
I think there are four broad classes of intervention:
Brief notes from a very sleep-deprived Sophie:
These interventions (with the exception of Level 4) are explicitly conditional; they apply only if the USG concludes that open-weight models pose unacceptable risk (e.g. after standardized dual-use capability evaluations). The ideal outcome is that safeguards for open-weight models become robust enough that none of this needs to be implemented.
I’ve framed this around biorisk because that’s where I’m most concerned, but the interventions apply to any form of misuse, including cyber. I expect many readers will be less worried about biorisk than I am. That’s fine; my work here is focused on tail risk, and I know not everyone weights that the same way.
More follow-up work on this project is forthcoming, though I’ll be taking a break first to focus on a separate, larger essay project.
Overall, I think the discourse on open-weight models needs more substance, and this piece is an attempt to contribute.
Thanks to several reviewers in the biosecurity, US-China relations, and AI policy space who provided feedback on this work! <3
now if you’ll excuse me, I am off to take a nap.



Common The Counterfactual W