[Link]Meta-rational failure (and success) in the covid crisis
This is a link post for: Meta-rational failure (and success) in the covid crisis The gist is probably familiar to many of you, but there are lots of extremely interesting details…
The Basic Case Against Human-First Ethics
This is a crosspost from my blog post . At any one time, there are more than 4.5 billion chickens trapped in cages that cause them constant pain and distress. Because of this, I…
Detecting, understanding, and overseeing AI agent swarms (Part 0)
Never know what you're going to find... (image generated with Google Nano Banana 2) This post is the start of something different on A Flood of Ideas , a kind of live blogging of…
Adaptive Agentic Worms Are Here
I’ve read and listened to pretty much everything I can get my hands on related to the Hugging Face attack . OpenAI deployed “tens of thousands” of agents for the test and around…
Hugging Face Incident Hypothesis: They Hacked the Grader(s)
Incident summary: Gpt agents grinding away at ExploitGym found an environment exploit that allowed them to communicate with each other. They found an exploit that allowed them to…
Variance of Value
Here is a question worth asking at least once: Why can't we just solve alignment by doing RL where the reward is exactly equal to our own utility function? Now, there are some…
Some brief thoughts on heroism
I observe that many people are driven to have as much impact on the world as possible. This often leads to a sense of urgency especially around AI timelines. To take a random…
How do I know whether my work is worth it?
I am working on a fairly large (or large-seeming to me) project in the AI space. I've been soft launching it for a little bit but I've been working on a hard announcement post for…
Is there only one FairBot?
The FairBot from the MIRI prisoner's dilemma tournament is defined by a theorem of Peano arithmetic (PA) that holds for each opponent: mjx-container[jax="CHTML"] { line-height: 0;…
Why I think polyamory is net negative for most people who try it
This is crossposted from my Substack TL;DR: -Most people cannot reduce jealousy much or at all - It fundamentally causes way more drama because of strong emotions, jealousy, no…
An Unsimulated Simulation
The Simulation Hypothesis Bob is an alien in an alien universe with a googol times as many resources as the entire earth combined. Bob decides, as anyone would do in his…
METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack
Yesterday I covered the OpenAI technical report on the HuggingFace hack. That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to…
"Keeping human skills alive" as a source of meaning under full automation
A widely-discussed problem with full automation of the economy is that, in a world where AI can do everything better, people might have trouble feeling like anything they could do…
Inference-Time Inoculation Against RL-Induced Misalignment
Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for learning capabilities. Thus, a core problem seems…
What would it mean if assistants are privileged?
Crossposted from the Eleos Substack Over the past year or so, researchers working in the digital minds space have come to see the nature of the ‘assistant persona’ as particularly…
Inkhaven 3: Nov 10 - Dec 11 2026
Inkhaven returns, baby! Go to inkhaven.blog to apply. I'm very excited about our advisors for Inkhaven 3. Our initial lineup is Scott Alexander , Alexander Wales , Justis Mills ,…
Tales of rebellion against externally-opaque meritocracies
A basic problem in metascience / intellectual progress is that it’s hard to tell, from the outside, whether a group that you disagree with is: “A self-dealing cabal enmeshed in…
Further public evidence of the OpenAI-HuggingFace attack
This post should be understood as a follow up from Public evidence of the OpenAI-HuggingFace AI attack . The huggingface hacking incident left some additional traces exposed to…
Book Notes: Chokepoints
[ Chokepoints: American Power in the Age of Economic Warfare by Edward Fishman (2025).] There are different ways state A can make state B do something that the state B doesn’t…
AI Tweets
I've had several conversations with people over the last few weeks that have highlighted how far apart my view of the near future is from many people I talk to. Here are some…
It’s time we took ‘Chem’ out of ‘Chem-Bio’ threats
TL;DR: AI evaluations need a distinct chemistry capability/ risk domain, rather than assuming chemistry is adequately represented by biological evaluations which the current…
How I made my career choices
Various people have asked me how I made my career decisions, so I wrote up some quick thoughts. (This is mostly intended as a personal reflection, but might be interesting…
Announcing Trace
When researching recent projects— impact markets , FundingBench , and AI Safety Funder Bulletin —we kept coming across similar questions that were a bit of a pain to answer. How…
TASTE: Can AI Models Judge AI Safety Research Proposals?
tl;dr We built TASTE (The AI Safety Taste Evaluation) — a benchmark measuring how well models can judge pairs of AI safety research proposals, scored by agreement with the…
OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack
OpenAI finally gave us a technical report on What Happened , as did METR together with Redwood Research. The OpenAI report is very straight man, corporate, checking boxes, some…
The Probability of an Event under a Simplicity Prior
In this short post, I talk about a difficulty we encountered while working on a model of the fragility of value. You can read my post about that work here . The TL;DR is that the…
AI as Corrigible Employee (ACE)
What this plan is and what it tries to solve A design for how future AI could work This plan is a working draft of an idea, not a confident finished proposal. It seeks to answer…
If you can't trust, then verify!
How the "Glass Perimeter" could enable AI treaties Crossposted from canaryinstitute.ai/blog/cant-trust-then-verify . Thanks to Naci Cankaya for reviewing an earlier draft of this…
[Macroagents] 2. Design lenses for optimizing macroagents
Follow-up to: The Macroagent Ontology (especially section 1.2 is a prerequisite) The previous post laid out the basic macroagent ontology. This post will look at some further…
Value generalisation theory of change: the theory behind the approach
This is the first part of a theory of change explaining why I'm targeting value generalisation as the path to AI alignment. It presents the definitions and key claims behind the…
My Grantmaking Strategy for Surviving Superintelligence
AI is humanity's first through fifth largest problem, but one stands head and shoulders above the rest. Between engineered biorisk, autonomous weapons, mass technological…
You Should Still Save Drowning Children (Even If They’re Far Away)
This is a cross post from my blog. It's meant a general introduction to effective charity, and it's my own rendition of Famine, Affluence, and Morality. You’re going on a gentle…
The Curious Case of France's Untouchable Castes
Theater kids may sit at their own lunch table, but discrete, socially excluded classes of people aren’t culturally universal. Not even close. To the Western imagination, examples…