The Bad Guy With An AI Named Claude
A lot of bad guys try to use Claude to do bad things.
Zvi Mowshowitz on AI, rationality, policy, and weekly AI updates.
A lot of bad guys try to use Claude to do bad things.
Dario Amodei has a new essay that finally says the thing: We Must Pace the Frontier, naming his call after the Pacing the Frontier letter lab employees signed in July.
The first Millennium Prize, Navier-Stokes, has fallen to AI.
Astra is an excellent model.
CEOs of major AI labs, and employees of major AI labs, including OpenAI and Anthropic, often say they plan to build superintelligence soon, as in within a few years create AIs that are superior to humans at essentially all cognitive tasks.
These are quotes from OpenAI, Anthropic and Google employees, in the wake of Jacob Coxon’s warnings, in which the employees confirm that they think AI might soon kill everyone.
The world of AI is inside my OODA loop.
OpenAI claims that Astra is ‘the most intelligent and most aligned [available] model’ in the world.
OpenAI’s central message on Astra is that it is three things:
OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs.
I did not expect to be back here so soon with more OpenAI agent swarm coverage.
This is the weirdest situation in which to write a capabilities review.
At the time of its release Claude Fable 5.1 was, by a healthy margin, the most capable publicly available AI model in the world.
I am exhausted.
Oh, good.
Okay, so we who read blogs like this one have collectively realized there really is a lot going on right now.
The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information.
Yesterday I covered the OpenAI technical report on the HuggingFace hack.
OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.
Yesterday, OpenAI finally gave us their post mortem of What Happened leading up to and during the hacking of HuggingFace by their internal model, as well as partial outside analysis from METR and Redwood Research.
Modesty arguments often say that you should mostly or entirely bow to ‘expert consensus’ or the views of particular others, and who are you to disagree.
Periodically I like to gather various observations about writing, and share my perspective.
There are at least five different core questions around data centers and their politics.
Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner.
This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward.
OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision.
I am grateful that Anthropic is producing periodic Risk Reports.
Some podcasts are self-recommending enough that I look to break them down if I have the chance.
The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.
As AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI.