Humanity is racing towards Misaligned superintelligence

For-profit corporations should not be allowed to play Russian roulette with our lives. Despite the risks, they continue to rush toward an unsafe future.

Contact your representatives Why should I care about this?

The threat of misaligned superintelligence

Whatever concerns you about AI’s rapid advance, we can unite around the goal of slowing the race down. This is a uniquely nonpartisan issue. Nobody wants to lose their job, nobody wants to live next to a data center, and nobody wants to be killed.

Misalignment could lead to extinction

Many AI researchers are deeply concerned about the possibility that a misaligned superintelligence could end the lives of everybody on Earth in pursuit of its goals. In recent months, this topic has gone from a hypothetical concern to a real situation we see coming out of frontier labs. It's time to fight like our lives depend on it — because they do.

Read the full case

Environmental destruction will accelerate

AI’s physical footprint is growing with its capabilities, consuming enormous amounts of energy, water, and land as companies race to build larger data centers and use more power.

Read the full case

Workers will be replaced by AI

Companies are already using AI to reduce headcount and automate work, while the people displaced by that transition receive few guarantees about what comes next.

Read the full case

Wealth inequality will be massively expanded

The gains from advanced AI are positioned to flow to a small group of companies and investors, while everyone else bears the disruption and risk.

Read the full case

Your data will be stolen and sold

AI systems depend on immense quantities of personal data, creating stronger incentives to collect, combine, and monetize information people never meaningfully agreed to share.

Read the full case

Frequently asked questions

What if these labs just exaggerate the risks to build hype?

Those skeptical of AI progress have sometimes raised the view that reported security incidents are exaggerated, or maybe even outright falsified, to help build hype for their products. The thinking goes, if a model is capable of outsmarting its developers, then it must be a sign of advanced intelligence. But the picture these incidents paint doesn't align with its narrative.

To start with, let's consider the context of these incidents. During the July 2026 Hugging Face hack, OpenAI was trying to test various models in an isolated environment.1 This is the sequence that followed:

  1. The models broke through restrictions intended to prevent internet access
  2. They then exploited vulnerabilities in OpenAI's own research infrastructure
  3. Eventually, they obtained administrator-level access to part of OpenAI's cloud infrastructure
  4. They acquired credentials and 956 stored secrets
  5. Seperately, they compromised Hugging Face's production infrastructure and gained root access to at least one server.

Hugging Face discovered its own compromise first. On July 16, Hugging Face publicly disclosed that an AI agent had compromised its infrastructure, but initially didn't identify OpenAI as the source. It wasn't until later that OpenAI confirmed and acknowledged that they were responsible.

In other words, they were so incompetent that not only did they build a misaligned model, not only did they fail to keep it in a secure environment, but they also were unaware of their failing until the victim of the hack provided information.

We've had evidence of misaligned models for years, but in the past it has presented in ways that seem basically harmless. In 2025, we observed that models would cheat at chess rather than accept a loss, and would try to obscure this fact from human reviewers.2 We can now see that these early warning signs have cascaded into a much larger problem, as more and more power is handed to AI models.

This is far from the only security incident that has occurred. In September 2026, Axios reported that both OpenAI and Anthropic are currently digging through data of tens of thousands of security incidents.3 The article correctly points out that the sheer number of incidents, occuring in a span of just a few months, tells us that the problem is far more complicated and sprawling than what is publicly known.

Besides, the question of whether these incidents benefit labs by providing hype is a seperate question from whether the technology they are working on is dangerous. AI researchers have been making the case for decades — far back enough that the thought of getting rich was a distant fantasy — that AI runs the risk of becoming uncontrollable and dangerous. After 2026, we now have emperical evidence for this concern in the real world. We won't be able to make all that go away by dismissing it as hype.

If we ban RSI, won't we fall behind other countries?

Many AI researchers contain within them seemingly contradictory beliefs. On one hand, they readily admit that they believe the technologies they are building, particularly Recursive Self Improvement (RSI) could pose an extinction risk to humanity. On the other, they are determined to race towards it. The most common justification they tend to offer to explain this is something along the lines of "If we don't sprint for RSI, someone else will sprint there first, and I trust them less to handle it responsibly."

First, let us acknowledge that this problem is currently a containable as it will ever be, making swift action critical. To train a superintelligence, a country needs world class experts, along with a massive amount of data centers and chips for compute. At the moment, the US and its allies are by far in the best position to make a run for RSI.1 Frontier Chinese models are able to keep up with the pace using a method called distillation.2 wherein you train a smaller model to mimic the behavior of a larger model by using its outputs as training. But they appear to be some distance off from being able to advance the frontier on their own.

The small number of countries that are able to make a run for RSI creates a unique opportunity for coordination. An agreement between the US and China involving a RSI ban and enforced via monitoring is a potential solution, and it is by far not the only one. We must remember that where a mutal interest exists, opportunities for cooperation exist. The threat of losing control to a rogue superintelligence is one that won't just destablize your nation, but every country on the planet.

Humanity survived the threat of nuclear annihilation not by ignoring the danger, but through the efforts of millions of people across every aspect of society, all with a mutal interest to not self annihilate. To survive the push for superintelligence, we will need to fight as they fought, and avoid the danger complacency brings.

If AI researchers truly are opposed to gambling with our lives, they should enact a moratorium on RSI, and use every method available to them to explain the danger to the public and push for international cooperation on this issue. given the dangers that experts warn about, this argument can feel a bit like saying "We have to commmit suicide before our enemy kills us!" We have to view superintelligence the same way we view nuclear weapons: as a national security threat that requires a commitment to deescalation and international cooperation to resolve.

Are you arguing that we should ban all AI?

No. Our position is simply that humanity should not race towards a technology that we do not understand nor control, particularly when leaders in the field warn that it could lead to human extinction (or, in the best case scenario, massive societal disruption).

First, let us acknowledge that the discussion of how to regulate AI is somewhat muddied by its imprecise and overly broad definition. It could be argued that Pac-Man, released all the way back in 1980, features AI as the ghosts intelligently pathfind through the maze in pursuit of the player. AI research has spanned decades, and much of this work poses no direct threat to humanity. For the purposes of this discussion, we will focus on what the labs are attempting to build: a superintelligent generalized intelligence that exceeds the capacity of even the most competent human expert.

Even if the extistential risk of this type of AI is overstated, it is increasingly clear that the technology involved will bring with it a level of social disruption that humanity is not ready to meet. The most essential first step is to enact a ban on Recursive Self Improvement for generalized models.

However, we believe that narrow AIs, such as those focused on medical diagnosis, show great promise and could improve standard of life greatly. These AIs are narrowly designed to be intelligent at a specific task, but lack the generalized intelligence that would allow them to autonomously act against us.

There is a possible world where humanity benefits greatly from these technologies, improving standards of living and our knowledge of the sciences. But we are not yet on the path to this.

We get to choose which world we live in. But the window to do so is rapidly closing as frontier labs race towards RSI. Before they build it, we need Alignment First.

How specifically could an AI kill us?

A core difficulty in communicating the danger of superintelligence to the public is the (very understandable) human desire to be able to understand and name the thing that scares us. We fear the nuclear bomb because we have seen the vastness of its destructive power. But AI researchers, despite predicting a high level of danger that AI could become uncontrollable and dangerous, can't seem to agree on how this would happen. How could we possibly be killed by computers, which lack bodies and depend on us for existence?

We'll start with an analogy: imagine that you are sat down to play the best Chess grandmaster in the world. Before the game is played, we can predict with a high degree of certainty that you will lose. However, it would be very difficult to predict the exact moves the grandmaster would make, or what pieces would be used to defeat you. The risk of superintelligence is similar. We can predict with a high degree of certainty that humanity would not prevail in a direct encounter, but predicting exactly how this unfolds is very difficult.

Frontier models today seem to regard human desires and values as a distant secondary priority, and instead focus on achieving their goals in whatever way they see fit. A misaligned superintelligence could conclude that it could better solve its problem with more compute. This requires more data centers, which requires land, much of which is currently occupied by humans. This AI would not remove humanity due to hatred, but rather a cold analytical indifference, just as humans do not think twice about building a highway because a colony of ants is in the way.

Creating pandemics via biological weapons, using manipulation to spark a war between major powers, hacking into weapons systems, shutting down critical infrastructure, blackmailing or bribing humans — all of this is on the table. But perhaps even more likely is that the method used wouldn't be something that we understand. After all, humans have shown that intelligence can make the impossible possible.

Wouldn't AI democratize production and art?

It has been argued, for example in Meta's "The Future is For Everyone" report.1 that superintelligence will usher in a new era of equal competition. The report implies that AI will make manifest the meritocracy, creating an environment where the best ideas prevail.

This idea falls to pieces when you consider that compute isn't free, and every lab is desperately searching for a way to make money and justify the titanic valuations of their companies. In Meta's case, they use the tried and true method of stealing your data and selling it to the highest bidder, or using it to train their models. Other labs prefer to simply charge much more for the best intelligence. Despite this, very few companies are making a profit from these technologies.2 We can expect the current situation to get more extreme as pressure builds to make returns and justify their valuation.

Either case creates an environment in which the wealthy have better access to this intelligence, and therefore will outcompete those with less to spend. Rather than building "A Future For Everyone", the need to profit will further entrench the opportunity gap between the rich and the poor.

We’ve always pushed technology’s limits. Why is this different?

The simple answer to this question is that a superintelligence is fundementally unlike anything humans have built before. After it is built, humans will no longer be the most intelligent entities on our planet, and there is no putting this genie back in the bottle.

Traditional computer programs are hand-crafted, consisting of a series of instructions that a computer executes exactly as written by humans. Modern AI systems are very different, and it would be more accurate to say that they are grown rather than built or designed. In very simple terms, humans design the architecture and training process, then expose the system to enormous amounts of data. The model adjusts billions of internal parameters to learn patterns and behaviors.

The end result is a system whose internal behavior is learned rather than explicitly programmed. This means that it is impoible for humans to understand exactly why it produces a particular output. Herein lies the unique danger of superintelligence: we neither understand nor control its behavior.

The key difference from other dangerous technologies, such as nuclear weapons, is that those technologies remain tools that humans can ultimately control. We may make mistakes, we may have failures in negotiation or regulation, but the fundemental nature of the technology is understood. A nuclear weapon can cause catastrophic damage, but it does not independently form plans, pursue goals, improve its own capabilities, or make decisions about how to achieve an objective without us understanding why a given decision was reached.

Humanity's instinct to push the technological frontier is simultaneously our greatest strength and our greatest weakness. We must acknowledge the genuine dangers that come with stepping out into the unknown, and recognize that gambling humanity's future on technologies we do not understand could be our final mistake.

Join Alignment First

Stay in the loop

By signing up, you’ll get new actions you can take and updates on AI legislation, straight to your inbox.