Anthropic Founding, 2021: The OpenAI Schism
On this page7 sections
Anthropic Founding, 2021: The OpenAI Schism
- Date:
- May 2021
- Location:
- San Francisco, California
- Lab/Organisation:
- Anthropic (founded by Dario and Daniela Amodei)
- Significance:
- Most significant schism in AI-industry history; created safety-focused competitor to OpenAI
An AI safety company founded in 2021 by Dario and Daniela Amodei and other former OpenAI researchers. Anthropic develops the Claude language-model family and pioneered Constitutional AI, an alternative to RLHF that aligns models to a written constitution rather than human raters.
Why They Left
The reasons for the Amodeis’ departure from OpenAI are not fully public, but the broad outlines are clear from reporting and from the Amodeis’ own statements.
Dario Amodei had joined OpenAI in 2016, coming from Google. He rose quickly, becoming Vice President of Research and playing a key role in the development of GPT-2 and GPT-3. The disagreements that led to the departure were about several issues. According to The New York Times, the Amodeis left “after a series of disagreements with OpenAI executives over how its A.I. should be funded, built and deployed.” The specific points of disagreement appear to have included:
Safety. The Amodeis believed that OpenAI was not taking safety seriously enough — that the rush to build more powerful models was outpacing safety research. They wanted to build AI more slowly, more carefully, and with more attention to the risks.
Commercialisation. The Amodeis were concerned about OpenAI’s growing commercial relationship with Microsoft. Microsoft’s $1 billion 2019 investment had given Microsoft significant influence over OpenAI’s direction, and the Amodeis believed this influence was pushing OpenAI toward commercial priorities at the expense of safety.
Governance. The Amodeis had concerns about OpenAI’s capped-profit model — the structure that gave the nonprofit board control over the for-profit subsidiary. They believed the structure was not working as intended, and that commercial pressures were overriding the nonprofit mission. (This concern would be dramatically validated two years later, in the November 2023 board crisis — see B39.)
Dario Amodei has said that safety was “not sufficient to leave” — suggesting the departure was about a broader pattern of concerns, not a single issue. The Amodeis were not leaving because of a specific disagreement about a specific decision. They were leaving because they had concluded that OpenAI, as an institution, was not going to prioritise safety in the way they believed was necessary. The only way to build AI the way they wanted to build it was to start their own company.
What They Built
Anthropic was founded in May 2021 and incorporated as a Delaware Public Benefit Corporation — a legal structure that requires the company to consider public benefit, not just shareholder returns. The founding team included Sam McCandlish, Tom Brown, Jared Kaplan, Chris Olah, Ben Mann, and Jack Clark (formerly OpenAI’s Director of Policy) — several of whom had been co-authors on the GPT-2 and GPT-3 papers at OpenAI. The team was, by any measure, one of the most experienced groups of AI researchers ever assembled in a single startup. Tom Brown had led the GPT-3 project. Jared Kaplan had co-authored the Kaplan scaling laws paper (see B73). Chris Olah had become one of the most respected researchers in the field of interpretability — the study of what neural networks actually do, internally, when they process information.
The funding came quickly. In May 2021, Anthropic raised $124 million in a Series A round, with investors including Jaan Tallinn (a co-founder of Skype and a major donor to effective altruism causes). In 2022, the company received approximately $500 million from Alameda Research — the trading firm associated with Sam Bankman-Fried. This investment would become a legal complication after the collapse of FTX in November 2022, with the funds held in frozen accounts to be returned to the FTX bankruptcy estate.
In late 2022 (disclosed February 2023), Google invested approximately $300 million, taking roughly a 10% stake. On September 25, 2023, Amazon announced a commitment of up to $4 billion (with an initial $1.25 billion and an option to increase). By March 2025, Anthropic raised approximately $4 billion in a Series E at a post-money valuation of approximately $61.5 billion, led by Lightspeed — making it one of the best-funded AI startups in the world.
The funding trajectory was itself a signal. Anthropic had positioned itself as the safety-focused alternative to OpenAI, and the major cloud companies — Google, Amazon — had chosen to invest in Anthropic as a hedge against Microsoft’s exclusive partnership with OpenAI. The safety positioning was, it turned out, also a commercial positioning. By 2025, Anthropic was one of the most valuable private companies in the world, and its safety-focused brand was a significant part of its competitive advantage.
The Early Technical Work
Before Claude, Anthropic spent nearly two years — from May 2021 to March 2023 — doing research that was not, in the conventional sense, product development. The company did not release a chatbot, did not launch a consumer product, and did not compete with OpenAI in the public eye. Instead, it focused on two areas of research that it believed were prerequisites for building safe AI: interpretability and alignment.
The interpretability work, led by Chris Olah, produced some of the most influential research in the field. Olah and his collaborators developed the technique of “circuits” analysis — a method for reverse-engineering the internal representations of neural networks, by identifying the individual features that individual neurons detect and the ways those features combine to produce the network’s output. The circuits work, published in a series of papers on the Distill journal and on Anthropic’s research blog, was not, in itself, a product. But it was, by most accounts, some of the most important research being done on understanding how neural networks actually work — and understanding was, in Anthropic’s view, a prerequisite for safety. You could not make a system safe if you did not know how it worked.
The alignment work, led by Sam McCandlish and others, focused on the problem of how to train AI systems to pursue the goals their designers intended, rather than misspecified proxies. This work built on the RLHF technique (see B37 — Paul Christiano) that had been developed at OpenAI, but it extended it in several directions. Anthropic researchers explored alternatives to RLHF, including methods that used AI feedback instead of human feedback, and methods that trained models to evaluate their own outputs against a set of principles. The latter approach would eventually become Constitutional AI, Anthropic’s most distinctive technical contribution.
The early period also saw Anthropic invest heavily in scaling laws research. Jared Kaplan, who had co-authored the original Kaplan scaling laws paper at OpenAI, continued this work at Anthropic. The company’s research on the relationship between model size, data size, and performance helped to inform its decisions about how large to build its models and how to train them efficiently. This research was not, in itself, a product — but it laid the groundwork for the Claude models that would follow.
The FTX Complication
Anthropic’s early funding included approximately $500 million from Alameda Research, the trading firm associated with Sam Bankman-Fried. Bankman-Fried, the founder of the cryptocurrency exchange FTX, had become one of the largest funders of effective altruism causes, and his investment in Anthropic was, in part, an expression of his commitment to AI safety as an effective altruism priority.
The investment became a major complication in November 2022, when FTX collapsed and Bankman-Fried was arrested and charged with fraud. The collapse of FTX froze Alameda’s assets, including its stake in Anthropic. The funds that Anthropic had received from Alameda were held in accounts that became subject to the FTX bankruptcy proceedings, and the FTX estate — led by bankruptcy trustee John Ray III — announced its intention to sell Anthropic shares to recover funds for FTX’s creditors.
The FTX complication was, for Anthropic, both a financial distraction and a reputational problem. The company had accepted money from a source that turned out to be fraudulent, and the association with Bankman-Fried — who was convicted on seven counts of fraud in November 2023 and sentenced to 25 years in prison — was not one that a safety-focused company wanted. Anthropic did not return the money (it had already been spent on research and operations), but the FTX stake was eventually sold off by the bankruptcy estate. In March 2024, FTX sold two-thirds of its Anthropic stake to a group of investors led by G Squared for approximately $884 million. The remaining stake was sold later in 2024.
The FTX episode was, in some ways, a test of Anthropic’s safety-focused brand. The company had taken money from a source that turned out to be unsafe, and the question was whether this would damage its credibility. In the event, it did not — at least not permanently. Anthropic’s subsequent funding rounds, from Google and Amazon, were large enough to make the FTX investment a relatively small part of its financial history, and the company’s safety-focused research output was substantial enough to maintain its credibility. But the episode was a reminder that the financing of AI safety was, itself, not always safe.
Constitutional AI
Anthropic’s most distinctive technical contribution is Constitutional AI — a method for training language models to be safe and helpful, without relying solely on human feedback. The method, described in a December 2022 paper (arXiv:2212.08073, “Constitutional AI: Harmlessness from AI Feedback”), works by giving the model a “constitution” — a set of principles that the model is asked to follow. The model is then trained to evaluate its own outputs against the constitution, and to revise them if they violate the principles.
The constitution draws on sources like the United Nations Universal Declaration of Human Rights and the terms of service of major technology companies. The model is prompted to ask itself: “Is this response harmful? Is it deceptive? Is it helpful?” If the answer is no, the model revises the response. This process, repeated many times, trains the model to produce outputs that are consistent with the constitution.
Constitutional AI was a significant innovation. It reduced the reliance on human raters — who are expensive, slow, and potentially exposed to harmful content during the rating process. It also made the safety training more transparent: the principles were explicit, and anyone could read them and understand what the model was being trained to do. This transparency was a deliberate contrast with OpenAI’s approach, which relied on RLHF (Reinforcement Learning from Human Feedback — see B37 — Paul Christiano) — a process that was less transparent and that depended on human raters whose judgments were not always consistent.
Constitutional AI was not a replacement for RLHF — Anthropic’s models still used RLHF as part of their training. But it was a complement, and it addressed some of RLHF’s known limitations. The technique has been influential, and variants of it have been adopted by other AI labs. It also reinforced Anthropic’s safety-focused brand: the company was not just building capable models, it was developing new techniques for making models safe, and it was publishing those techniques for the broader community to use.
The Claude Models
Anthropic’s first public model, Claude, was released in March 2023. Claude was competitive with OpenAI’s GPT models on most benchmarks, and it was widely praised for its safety features — its tendency to refuse harmful requests, its transparency about its limitations, and its conversational style. The Claude 2, Claude 3, and Claude 3.5 releases followed, each improving on the previous version. By 2025, Claude was widely regarded as the most capable alternative to OpenAI’s models, and Anthropic had established itself as OpenAI’s most serious competitor.
The Claude model family evolved rapidly. Claude 1 (March 2023) was Anthropic’s first public release, trained using Constitutional AI and demonstrating the company’s safety-focused approach. Claude 2 (July 2023) improved on Claude 1’s capabilities and introduced a longer context window, allowing the model to process longer documents. Claude 3 (March 2024) was released in three sizes — Haiku, Sonnet, and Opus — and the Opus version was, by some measures, the most capable model available at the time, surpassing GPT-4 on several benchmarks. Claude 3.5 Sonnet (June 2024) and Claude 3.5 Opus (later in 2024) continued the rapid improvement, and by 2025, Claude was the model of choice for many developers, particularly for tasks that required careful reasoning or long-context understanding.
The flow of safety researchers from OpenAI to Anthropic continued through 2023 and 2024. In May 2024, after OpenAI disbanded its Superalignment team (which had been co-led by Jan Leike and Ilya Sutskever), Jan Leike resigned from OpenAI and joined Anthropic. This flow reinforced Anthropic’s position as the safety-focused alternative to OpenAI, and it gave Anthropic a steady stream of experienced researchers who had been trained at the frontier of the field.
What It Means
The founding of Anthropic was the most significant schism in the AI industry. It created a credible competitor to OpenAI, it gave the AI safety community an institutional home, and it gave the market a choice between a company that prioritised speed (OpenAI) and a company that prioritised safety (Anthropic).
The schism also illustrated the central tension in the AI industry: the tension between building AI quickly and building it safely. OpenAI, under Sam Altman, has generally prioritised speed. Anthropic, under Dario Amodei, has generally prioritised safety. Both companies are building frontier AI, but they are doing it with different philosophies, different priorities, and different risk tolerances.
The competition between OpenAI and Anthropic is, in some respects, beneficial — it provides alternatives, drives innovation, and provides a check on the power of either company. But the competition also has risks. If the two companies are racing to build more capable AI, and if one is willing to take more risks than the other, then the competition could push both toward less safe practices — the “race to the bottom” dynamic that safety researchers have warned about. Whether this dynamic will materialise is an open question. The answer will depend on the choices both companies make in the coming years — and on whether the safety culture Anthropic was founded to embody can survive the pressures of competition.
The longer-term question is whether Anthropic’s safety-focused model is, in fact, viable. The company has raised billions of dollars from investors who expect returns, and the returns will come from building and selling capable AI systems. There is a tension, at the heart of Anthropic’s structure, between the safety mission and the commercial imperative. The company’s Public Benefit Corporation structure is designed to manage this tension, but the structure is untested. The November 2023 OpenAI board crisis (see B39) was, in part, a story about what happens when a safety-focused governance structure collides with commercial reality. Anthropic has, so far, avoided a similar crisis. But the pressures that produced the OpenAI crisis — the pressure to commercialise, the pressure to move fast, the pressure to compete — are present at Anthropic too. The question of whether Anthropic can remain safety-focused while building frontier AI is one of the most important questions in the AI industry. The answer is not yet known.
- “Anthropic” — Wikipedia — A well-sourced overview.
- “Constitutional AI: Harmlessness from AI Feedback” — Anthropic, December 2022 (arXiv:2212.08073) — The Constitutional AI paper.
- “They Created a Company to Build Safe A.I. Now They’re Racing Against Their Former Employer” — The New York Times — Profile of the Amodeis.
- “Dario & Daniela Amodei: Building AI You Can Trust” — Minds & Machines, P18 — The main series’ profile.
- “Zooming in on Neural Nets” — Chris Olah, Distill, 2017 — Early interpretability work by Anthropic’s head of interpretability.
This piece is part of Minds & Machines: Beyond the Series. The companion piece B39 — The OpenAI Board Crisis covers the November 2023 crisis that validated the Amodeis’ governance concerns. The main series profiles the Amodeis in P18 — Dario & Daniela Amodei.
Was anthropic’s founding inevitable — the product of forces too large to redirect — or was it a series of choices, each of which could have gone differently? The answer matters, because it determines whether the future is something that happens to us or something we make.
Subscribe
Get new articles delivered to your inbox. No spam — just the story behind the screen.