P Doom: Who Gets to Put Odds on the Apocalypse?
P doom, usually written p(doom), means someone's estimated probability that advanced AI causes an existential catastrophe, such as human extinction. It is a...
P doom, usually written p(doom), means someone’s estimated probability that advanced AI causes an existential catastrophe, such as human extinction. It is a personal judgement under uncertainty, and its meaning changes with the outcome, timescale and assumptions behind the number.
Quick summary: A 2023 survey of AI researchers found that 38% to 51% of respondents assigned at least a 10% chance to outcomes as bad as human extinction. METR measures AI capabilities, while the argument about its independence concerns who should judge their risks.
The strange part of the latest argument is how much depends on a footnote. Draw a relationship map without dates and a former trustee becomes a current power broker. Add the dates back and the same network tells a different, more complicated story.
What does p doom actually mean?
P doom means an estimate of the probability of an AI catastrophe, usually one threatening humanity’s survival or long-term future. The letter p stands for probability. The troublesome word is doom, because people can use it for different outcomes without saying so.
One person might mean human extinction. Another might include humanity permanently losing control over its future. Someone else may be speaking loosely about enormous disruption. Those definitions describe different events, so their percentages cannot be compared sensibly until the event is made explicit.
The timescale matters just as much. A forecast about the next decade and a forecast about any point in the future answer different questions. A small annual risk can also accumulate over time, so shortening or extending the window changes the calculation.
Then ask what the estimate assumes. Does development continue at its present pace? Do governments intervene? Do safety techniques improve? A forecast conditional on weak safeguards cannot simply be presented as the expected result after substantial new safeguards have been introduced.
The useful version of a p doom number therefore comes with a sentence: the estimated probability of this particular outcome, within this period, under these conditions. Without those qualifications, a precise-looking percentage can conceal a surprisingly vague claim about the future.
Why there is no official p doom score
There is no agreed instrument that measures a universal p doom score. An individual’s estimate might draw on technical work, historical analogies, judgement or formal modelling. Giving it a decimal place does not establish how accurate it is or make other estimates comparable.
The researcher survey above records people’s beliefs about possible outcomes. It does not observe repeated civilisational experiments or measure an extinction frequency. Its percentage range also reflects different question versions, a reminder that wording matters even when respondents have relevant technical expertise.
That does not make expert judgement worthless. It makes its assumptions part of the evidence. Ask why someone chose their number, which developments would change their mind, and whether their definition matches the definition used by the person they disagree with.
| Evidence | What it tells you | What it cannot establish alone |
|---|---|---|
| A person’s p(doom) | Their estimate under stated assumptions | An objectively measured extinction probability |
| A researcher survey | How respondents judge specified risks | Whether their collective forecast is correct |
| A METR task evaluation | Performance under particular test conditions | A universal measure of real-world autonomy |
| A donor or employment disclosure | A documented financial or professional connection | Control over an evaluator’s conclusions |
| A conflict policy | The organisation’s stated safeguards | Whether every safeguard works in practice |
The resignation that pulled the argument into public
The resignation matters because it put an insider’s warning beside a much larger argument about who should oversee advanced AI. It also produced several different stories at once: a researcher’s objections, proposals for external scrutiny, and allegations about the people supporting that scrutiny.
Jacob Coxon
The researcher is Jacob Coxon, who posts as @hilbertspaess. In the resignation post itself he wrote that after three years of pretraining research at OpenAI and Anthropic, “neither company is acting responsibly”, and that both are “racing straight to self-improving superintelligence and gambling with our lives”. Coxon set out his reasoning at greater length in a PBS News Hour interview on 17 September 2026.
What made the post travel was the response from inside the company. Evan Hubinger, an alignment lead at Anthropic, publicly agreed, writing that “we really do earnestly believe AI could kill all humans” and putting his own figure above 10% within the next decade. A serving researcher endorsing the warning is what carried it beyond AI circles.
Coxon’s central concern was recursive self-improvement: AI systems helping develop increasingly capable successors. He argued that reducing the need for human judgement in that process could make progress harder to control. That is his assessment of a possible trajectory, rather than an observed catastrophe.
What is METR AI?
METR, in the AI safety context, is Model Evaluation and Threat Research, a nonprofit that evaluates AI capabilities and risks. Its work includes studying what systems can accomplish autonomously and developing evidence that researchers, companies and policymakers can use when assessing those capabilities.
METR’s own description sets out a research and evaluation organisation, pronounced “meter”. Its role is easy to misunderstand when the name appears beside an alarming forecast. METR does not supply an official p doom number or turn a benchmark into an extinction prediction.
The evaluator’s position still deserves scrutiny. Access to unreleased systems can improve the evidence available, while dependence on that access can create pressure. That tension is a reason to inspect contracts, publication rights and funding arrangements alongside the technical results an evaluator publishes.
What METR’s AI safety tests can actually tell you
METR’s task time horizon methodology describes task difficulty using the time a human expert would need to complete the work. A horizon at a specified success rate estimates the human-duration difficulty at which an AI agent reaches that rate on the evaluated tasks.
Consider an illustrative two-hour horizon at 50% reliability. That would mean success on about half the relevant tasks that take a human expert two hours. It would not mean the agent reliably works for two hours unattended, or completes half of every workplace job.
The task collection matters. METR’s published work focuses heavily on software engineering, machine learning and cybersecurity tasks with assessable outcomes. Normal work can involve missing context, ambiguous priorities and consequences that a test environment does not reproduce. Generalising requires further evidence beyond the benchmark.
A capability measurement can nevertheless matter for risk. Better autonomous performance may expand what an AI can attempt. Judging danger also requires examining access, incentives, oversight and the consequences of failure. A chart of capability growth leaves those additional questions open for investigation.
The proposal that made the evaluator matter
The dispute becomes easier to follow once research is separated from proposed authority. A group can publish valuable evaluations without having the right to stop another company’s work. Giving evaluators greater access or influence is a governance choice that deserves its own justification.
Dario Amodei
In his September 2026 essay, “We Must Pace the Frontier”, Anthropic chief executive Dario Amodei proposed embedded evaluators with ongoing, employee-like access. He named METR as an example and argued for coordination over the pace of increasingly capable AI development.
The proposal covers scrutiny of safety practices, incidents and development processes, alongside model assessment. It also describes publication rights, subject to limited redactions. Those details matter because an evaluator’s ability to inspect a company and its ability to report criticism are separate practical requirements.
Naming an organisation as an example does not establish an exclusive appointment or a legal veto. The harder question is what powers any eventual arrangement would confer, who would select the evaluators, and whether competitors and the public could challenge their findings.
There is also a straightforward incentive question. A company advocating stricter oversight may sincerely believe it is necessary while benefiting from particular rules. Both possibilities can be examined together. Evaluating the proposal requires looking at its effects, beyond guessing whether its author’s motives are pure.
Follow the funding, then check the dates
The relationship web contains real connections. The task is to describe each one accurately: investment, charitable support, employment, marriage or governance. Those relationships can have different consequences, and none should quietly become another kind of relationship as a diagram travels around the internet.
Jaan Tallinn and Dustin Moskovitz
Anthropic’s May 2021 funding announcement identifies Jaan Tallinn as the investor who led its Series A and Dustin Moskovitz among the participants. That documents investment in Anthropic. It does not, on its own, establish either person’s authority over METR’s research or its published conclusions.
The Survival and Flourishing Fund’s 2025 recommendations also list support for METR from Tallinn, including a conditional matching pledge. That is a relevant funding connection to disclose. A recommendation, a conditional pledge and money already received should be described according to their actual status.
Moskovitz also holds governance roles at Coefficient Giving, formerly Open Philanthropy. Its governance explanation distinguishes the organisation’s recommendations from grants awarded through several entities, including Good Ventures. Collapsing these organisations into a single donor can obscure who actually made a particular funding decision.
Holden Karnofsky and Daniela Amodei
Holden Karnofsky co-founded Open Philanthropy and later joined Anthropic. His July 2025 interview with 80,000 Hours discusses his work there and his marriage to Daniela Amodei, Anthropic’s co-founder and president. The relationship and resulting financial interest are disclosed, rather than inferred from a photograph.
That connection is relevant when considering interests in this network. It does not tell us who designed a particular METR test, approved an assessment or could suppress a result.
METR’s August 2026 funding update reports around $71 million in commitments raised during the preceding six months. It says it accepts no frontier AI company funding or donations directed by their staff, while acknowledging significant access to free model tokens from AI companies.
Cash funding, donated access and a donor’s investments are different relationships. A credible critique identifies which channel of influence it means, rather than merging them.
The tinfoil-hat case, in its strongest form
The sceptical case is worth stating properly rather than waving away. Its most widely shared version, laid out at length in a September 2026 breakdown by ThePrimeagen, rests on three things: the timing, the money, and one very unusual account.
The account and the timing
The resignation post was reported to pass 120 million views within roughly a day, from an account created that year, carrying two posts and effectively no follower history. Elon Musk remarked publicly that he had never seen an account go that viral with no prior interactions, a point repeated throughout the commentary.
Coxon has said he gave journalists exclusive access before resigning, which is offered as the explanation for a Wall Street Journal piece appearing minutes ahead of the post itself. Critics further note a Netflix documentary on AI danger and a Time magazine cover about pausing AI landing in the same week.
Coincidence is the boring explanation and remains entirely possible, particularly since a documentary takes years to produce. Coordinated communications is the interesting one. Neither has been demonstrated, and a synchronised media week is evidence of media planning rather than of a conspiracy to capture regulators.
The cartel accusation
The sharpest public version came from investor David Sacks, quoted in that breakdown, arguing that companies should “stop pretending antitrust law has to be suspended so you can form a cartel” and “stop pretending METR is independent when it is intertwined with Anthropic’s investors and staff”.
That is an allegation about structure and incentives, not a finding of fact, and it is the fair version of the worry. Amodei’s proposal does invite frontier companies to coordinate on the pace of development, and competitors coordinating on output is the exact shape competition law exists to examine.
The funding history gives the accusation something concrete to point at. FTX, before its collapse, was reported to have given over a million dollars to METR, and Sam Bankman-Fried was an early Anthropic investor. Those are historical relationships involving a now-imprisoned donor, which is different from present-day control.
The contradiction that is harder to wave away
One detail resists the charitable reading. Coxon stated he worked with no third parties in arranging his resignation, while reporting around it indicates he had been in contact with policy figures opposed to frontier AI beforehand. A discrepancy about coordination, inside a story about coordination, is a reasonable thing to flag.
It does not follow that his technical concerns are wrong. People can present a genuine warning and also manage its release carefully. It does mean the presentation was more organised than described, which is precisely the sort of claim that should be checked rather than assumed in either direction.
The counter-theory nobody checks
There is a simpler reading that cuts the other way entirely. Sam Altman and Elon Musk both publicly agreed with Amodei’s call to pace the frontier, which is unusual enough to notice. The cynical interpretation is not that they joined a cabal, but that they encouraged a competitor to restrain himself.
Unilateral self-restraint by one laboratory is a straightforward gift to the others. That reading requires no secret meetings and no shared ideology, only ordinary commercial interest. It is worth holding alongside the cabal theory, because it explains the same observed agreement with considerably fewer assumptions.
Does the relationship web prove a cabal?
The relationship web does not prove a coordinated cabal controlling AI oversight. It demonstrates overlapping financial, professional and personal connections that warrant scrutiny. Establishing the stronger allegation would require evidence of coordinated decisions or control, beyond the existence of people who know one another.
Paul Christiano and Ajeya Cotra
METR’s origins help explain another cluster of connections. Its updated spinout announcement describes its separation from the Alignment Research Center. It also states that Paul Christiano declined the planned board and adviser roles because of his new role at the US AI Safety Institute.
Anthropic’s Long-Term Benefit Trust announcement includes a crucial later footnote: Christiano stepped down as a trustee in April 2024. A chart presenting him as a current trustee therefore misstates the relationship. Historical service belongs on the map with its dates clearly attached.
Ajeya Cotra’s METR biography identifies her as technical staff working on threat modelling and risk assessment, following work at Coefficient Giving. That documents a professional connection between organisations in the network. It does not establish that her former employer directs her present research conclusions.
These corrections neither erase the network nor certify its independence. They narrow the claims to what the records support. An old role, a declined appointment and a current job are three different facts, even when putting them into identical boxes makes a simpler story.
Shared beliefs can still matter without a conspiracy. People who agree about the most serious risks may favour similar methods, hire from similar circles and overlook similar objections. Testing for that possibility requires methodological criticism and independent review, alongside accurate disclosure of financial interests.
How would you check whether an evaluator is independent?
You check whether an evaluator is independent by examining funding, decision rights, publication freedom, conflicts and responses to criticism. Independence should be demonstrated through arrangements and behaviour. An organisation’s name, nonprofit status or stated commitment to safety cannot answer all those questions by itself.
Start with the agreement governing access. Can the evaluator choose tests, inspect failures and publish an unfavourable conclusion? Can the assessed company delay publication or narrow the scope? Useful disclosure would explain those permissions clearly enough for outsiders to understand where the evaluator’s freedom ends.
Then inspect conflicts at the level of the assessment. METR’s August 2026 conflict policy addresses financial holdings, employment and close personal relationships, with disclosure and mitigation requirements. A written policy is useful evidence of intended safeguards; published assessments show how those safeguards are applied.
The practical picture is already broader than a single proposed evaluator. On 18 September 2026, Anthropic announced an embedded evaluation partnership with Accenture, involving its Faculty business. Anthropic says it will fund that work directly and describes the arrangement as nonexclusive, with additional evaluators envisaged.
Finally, look for the possibility of disagreement that has consequences. Could another team reproduce the work, challenge the interpretation or publish a competing assessment? An oversight system becomes more credible when its conclusions can be contested through evidence, including evidence inconvenient to its supporters.
Frequently Asked Questions
What does a high p doom number mean?
A high p doom number means someone assigns a relatively high probability to the catastrophic outcome they have in mind. Its significance depends on their definition, deadline and assumptions. Ask for those details before comparing it with another person’s lower estimate.
Is p doom a scientific measurement?
P doom is generally a forecast or subjective probability, sometimes informed by scientific research and explicit models. It has no universally agreed measuring instrument. The useful questions concern the evidence behind the estimate, its assumptions and what observations would cause it to change.
What is METR AI used for?
METR’s AI evaluations help examine capabilities, autonomous task performance and potential risks under specified conditions. Researchers and decision-makers can use that evidence when assessing systems. Interpreting a result still requires reading the methods, understanding the task sample and respecting the limits of the test.
Does Anthropic own METR?
The public sources discussed here identify METR as a separate nonprofit, not an Anthropic subsidiary. There are documented connections through investors, funders and professional networks. Those connections justify questions about independence, but they do not establish ownership or control over its assessments.
Is the AI safety cabal theory true?
No evidence presented so far establishes a coordinated group controlling AI oversight. The documented connections are real and worth disclosing, and several widely shared charts misstate dates and roles. Overlapping networks and shared beliefs are a legitimate concern without amounting to a proven conspiracy.
Can METR calculate the chance of AI causing extinction?
METR’s capability tests do not directly calculate the probability of human extinction. They can inform risk assessments by showing what systems accomplish under test conditions. Moving from those results to a p doom estimate requires additional assumptions about deployment, safeguards and future behaviour.
Trust the source of truth before the confident answer
For a business, the immediate question is whether you can trust what an AI hands you today. A purchasing recommendation can sound convincing while using an obsolete stock figure. A production summary can read perfectly while overlooking the latest change to an order.
In the systems we build, inventory, orders, purchasing, production and reporting need a shared source of truth. The practical test is traceability: which records support this answer, when were they updated, and can the person responsible verify them before taking action?
That applies to familiar operational decisions. A safety stock calculation depends on the demand and lead-time figures supplied. A bill of materials needs the correct components and quantities. AI can help explain either, but fluent wording cannot repair incorrect inputs or missing records.
The next useful step is to identify where your team loses that connection between an answer and its evidence. Our operations demos show the kinds of workflows we build. Start with the actual gap, then choose a system proportionate to the problem.