Who this article is for
Anyone deciding whether and how to put AI-risk messaging in front of an X audience specifically, and what changes when the messenger is quitting rather than reporting on someone else's incident.
TL;DR
- Of 1,840 replies taking a readable position, 42.5% agree with the central claim, 39.9% disagree and 17.6% are mixed. Weighted by likes the gap widens only to 50.7% against 37.5%, nothing like the 90%-plus agreement seen elsewhere in this series.
- The most-liked reply in the whole dataset took no position at all: “Well this was a pleasant read right before bed.” drew 28,913 likes, more than double the next comment. Reactions with no argument in them (jokes, one-liners, agreement without elaboration) account for under a quarter of replies but more than half the thread's total engagement.
- Capability scepticism, not geopolitics or credential attacks, is the single largest source of disagreement: one in five sceptical replies doubts AI can do what the thread describes.
- Emotion splits cleanly along stance. Agreeing replies run mostly fearful or neutral; disagreeing replies run mostly angry or neutral. Rejecting the claim reads as anger, not dismissal.
- A newly identified frame, replies that accept the danger but locate the fault in human greed or the desire for control rather than the technology itself, is the strongest agreement signal in the dataset after outright endorsement.
Introduction: a resignation, not a third-party account
On 9 September 2026, Jacob Coxon, a former pretraining researcher at both OpenAI and Anthropic, posted a seven-tweet thread on X. The first line: “I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
Six more posts followed, elaborating the argument: the technology's speed and power, that the fear inside labs is sincere rather than marketing, why researchers stay despite believing the risk is real, that entering what he calls the “endgame” is a gamble that should not be launched from a private company's Slack, qualified optimism about coordination between labs set against pessimism that a global race can be averted without a costly step such as a temporary ban, and a direct appeal to lab researchers to weigh what the next few years will feel like.
The root post alone drew 20,553 replies. This is the fourth artefact in this series and the first from X: earlier studies covered an Instagram reel, a YouTube podcast episode and a Reddit before-and-after natural experiment around the same OpenAI and Hugging Face incident. It is also the only artefact in the series delivered by the person it is about, in his own voice, rather than reported or discussed by someone else, and the only one where the messenger and the message are the same person.
The original thread: x.com/hilbertspaess/status/2097476196791709843.
Method
X gives no API or export route, so every reply was retrieved through a signed-in browser session reading the on-page feed directly, across all four sort views the platform offers. Each view independently stops rendering new items after a few hundred, regardless of how many replies the post has, so the 25,219 replies across all seven posts were never fully reachable. 2,639 items were retrieved; 2,480 remained after excluding promoted content and Coxon's own thread posts (which the scraper also picks up); 2,391 carried text to code; 1,840 took a readable position, the base for stance percentages below. Coverage per post ranged from 2.6% (the root post, with 20,553 replies) to 93% (the smallest follow-up post, with 250), tracking an apparent rendering cap of roughly 350 to 450 items per post rather than a fixed share.
Every reply was coded by Claude, in parallel batches of around 190, for stance towards Coxon's central claim, dominant frame, reaction to format, emotional register and eight mention flags, against a written codebook drafted from the thread itself. Ben Matthews spot-checked a random 40-reply sample at 38/40 agreement before publication.
One sentence of limits: this is a self-selected sample, and a small, algorithmically selected slice of a much larger reply pool rather than a near-complete capture, X's own ranking chose which replies were even reachable, and likes measure salience, not persuasion. Full method, codebook and dataset at the Common Signals method page.
Coded by Claude and spot-checked by a human against the codebook. Browse every coded reply and the full codebook: explore the data and codebook.
The numbers
Of the 1,840 replies taking a position, 42.5% agree with Coxon's central claim, 39.9% disagree and 17.6% are mixed. Weighted by likes, agreement carries 50.7% of engagement against disagreement's 37.5% and mixed's 11.8%.
Every earlier artefact in this series drew lopsided agreement: 67.7% on the Amanpour segment, 98.2% among comments that engaged with the core claim on the Soares reel, rising further once weighted by likes. Here the count-based gap is under three points and the weighted gap, while real, does not come close to reproducing that pattern. The most plausible explanation is the audience, not the artefact: X repliers on this thread skew elite, industry-adjacent and combative, closer to a peer audience for Coxon than a general public one, and peers argue back.
A joke outperformed the other arguments
The single most-liked reply in the dataset is five words long and takes no position: “Well this was a pleasant read right before bed.” 28,913 likes, more than double the next comment and more total engagement than most entire threads elsewhere in this series.
It was not alone. format_only, the frame for replies reacting only to tone, length or format rather than content, is 13.0% of coded replies but 32.6% of all engagement. Add general_agreement, endorsement without a distinct argument (9.5% of replies, 25.7% of engagement), and under a quarter of the thread's replies account for well over half its total engagement.
“I wonder why progress looks so much like destruction” (attributed to John Steinbeck)X reply, 19,235 likes
“Nobody in the replies even remotely understands what he's saying here. Self improving intelligence means in the near future no human on earth will ever be able to understand it.”X reply, 11,081 likes
The pattern is not that argument fails to attract engagement. It is that pithy reaction, whether a deadpan joke or a one-line endorsement, consistently outcompetes it. A communicator judging what “landed” by likes alone would conclude the audience mostly agreed and mostly said very little, which is technically true and would miss most of the actual disagreement happening lower down the thread.
What the sceptics said
Among the 1,058 replies coded disagree or mixed, capability scepticism, doubt that current or near-term AI can do what Coxon describes, is the largest single frame at 19.8%, ahead of china_race (12.8%) and everything else.
“With all respect, this is bizarre. It's a ridiculous take. Humans have evolved over hundreds of thousands of years. We aren't going to die out because a token prediction model gained sentience. Get a grip.”X reply, 8,285 likes
“Dude all we have to do is turn off the Data Centers and then they are gone. How is this the end of civilization?”X reply
The geopolitical objection, that unilateral restraint just hands the advantage to China, is present and substantial (66.9% disagree within the frame) but plays second to a more basic argument: many repliers simply do not believe the danger described is technically real, independent of what anyone should do about it.
ipo_marketing_cynicism, reading the whole thread as promotional timed to Anthropic's pending IPO, is smaller in volume (7.8% of disagree and mixed replies) but punches above it in engagement, carrying 6.9% of all likes from only 3.6% of replies.
“okay so buy the IPO?”X reply, 7,802 likes
Fear and anger
Among replies that agree with Coxon, 33.2% are coded fearful and 35.5% neutral, with humour rare (4.3%). Among replies that disagree, 34.4% are coded angry and 45.4% neutral, with fear almost absent (1.0%).
Accepting the danger claim reads as frightened. Rejecting it reads as angry, not bored or dismissive. That distinction matters for anyone drafting a rebuttal: an angry sceptical audience is not a disengaged one, and treating disagreement as apathy misreads what is actually a live, emotionally invested argument on the other side.
Accepting the danger
One frame added specifically for this study, after the first pass under-covered the six follow-up posts, captures a pattern distinct from either straightforward agreement or disagreement: blame_people, replies that accept the danger is real but locate the fault in human greed, elitism or the desire for control rather than in the technology itself. 73.8% of replies in this frame agree with the central claim, nearly as strongly as outright endorsement.
“This is exactly why AI needs to be people-owned, not controlled by a handful of billionaires. AI needs democracy. The future of AI shouldn't be shaped solely around how much profit Big Tech can extract from it.”X reply, 787 likes
scifi, framing the claim through film or fiction, shows the same shape even more strongly (72.2% agree). On this thread, reaching for a familiar story or a structural explanation about human nature accompanies acceptance far more often than it accompanies mockery.
If you believed it, why did you just leave?
hypocrisy_critique, the argument that quitting is the wrong response if Coxon genuinely believes what he says, runs 48.1% mixed, 43.3% disagree and only 8.7% agree, the clearest bridging frame in the dataset.
“if you really believe this you stay on the job and sabotage the progress, what does resigning do? you think it's going to kill everybody and you decided not to fight?”X reply, 6,612 likes
This is not scepticism about the danger. It is a dispute about the right response to it, made by people who have already conceded enough of the premise to demand a different action from Coxon, not evidence against his claim. A rebuttal aimed at the underlying danger claim will not move this bloc; one aimed at why leaving, rather than staying and pushing for change, was the right call might.
What this means for communicators
X, at least on this thread, did not behave like the rest of this series. Every earlier artefact drew an audience that mostly agreed, more so once weighted by engagement. This one split close to evenly, and the audience that disagreed was substantive rather than dismissive: capability scepticism, not credential attacks or geopolitics, carried the argument.
The engagement pattern is the second surprise. On every earlier platform in this series, the most-liked comment was also one of the most substantive. Here it was a joke that said nothing about the argument at all, and pithy non-arguments as a class outengaged every camp of substantive discussion combined. Likes on X, on a thread this size, measure something closer to “this was quotable” than “this was persuasive,” more so than elsewhere in the series.
The messenger question also looks different when the messenger is the subject of the story rather than its narrator. Coxon's own credibility was contested, but not as the leading objection the way it was for Khlaaf on the Amanpour segment. The leading objection here was about the technology's capability, a claim that can be answered with argument rather than with proof of who someone is.
How to use this
-
1
Expect capability scepticism as the default objection on X, and answer the mechanism, not the messenger.
Nearly one in five sceptical replies here doubted AI could do what was described, ahead of geopolitics or credential attacks. A rebuttal built to answer “prove you're real” will miss the audience actually asking “prove this is possible.”
-
2
Do not read engagement as agreement on X threads at this scale.
The single most-liked reply took no position, and pithy non-arguments carried more total engagement than every substantive camp combined. Judging reception by likes alone on a viral X thread will systematically overstate how much of the audience actually agreed, and understate how much serious disagreement is sitting further down.
-
3
Distinguish “this can't happen” from “you should have stayed and fought,” because they need different answers.
The
hypocrisy_critiquebloc has already accepted much of the premise. Engaging it with more evidence for the danger claim wastes an argument on people who are not contesting that point; engaging it on why leaving, rather than staying to change things from inside, was the right call has a chance of landing. -
4
Treat “it's just human greed” acceptance as a real ally, not a distraction.
Nearly three-quarters of replies that reframed the danger as a story about human nature rather than the technology still agreed with the underlying claim. That is closer to a persuadable ally with a different emphasis than to an opponent, and worth a different kind of response than either agreement or disagreement gets by default.
Data and citation
The coded dataset and codebook are published with the Common Signals method page. Cite as: Common Signals, '“Neither company is acting responsibly”: How Jacob Coxon's resignation from Anthropic was received on X', September 2026. Companion pieces in the same reception series: the Soares reel analysis, the Dwarkesh episode analysis and the METR report before/after analysis. See also our theory of change and the full glossary of terms used across this series.
Disclosure
An LLM was used to structure and review this article, with the first and final edits made by a human.