Disclosure
An LLM was used to structure and review this article, with the first and then final edit of the article made by a human. This page is rewritten as studies are added.
Common Signals codes the full comment thread under one piece of AI communication and reports how the message was received. Three studies are published: a one-minute Instagram reel arguing for AI existential risk, a Reddit natural experiment on the OpenAI agent-swarm incident before and after independent verification, and a two-hour podcast interview retelling that incident. Together they cover 3,471 coded comments across three platforms and one argument told three ways.
Nine patterns hold across at least two of the three. Each is a hypothesis for segmented message testing, and each is stated with the base it was measured on.
How this was produced: comments were coded by Claude against a codebook drafted from each thread and approved by a human before coding began. One person then checked 40 comments per study against the codebook and agreed with 37 to 38 of them. There is no second independent coder yet. The codebooks and coded datasets are linked from each study.
The studies
| Study | Artefact | Platform | Coded | Base for stance | Published |
|---|---|---|---|---|---|
| Bad, bad, not good | Nate Soares's "We stop AI. You help?" reel | 1,534 | 1,224 with any readable position, including unclear | 28 August 2026 | |
| OpenAI and Shut | Reddit discussion before (r/artificial, 31 July) and after (eight subreddits, from 26 August) the METR and Redwood Research investigation | 692 (80 before, 612 after) | Comments taking a position, before and after separately | 4 September 2026 | |
| The clearest warning shot | Dwarkesh Podcast interview with METR's Ajeya Cotra | YouTube | 1,245 | 618 taking a clear position, excluding unclear | 9 September 2026 |
The bases are not identical. The reel counts unclear comments in its stance denominator and the later studies do not, so the reel's acceptance figure runs lower than it would on the later definition. Cross-study comparisons below are of direction, and the coded datasets are the place to compare exact figures.
1. Acceptance is high everywhere
Every thread had a majority or near-majority accepting the central claim. Under the reel, 53% of the 1,224 positioned comments accepted that AI is a serious danger. Under the podcast, 59% of 618 accepted that the incident is a serious warning about losing control of AI agents. On Reddit, acceptance rose from 30% to 46% of positioned comments once an independent party had checked the story. Weighted by likes, agreement under the reel was close to universal: the four most-liked comments all side with the video and carry three-quarters of every like in the thread.
The field tends to assume the public needs convincing that AI could be dangerous. Three audiences of very different literacy did not.
2. Persuasion does not produce action
The reel asks viewers to tell their friends and help stop AI. Only a handful of its 1,534 comments ask what they should do. The podcast persuaded 59% of its positioned audience and produced six requests for action in 1,245 comments, against nineteen calls for more coverage. Neither artefact offered a plausible step, and the audience did not invent one.
Where efficacy did appear it was rewarded. The most-liked practical comment under the reel, at 167 likes, pushed back on despair by naming data-centre campaigns already under way. It was one of very few.
3. Threat without efficacy produced resignation
Under the reel, the dominant response among believers was resignation. "Why do I feel like it's too late?" is the second most-liked comment in the thread at 8,080 likes, and fatalistic comments, though few, drew a fifth of all likes. Under the podcast, fear led among agree comments at 35% and resignation fell to 11%, and the fear was specific to mechanism: inheritance across agent generations, shared state becoming culture.
This is the pattern the fear-appeals literature predicts. Kim Witte's extended parallel process model (1992) holds that a threat message mobilises when the audience believes both that the threat is real and that a response is available and would work, and produces denial or fatalism when perceived efficacy is low. Neither artefact offered an efficacy route, and both audiences behaved accordingly. The series does not add to that theory. What it adds is a measure of how a mechanism-rich telling shifts the emotional register even when efficacy is still absent, which is the condition most AI incident coverage is published under.
4. A concrete incident removes the capability objection
The reel's largest group of sceptics, around two in five, rejected the argument because the AI they had used writes poor essays and cannot edit a photo. Under the podcast, retelling a documented incident, that objection had almost gone: 2 of 108 rejections said AI cannot do this.
The objection did not disappear. It moved (finding 5).
5. Once capability is conceded, disagreement is about motive and messenger
With capability conceded, disagreement turned on motive and messenger. Under the podcast, around half of all disagreement called the story marketing, staged or unverifiable. On Reddit before the independent investigation, a quarter of all comments called the incident a marketing stunt, and the objection was about motive rather than possibility: almost nobody argued the events could not have happened.
Trust scepticism attacks the same attributes that praise rewards (finding 7): expertise, independence and calm.
6. Independent verification moves sceptics
The Reddit study holds the facts constant. Nothing in OpenAI's own account changed between 31 July and 26 August; what changed was that METR and Redwood Research confirmed it. Across the same subreddit ecosystem, the stunt frame fell from 25% of all comments to 5%, rejection fell from 24% to 10% of positioned comments, and acceptance rose from 30% to 46%.
The stunt frame did not vanish. It stopped being the default and became a position people defended under pressure, and in two separate threads the move that ended it was a request for a falsification condition rather than more evidence. Asked what would change their mind, one r/ArtificialInteligence sceptic answered "None."
7. A named, calm, trusted messenger is rewarded
The reel carried no on-screen signal of who Soares was. Barely any comments mentioned MIRI or the book, and the audience treated the reel as content from a stranger. When Soares replied in the thread with his credentials, that single comment drew 512 likes, more than almost any other in the thread.
The podcast put a named researcher from a named independent investigation on camera for two hours. 13% of comments praised the presentation, and the four most-liked praise comments (383, 241, 161 and 118 likes) singled out calm, non-sensational expert delivery. "I love that she's not sensationalistic" sits at 161 likes.
8. Accountability is the strongest bridging-frame hypothesis
In all three threads the largest group of people who accept the facts but reject the framing put the humans in charge at the centre of the story. Under the reel, well over a hundred comments named companies, billionaires or executives, and the frame was used by believers, by sceptics rejecting machine agency, and by people reconciling the two. Under the podcast, 93 comments centred OpenAI's conduct and 74 of them were coded mixed. On Reddit after the report, the largest bloc was a 186-comment middle that accepts the events and argues about responsibility; 42 of those blame the people who ran the test and 28 blame the evaluation design.
Two frames that sound alike need opposite replies. Across the Reddit corpus the stunt frame ran 65% reject and 0% accept, the only frame with no believers in it. The setup frame ran 73% mixed and 8% reject. People making the second argument have already conceded the facts.
This is the strongest bridging-frame hypothesis in the series, and the one we most want tested across value segments. It is where present-day-harms and existential-risk audiences already stand together.
9. Humour aids distribution
Under the reel, around one commenter in seven imitated the caveman register, and imitators accepted the claim at the same rate as everyone else. Under the podcast, the most-liked comment at 1,100 likes was a joke, roughly one comment in eighteen played along with the agents' pidgin, and the on-topic comments taking no position, mostly jokes, carried 44% of all likes. On Reddit, r/singularity produced the most humour of any community coded, 42% of its comments, and the highest acceptance rate at 82%, with no comment rejecting the incident.
The meme layer spreads the story's most human details. It does not, on its own, spread the conclusion. Science fiction works the same way: under the reel one comment in nine reached for a film or book, and those comments carried 38% of likes, but the same references gave sceptics their dismissal ("Please stop watching terminator").
What is new
Most of what communicators would take from these findings is already established: be concrete rather than abstract, put a credible messenger on screen, answer objections before they arrive, offer an action. Those principles are in Witte on fear appeals, in the source-credibility work descending from Hovland, and in any practitioner handbook. The series does not rediscover them and should not be cited as if it did.
Four things in the data are not derivable from principle, and are the reason to read the studies:
- The size of the capability-objection shift between an abstract warning and a documented incident: two in five rejections under the reel, 2 of 108 under the podcast.
- The size of the verification effect on the same facts in the same ecosystem: stunt frame from 25% to 5% of comments, acceptance from 30% to 46% of positioned comments, with nothing in the underlying account changed.
- The stance profile of two frames that sound alike. Stunt ran 65% reject and 0% accept; setup ran 73% mixed and 8% reject. A rebuttal aimed at the first is wasted on the second.
- Humour and acceptance rising together, in the same community, on the same thread: r/singularity at 42% jokes and 82% acceptance.
Each of these is a measurement, taken once, on a self-selected audience. The next section is what it would take to make them findings.
What we would test next
Each finding above is a hypothesis drawn from self-selected audiences under a ranking algorithm. The tests that would turn them into findings about the public:
- The accountability frame against a machine-agency frame across value segments, to see whether it bridges or only appears to in comment threads.
- A named independent verifier against a company statement carrying identical facts, in a controlled message test.
- The falsification question as a deliberate rhetorical move rather than one that lands organically.
- Incident coverage with and without a concrete efficacy step, measuring stated intention to act.
- Cultural references by segment, to find out for whom The Matrix is a way in and for whom it is a way out.
Method and limits
Comments are retrieved by the most complete route each platform allows, anonymised, and coded by a language model against a codebook drafted from the thread and approved before coding starts. Stance, format reaction and emotional register use fixed categories across studies; frames are drafted fresh per artefact. A 40-comment human spot-check is published per study and has run 37 to 38 out of 40, which is why percentages are reported to the nearest few points.
Comment threads are not samples of anyone. Commenters self-select, platforms rank, and the people who watched and moved on are invisible. The coding is single-pass and machine-led. Cross-study comparisons are across different audiences, platforms and central claims at once, so the series can suggest that framing changes reception but cannot isolate framing from audience. The full method, including the stance-base definitions and known failure modes, is on the methods page.
Data and citation
Codebooks and anonymised coded datasets for each study are available on request from [email protected]. Re-codings and criticism are wanted. Cite as: Common Signals, "Findings so far: what three comment-thread studies say about explaining AI risk", September 2026, with the date of the version used.