This page documents the method behind the Common Signals comment analyses, currently the Soares reel study, the Dwarkesh Cotra episode study, the Amanpour Khlaaf segment study and the METR report before/after study. It exists so that the work can be checked, criticised and replicated. Where the method has weaknesses, they are stated here rather than smoothed over.
What these studies are, and are not
Each study codes the full comment thread under one piece of AI communication and reports how the message was received: who accepted its central claim, which frames commenters used, how they reacted to the format and messenger, and what emotional register dominated.
They are hypothesis generation, not findings about the public. A comment thread is a self-selected sample of one audience under one artefact, ranked by a platform's algorithm. Nothing in these pieces supports a claim about what any population believes. Their output is a set of specific, falsifiable hypotheses for the segmented message testing that Common Signals exists to run.
The METR report study is a partial exception to "one artefact": it compares the same Reddit ecosystem's reaction to OpenAI's own account of the incident against its reaction to the independent investigation six weeks later, so it can say something about a shift, not just a snapshot. It is described separately below.
The pipeline
1. Acquisition
Comments are retrieved by the most complete route each platform allows: platform APIs called from a logged-in browser session (Instagram), a yt-dlp comment export (YouTube, complete with threading and like counts), or a page capture where nothing better is available. Every study reports the number retrieved against the platform's on-page count and explains the gap.
Before anything leaves the collection environment, author identifiers are replaced with salted SHA-256 hashes and the salt is discarded, so the mapping cannot be reversed. @handles inside comment text are replaced with @user. Quotes in the published articles are reproduced verbatim, spelling intact; they are public comments and can be found by searching the platform, but we never publish usernames.
2. Codebook
Three dimensions are fixed across every study, with identical categories: stance towards the artefact's central claim (agree, disagree, mixed, unclear, na), reaction to the format (praise, mock, condescended, imitates, none), and emotional register (fear, resignation, anger, humour, hope, neutral), plus a small set of 0/1 mention flags.
The frame list is drafted fresh for each artefact from a sample of 150 random comments plus the 50 most engaged, merged into 10 to 16 frames, each with a definition and a verbatim example from the thread. The artefact's central claim is stated explicitly, and stance is coded against it; when the artefact argues against a prior narrative, as the Khlaaf interview does, agree means accepting that reframing, and the codebook says so.
The frame list is reviewed and approved by a human before any coding starts. Each study's full codebook is published alongside its dataset.
The METR report study departs from this in one respect: its frame list (17 frames) was drafted from its own two-thread corpus rather than the standard 150+50 sample from a single artefact, because the study spans nine separate threads across two time periods rather than one comment section. Its stance and format categories still follow the shared scheme.
3. Coding
Coding is done by a large language model (Claude), in parallel batches of roughly 190 comments, each coder reading the full codebook and, for replies, the parent comment's text. This is the method's biggest limitation and it is worth being precise about.
It is a single pass by a single model. There is no second coder and therefore no inter-rater reliability statistic. LLM coding has known failure modes: drift from the codebook over a long batch, over-use of residual categories, difficulty with irony and with comments that fit two frames, and a tendency to impose coherence on ambiguous text. We mitigate mechanically (validation that every coded value is from the allowed set and every comment is coded exactly once) and by review: a random 40-comment sample from each study is checked by a human against the codebook, and the agreement rate is published per study. Where a high-engagement comment materially affects engagement-weighted results, its codes are individually reviewed, and any correction is noted. Spot-check agreement has run 37 to 38 out of 40 across the three video-platform studies, which is why percentages are reported as accurate to a few points, never to decimals.
The METR report study has not had a spot-check hit rate published or logged. Until one is run, treat its reliability as consistent with the rest of the series in method but unverified in practice.
A standing invitation: the codebooks and coded datasets are published, and we would welcome anyone re-coding a sample and telling us where the codes are wrong. That is the fastest way this method improves.
4. Analysis
Every figure is computed from the merged coded table by published scripts: stance shares on stated bases, engagement-weighted shares, frame distributions, disagreement broken down by frame, emotion by stance, mention flags, engagement concentration, and top comments per frame. Every quote in an article is matched verbatim against the dataset before publication.
Two definitional notes, for anyone comparing across studies. First, the stance base: the reel study reported stance out of all comments with any readable content including unclear (1,224), while the later studies report stance out of comments taking a clear position, excluding unclear (618 and 191). The articles state their base wherever a percentage appears, but the bases are not identical across the series and cross-study comparisons should use the underlying datasets. Second, engagement weighting: likes measure salience under a ranking algorithm, not persuasion, and in every thread so far the top ten comments carry half or more of all likes, so weighted figures describe what the algorithm surfaced.
The METR report study breaks the second convention: Reddit page captures carried vote counts on only one of its nine threads, so almost all of its figures are shares of comments, not shares of what most readers saw, and engagement-weighted results are not available as a check the way they are for the other three studies.
The studies so far
| Study | Platform and route | Counted on page | Retrieved | Coded (with text) | Clear position | Spot-check |
|---|---|---|---|---|---|---|
| Soares reel, August 2026 | Instagram, comment API via logged-in session | 1,703 | 1,666 (98%) | 1,534 | 1,224 incl. unclear | 37/40 |
| Dwarkesh Cotra episode, September 2026 | YouTube, page capture | 1,569 | 1,249 (80%) | 1,245 | 618 excl. unclear | 37/40 |
| Amanpour Khlaaf segment, September 2026 | YouTube, yt-dlp export | 283 | 283 (100%) | 283 | 191 excl. unclear | 38/40 |
| METR report, Reddit, before/after (31 July and 26 Aug to 3 Sept 2026) | Reddit, page captures across 9 threads (1 baseline, 8 comparison) | not published per thread | 694 | 692 | not published as a single base; reported as % accepting/rejecting/mixed by period | not published |
The Dwarkesh study's shortfall is unexpanded reply threads in the page capture, which also lost threading, so that dataset leans towards top-level comments; its article says so. The reel's shortfall is deleted or hidden comments. The METR report study isn't really commensurable with the other three rows: it's nine threads across two time periods rather than one thread under one artefact, most of its threads carry no vote data at all, and its published article doesn't state a single retrieved-per-thread or spot-check figure the way the others do.
Known limitations, in order of importance
Comment threads are not samples of anyone. Commenters self-select, platforms rank, and the silent majority who watched and moved on are invisible.
The coding is single-pass and machine-led. Until a second, independent coding of a full study exists, agreement rates from 40-comment spot-checks are the only reliability evidence, and they are thin. For the METR report study, even that thin evidence doesn't yet exist: no spot-check has been published or logged, and it should be run before that study's findings are treated as equivalent in reliability to the other three.
Cross-study comparisons are across different audiences, different platforms and different central claims at once. The series can suggest that framing changes reception; it cannot isolate framing from audience. The METR report study's before/after design is the exception that can say something about change over time within one platform, but it compares one subreddit against eight, not a matched panel, so its exact percentages should be read as directional too.
Where an artefact could not be watched directly, its argument was reconstructed from the transcript or from what commenters echoed, and the study says so.
The moves recommended in the articles' closing sections are largely consistent with established communication research, including the fear-appeals literature, which finds that threat messaging helps or backfires conditional on perceived efficacy rather than being reliably one or the other. The contribution of these studies is not those principles but the specific, comparable evidence of how AI-risk framings are received, which the field currently lacks.
Data and contact
Each study's codebook and coded dataset are published with this page, and browsable comment by comment at /data/. Errors, re-codings and methodological criticism are actively wanted: [email protected].
Takedown policy: if a comment's author asks for it to be removed, that row's text is replaced with a removal marker while its codes (stance, frame and the rest) are kept, so published counts and article figures stay reconciled with the dataset rather than drifting out of sync afterwards. This is a policy statement, not a request form; removal requests go to the same address above.
Disclosure
An LLM was used to structure and review this article, with the first and final edits made by a human.