Flattened Conversations: Reply-Structure Loss in YouTube Comment Data
Research on YouTube comment sections routinely reconstructs who replied to whom. That edge list comes from one of two collection methods, and they do not agree. This report measures the disagreement on identical material. Two videos were collected both through the official YouTube Data API v3 and through interface-level extraction (yt-dlp). In the first, 3,032 of 6,248 replies are assigned different parents by the two methods; in the second, 17,704 of 41,885. Every divergent case has the same shape. Where the interface export records that C answered B, the API records that C answered A, the comment at the top of the thread. For any analysis conditioning on the previous turn, C's prior turn changes identity. This is not an API defect. The API documents a two-level comment model and behaves accordingly: across 43,896 replies retrieved via comments.list, parentId points to a top-level comment in 100.0% of cases, without exception. The flattening follows from the specification. What has not been measured, to our knowledge, is its size against interface-level data on the same comments. Across four corpora (n = 267,505) the share of replies whose parent is itself a reply runs from 10.8% to 65.5%, and is highest in precisely the conversational corpora that interaction research selects for. What the interface method recovers, and where it stops. The structure it returns is richer than the API's and is itself bounded: every corpus here has a maximum depth of exactly 3. That bound was reported as an observation in version 3.0 and is now known to be a platform nesting cap, so the deepest layer of every depth distribution in this report is a censored layer (Sedlmair 2026d). Restricting the comparison to edges below the cap, the usable gain over the API falls from 42–49% of replies to 8.6–24.3%. That reduction is substantial and is stated here rather than left to §6. What remains is not residue: depth-2 edges agree with the platform-inserted address marker in 99.97% of cases, and the API supplies none of them. A blind coherence validation on 50 hand-checked parent–child pairs supports, without proving, that the interface-level parent is the comment actually responded to: 36 of 40 decidable pairs were judged coherent (90.0%, 95% CI 76.9–96.0). The test was designed before the cap was known and was not stratified by depth, so it provides no evidence about the censored layer. A depth-stratified re-run is the single most useful thing that could still be done to this report.
Authors
- Max Sedlmair (ORCID: https://orcid.org/0009-0007-2471-5595)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.22842406
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- preprint