The History and Labor of Web Scraping

Abstract Both the free labor of generating data online and the importance of data work which cleans, labels, and verifies these data are required for producing artificial intelligence (AI) systems. Between these two essential moments in AI production, we identify web scraping as the intercalary process of extracting data from the internet at scale. In this chapter, we highlight the essential contributions of web scraping to the production of machine learning training datasets and show that scraping is a distinct category of software labor whose understanding is important for the political economy of AI. The chapter characterizes the labor involved in both the production of scraping software and the labor required to conduct scraping operations. We provide a critical historical assessment of scraping dating back to the 1990s, tracing the development of scraping from a technical requirement for indexing the web in early communication software, to its commercialization and contemporary use in the development and production of large language models. We further describe the conditions of the labor of scraping through the analysis of contemporary documentation: industry-specific blogs, discussion boards, market reports, and conferences. We argue that the economic utility of data is dependent on its transformation from web data into datasets, and that this transformation is achieved through the labor of scraping. We suggest that the increased legal and technical barriers to the practice frame the current web scraping labor ecosystem, continuously demanding new kinds of human expertise.

Authors

Institutions

Publication Details

Journal
International Labor and Working-Class History
Published
2026-10-09
DOI
https://doi.org/10.1017/s0147547926100350
Primary Topic
Digital Economy and Work Transformation
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

The History and Labor of Web Scraping

James Steinhoff, Sebastião Quelhas Freire
International Labor and Working-Class History
Digital Economy and Work Transformation
article

The History and Labor of Web Scraping

James Steinhoff, Sebastião Quelhas Freire
article en

Abstract

Abstract Both the free labor of generating data online and the importance of data work which cleans, labels, and verifies these data are required for producing artificial intelligence (AI) systems. Between these two essential moments in AI production, we identify web scraping as the intercalary process of extracting data from the internet at scale. In this chapter, we highlight the essential contributions of web scraping to the production of machine learning training datasets and show that scraping is a distinct category of software labor whose understanding is important for the political economy of AI. The chapter characterizes the labor involved in both the production of scraping software and the labor required to conduct scraping operations. We provide a critical historical assessment of scraping dating back to the 1990s, tracing the development of scraping from a technical requirement for indexing the web in early communication software, to its commercialization and contemporary use in the development and production of large language models. We further describe the conditions of the labor of scraping through the analysis of contemporary documentation: industry-specific blogs, discussion boards, market reports, and conferences. We argue that the economic utility of data is dependent on its transformation from web data into datasets, and that this transformation is achieved through the labor of scraping. We suggest that the increased legal and technical barriers to the practice frame the current web scraping labor ecosystem, continuously demanding new kinds of human expertise.

International Labor and Working-Class History
University College Dublin (IE)
Openalex Percentile: Top 5%
Digital Economy and Work Transformation
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

The History and Labor of Web Scraping — James Steinhoff, Sebastião Quelhas Freire · International Labor and Working-Class History (2026) | TGRS Research Map | TGRS