Ask Paul: May 31 (Premium)

Happy Friday! It's been a busy few weeks, but I'm not surprised that many of you are thinking about the same things as I am. So let's jump in and kick off the weekend a bit early with some terrific questions.
AI PC, more like DOA PC
A couple of related questions about AI PCs…

madthinus asks:

Between Ai PC and Copilot+ PC arriving less than a year apart, both proclaimed to be the next thing, is the 40 TOPs arbitrary like 8gen Chips was for Windows 11 or are they waking up to the reality of on Chip AI?

40 TOPS doesn't feel arbitrary to me. And while this "session" isn't as on-point as his Build 2023 appearance, Stevie Bathiche gave a very short talk at Build 2024 that neatly explains Microsoft's focus on the NPU (among other things).

We all know that GPUs can handle graphical workloads more efficiently than CPUs, and that using one can thus provide this magical combination of better performance and battery life. But the way Stevie describes these different processors is fascinating to me. CPUs are optimal for scalars, like numbers. GPUs are optimal for vectors, like arrays of numbers. But NPUs are optimal for tensors, like arrays of arrays.

"NPUs ... are basically tensor accelerators," he says. "NPUs are purpose built to handle tensors. All the way from the entire Silicon architecture to the software stack is all about managing this data type. It gives us a tremendous amount of efficiency as a result."

I try to explain things in simple terms and AI is particularly challenging because there's a lot of new and unfamiliar terminology, and because it's, well, complicated computer science. So I typically boil down the NPU to something like "hardware accelerated AI." But Stevie gives a great real-world example of how and why the NPU is so important for generative AI tasks by visually showing how much more efficient it is than a CPU or GPU. It's dramatic.

But the basics still apply. In the same way that an GPU can offload tasks from the CPU and improve performance and battery life, the NPU can perform certain tasks in the background so efficiently that the net impact to the system is negligible. To date, the big example of this type of task is Windows Studio Effects: You're on a video call for 30, 60 minutes, whatever, and it's in the background doing its thing, blurring or changing your surroundings, killing background noise, and so on, and there's basically zero impact on system performance or battery life. The CPU and GPU could both handle that work, but harming performance and battery life. GPUs can have higher TOPS scores than NPUs, but they are not even close on efficiency. Which is why NPUs are so important on Edge devices like laptops.

Windows Studio Effects works fine on a 7 TOPs AI PC. But to do the work that Microsoft showed off recently, like Recall, which runs in the background and uses multiple on-device SLMs simultaneously, raw performance isn't enough. Efficiency matters.

"The way we built Recall...

Gain unlimited access to Premium articles.

With technology shaping our everyday lives, how could we not dig deeper?

Thurrott Premium delivers an honest and thorough perspective about the technologies we use and rely on everyday. Discover deeper content as a Premium member.

Tagged with

Share post

Thurrott