What to Expect From Pure CPU-Only Local Inference and Hardware Limits
Understanding local AI speed limits on CPUs and RAM, and how to calculate tokens per second yourself.

Stock photo for illustration only, not from the actual event
- CPU-based local AI speed is strictly limited by memory bandwidth, not core count.
- Performance ceilings can be calculated by dividing usable memory bandwidth by model file size.
- Long prompt prefill phases on CPUs take significantly longer than token generation speeds.
Nobody publishes precise tokens-per-second figures for your specific model running on your RAM, and any page providing one has either measured a different machine or made it up entirely. However, you can establish a reliable performance ceiling using just two numbers that take five minutes to look up.
A request consists of a prefill phase, where the entire prompt passes through the model in a single parallel run, and a decode phase, where each output token requires its own pass. They face entirely different bottlenecks, and conflating them is why advice on CPU inference is frequently useless.
Decode is what people are referring to when they complain that a local model feels sluggish. To produce a single token, the CPU must read every active weight in the model out of RAM and into the cache, perform minor arithmetic on each, and move forward. With a batch size of one, there is no way to amortize that read cost—one pass over all weights yields precisely one token. Consequently, decode is constrained by memory bandwidth while arithmetic units sit idle waiting for data.
If every token requires reading the weights once, the governing rule is:
- Tokens/second <= Usable memory bandwidth (bytes/s) / Weight bytes
The numerator derives from the DDR standard your hardware utilizes. A DDR4 or DDR5 channel spans 64 bits wide, moving 8 bytes per transfer, with the nominal name indicating the transfer rate:
- DDR4-3200, one channel: 3200e6 transfers/s x 8 B = 25.6 GB/s
- DDR4-3200, dual channel: 51.2 GB/s
- DDR5-5600, one DIMM: 5600e6 transfers/s x 8 B = 44.8 GB/s
- DDR5-5600, two DIMMs: 89.6 GB/s
The denominator represents the size of the file on disk, viewable directly from the repository listing prior to downloading. This size serves as a reliable proxy for bytes streamed per token in dense models because GGUF files consist almost entirely of tensor data.

Stock photo for illustration only, not from the actual event
Apple publishes unified memory bandwidth specifications directly rather than requiring derivation. During Apple's October 2024 announcement for the M4 Pro and M4 Max, the figures provided were 273 GB/s for the M4 Pro and 546 GB/s for the M4 Max. This single hardware advantage explains why Apple Silicon excels at local inference—not because of superior CPU compute, but due to the memory bus.
Grasping memory bandwidth limitations allows for realistic hardware expectations. Upgrading CPU core counts will not accelerate large model inference if the memory bandwidth remains the bottleneck, which highlights why unified memory architectures perform exceptionally well for local AI workloads.
Dividing these bandwidth figures by published file sizes from Hugging Face GGUF repositories (verified on 2026-08-11) yields clear performance ceilings. For instance, the Phi-3.5-mini-instruct model at 2.39 GB achieves up to 21.4 tokens per second on DDR4-3200 and up to 228.1 on an M4 Max, whereas the larger Meta-Llama-3.1-70B-Instr model at 42.52 GB tops out at 1.2 tokens per second on DDR4-3200 and 12.8 on an M4 Max.
Reviewing the bottom performance bracket reveals the most practical takeaway. A 70B model quantized at Q4_K_M on a dual-channel DDR5 desktop cannot exceed roughly two tokens per second regardless of attached CPU power, because 42.5 GB must traverse the memory bus for every single generated token. No core count configuration overcomes this hardware constraint, which explains why users often mistakenly conclude that local inference is impractical when the mismatch lies in bandwidth ratios.
Conversely, the prefill phase is not bandwidth-bound because the entire prompt is available simultaneously and every weight read is reused across all prompt tokens. Instead, it is constrained by arithmetic operations, scaling at roughly two floating-point operations per parameter per token.
The practical implication is that pure CPU inference degrades far more drastically with prompt length than with output length, reversing the intuition developed from hosted APIs. Short prompts and compact models run comfortably, whereas long documents introduce noticeable initial delays that no quantization level can fix, since prefill costs scale with parameters and tokens rather than bits per weight.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment