Fundamentals of Redis Caching: Why and How It Works
Explore computer memory basics, concurrency challenges in production, and why modern distributed architectures rely on Redis caching.

Stock photo for illustration only, not from the actual event
- CPU executes instructions, RAM acts as fast temporary memory, and SSD provides permanent storage.
- Local hash maps work on a laptop but fail under concurrency in production environments.
- Building a manual cache requires complex handling of Mutex, LRU eviction, and TTL expiration.
- Separate RAM across multiple EC2 servers leads to cache inconsistency and duplicate storage.
Understanding caching starts with the fundamentals of computer architecture. The CPU (Central Processing Unit) serves as the core processor that executes instructions and runs code computations. Meanwhile, RAM (Random Access Memory) operates as fast, temporary storage where running programs keep active data like variables and arrays, and SSD (Solid State Drive) provides persistent storage for long-term files and databases. A key distinction is that RAM is volatile and loses data on power loss, whereas SSD is persistent.
Latency varies significantly across these hardware layers. Reading from RAM takes about 100 nanoseconds, reading from SSD takes roughly 100 microseconds, a network call to Redis takes between 0.5 and 1 millisecond, and database queries range from 5 to 50 milliseconds, with 1 millisecond equating to 1,000 microseconds or 1,000,000 nanoseconds. Standard hardware setups on laptops, EC2 instances, and database servers all share these fundamental components.

Stock photo for illustration only, not from the actual event
While local data structures like hash maps or dictionaries work seamlessly on a single laptop request, production environments differ vastly due to concurrency, which involves multiple tasks running at overlapping times, such as 100 simultaneous user requests. Servers handle this traffic using multiple threads or lightweight goroutines in Go, which can trigger race conditions when multiple threads access the same underlying data simultaneously.
In Go, writing concurrently to a map from multiple goroutines triggers a fatal crash. Developers typically apply a Mutex (Mutual Exclusion lock) to restrict access to protected code sections sequentially. Furthermore, maps have no built-in size limits, meaning continuous additions will eventually exhaust RAM and trigger an Out Of Memory (OOM) error where the operating system terminates the process, or cause memory leaks.
"A fast storage layer that keeps copies of frequently used data so you don't have to fetch it from the slow source again."
Cache Fundamentals
Production caching requires automatic eviction policies like Least Recently Used (LRU) to remove stale data, alongside Time To Live (TTL) timers to expire keys automatically. Handcrafting a reliable cache implementation involving Mutex locks, LRU eviction, and TTL management is notoriously difficult and prone to bugs.

Stock photo for illustration only, not from the actual event
Another major architectural challenge involves utilizing local EC2 or database RAM as a cache. In typical production setups utilizing load balancers to distribute incoming traffic across multiple application servers, RAM is never shared between machines. This isolation causes two critical issues: first, user data gets cached across multiple instances redundantly, leading to cache misses and inconsistent data states; second, server restarts, deployments, or crashes completely wipe out local memory caches.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment