Valkey and Redis Basics: Caching, Persistence and Not Losing Data

Valkey and Redis Basics: Caching, Persistence and Not Losing Data

An in-memory key value store, used as a cache, a session store, a rate limiter and a queue. Fast because everything is in RAM, and that same fact is the source of every operational problem it has.

The fork, briefly

Redis changed from BSD to a source-available licence in 2024. The community forked the last open version as Valkey, which landed under the Linux Foundation with backing from several large users.

For practical purposes they are the same thing: same protocol, same commands, existing clients work unchanged. Most Linux distributions now ship Valkey, which is why apt install redis may install something called valkey, or nothing at all.

# Debian and Ubuntu
sudo apt install valkey-server

# Fedora
sudo dnf install valkey

# Arch
sudo pacman -S valkey
valkey-cli --version
valkey-cli ping
# PONG

The command names are aliased, so redis-cli usually still works. Our licences guide covers why the fork happened and what source-available means.

First contact

valkey-cli

127.0.0.1:6379> SET greeting "hello"
OK
127.0.0.1:6379> GET greeting
"hello"
127.0.0.1:6379> EXPIRE greeting 60
(integer) 1
127.0.0.1:6379> TTL greeting
(integer) 57
127.0.0.1:6379> DEL greeting
(integer) 1

EXPIRE and TTL are the features that make it a cache rather than a slower database. A key with a TTL removes itself, so stale data has a defined lifetime you do not have to manage.

# the data types, briefly
LPUSH queue "job1"           # list
SADD tags "linux" "kernel"   # set
ZADD scores 100 "alice"      # sorted set
HSET user:1 name "Alice"     # hash
INCR pageviews               # atomic counter

INCR being atomic is why rate limiting works: no read-modify-write race, regardless of how many clients are counting.

Memory limits, the setting that matters most

# /etc/valkey/valkey.conf
maxmemory 2gb
maxmemory-policy allkeys-lru

Without maxmemory, it grows until the kernel kills it. That is the default, and it is the most common way one of these takes a server down with it. Our OOM killer guide covers what that looks like from the outside.

With maxmemory but the default noeviction policy, it stops accepting writes when full and returns errors. Correct for a database, wrong for a cache.

PolicyBehaviour
noevictionReject writes when full, the default
allkeys-lruEvict least recently used, any key
allkeys-lfuEvict least frequently used, any key
volatile-lruEvict least recently used, only keys with a TTL
volatile-ttlEvict the soonest to expire
allkeys-randomEvict at random

allkeys-lru for a cache. allkeys-lfu is often better for a cache with a stable hot set, since it keeps frequently used keys rather than recently used ones.

volatile-* policies are a trap if not every key has a TTL: keys without one are never evicted, so the instance fills with immortal keys and then behaves like noeviction.

# what is it using
valkey-cli INFO memory | grep -E 'used_memory_human|maxmemory_human|maxmemory_policy'

# is it evicting
valkey-cli INFO stats | grep -E 'evicted_keys|keyspace_hits|keyspace_misses'

The hit and miss counters tell you whether the cache is doing anything useful. A hit rate below about 80% usually means the cache is too small or the keys are wrong.

Persistence, and what it guarantees

Two mechanisms, and the difference matters when the machine dies.

RDB snapshots

# save after 900s if 1 key changed, 300s if 10, 60s if 10000
save 900 1
save 300 10
save 60 10000

dbfilename dump.rdb
dir /var/lib/valkey

A point in time snapshot of the whole dataset. Compact, fast to load, and you lose everything since the last one. With the defaults above, that can be up to fifteen minutes on a quiet instance.

The fork used to write the snapshot can also briefly double memory usage on a write-heavy instance, which is worth knowing when sizing the machine.

AOF, the append only file

appendonly yes
appendfsync everysec

Every write command is appended to a log, which is replayed on restart.

appendfsyncDurabilityCost
alwaysLoses nothingSlow, an fsync per write
everysecLoses up to one secondThe sensible default
noUp to 30 secondsFastest

everysec is the right answer for nearly everyone. always makes it as slow as a disk-backed database, which defeats the purpose.

# both, which is the common production choice
appendonly yes
appendfsync everysec
save 900 1

AOF for recovery, RDB for fast restarts and backups. They coexist.

# force a snapshot without blocking
valkey-cli BGSAVE

# rewrite a bloated AOF
valkey-cli BGREWRITEAOF

# when was the last successful save
valkey-cli LASTSAVE
valkey-cli INFO persistence | grep -E 'rdb_last_bgsave_status|aof_last_write_status'

Those status fields are worth monitoring. A failing background save, usually because the disk is full, is silent until you restart and discover the data is old.

For a pure cache, turn both off. Persistence costs I/O for data you do not need to survive a restart:

save ""
appendonly no

Securing it

Redis was designed for trusted networks, and the consequences of that assumption are well documented. Unprotected instances are scanned for constantly, and some commands let an attacker write files to disk, which turns exposure into remote code execution.

# bind to localhost only
bind 127.0.0.1 -::1

# a password, and make it long
requirepass <64 random characters>

# protected mode, on by default, keep it
protected-mode yes
openssl rand -base64 48
valkey-cli
127.0.0.1:6379> AUTH <password>
OK

# on the command line the password is visible in ps, so prefer the variable
REDISCLI_AUTH=<password> valkey-cli

Rename or disable the dangerous commands:

rename-command FLUSHALL ""
rename-command FLUSHDB ""
rename-command CONFIG ""
rename-command DEBUG ""

CONFIG SET dir plus SAVE is the primitive that lets an unauthenticated attacker write arbitrary files, and it is how instances get compromised. Disabling CONFIG breaks some management tooling, which is a reasonable trade on an internet-adjacent host.

For proper access control, use ACLs rather than one shared password:

valkey-cli ACL SETUSER appuser on '>password' '~app:*' +get +set +del +expire
valkey-cli ACL LIST

That user can only touch keys matching app:* and can only run those four commands. Far better than a single password shared by everything.

# firewall, if it must be reachable
sudo ufw allow from 10.0.0.10 to any port 6379

TLS is supported and requires certificates on both ends:

tls-port 6380
port 0
tls-cert-file /etc/valkey/server.crt
tls-key-file /etc/valkey/server.key
tls-ca-cert-file /etc/valkey/ca.crt

Across anything you do not control, a WireGuard tunnel is simpler than getting TLS right, and our security practices guide covers the wider posture.

Using it as a queue

# the old pattern, with a flaw
LPUSH jobs "task1"
BRPOP jobs 0

There is no acknowledgement. A consumer that pops a job and then crashes has lost it, and nothing knows.

Streams fix that:

# produce
XADD jobs '*' task "resize-image" id 42

# consumer group, created once
XGROUP CREATE jobs workers 0

# consume
XREADGROUP GROUP workers worker1 COUNT 1 BLOCK 0 STREAMS jobs '>'

# acknowledge when done
XACK jobs workers 1695200000000-0

# what was delivered and never acknowledged
XPENDING jobs workers

XPENDING is the point: a job delivered but not acknowledged is visible and can be reclaimed by another worker. For anything where losing a job matters, this is the minimum, and a purpose-built broker is worth considering above a certain complexity.

Keeping an eye on it

# overview
valkey-cli INFO
valkey-cli INFO clients
valkey-cli --stat

# slow commands
valkey-cli SLOWLOG GET 10
valkey-cli CONFIG SET slowlog-log-slower-than 10000

# watch commands live, expensive, never leave running
valkey-cli MONITOR

# what is big
valkey-cli --bigkeys
valkey-cli --memkeys

--bigkeys is the first thing to run when memory is higher than expected. It finds the one enormous list somebody appended to and never trimmed.

Never use KEYS * on a live instance. It is O(n) and blocks the single-threaded server while it scans. SCAN iterates in chunks instead:

valkey-cli --scan --pattern 'session:*' | head

That single-threaded design is worth internalising: one slow command blocks everything. A KEYS * on a million keys is an outage.

Our monitoring guide covers graphing the INFO output, and hit rate, evictions and memory are the three to watch.

Running it in a container

# compose.yaml
services:
  valkey:
    image: valkey/valkey:8-alpine
    command: >
      valkey-server
      --maxmemory 512mb
      --maxmemory-policy allkeys-lru
      --appendonly yes
      --requirepass ${VALKEY_PASSWORD}
    volumes:
      - valkey-data:/data
    ports:
      - "127.0.0.1:6379:6379"
    restart: unless-stopped

volumes:
  valkey-data:

127.0.0.1:6379:6379 rather than 6379:6379. The short form binds to every interface, and on a host where the firewall rules are managed by the container runtime that can mean exposing it to the internet without ever intending to. Our Docker Compose guide and container networking guide cover why that happens.

Passing config as command arguments rather than mounting a file is convenient for a few settings and becomes unreadable past about six. Mount a config file at that point.

Frequently Asked Questions

What is the difference between Valkey and Redis?

Valkey is a fork of Redis created when Redis changed to a source-available licence, and it is maintained under the Linux Foundation. It began as a drop-in replacement with the same protocol and commands, so existing clients work unchanged. Most Linux distributions now ship Valkey in place of Redis for that licensing reason.

Does Redis lose data if the server restarts?

It depends on your persistence configuration. With snapshots only, you lose everything since the last snapshot, which can be minutes. With append only file persistence and per-second flushing, you lose at most a second. With no persistence, you lose everything, which is correct for a pure cache.

What is the difference between RDB and AOF persistence?

RDB writes a point in time snapshot of the whole dataset periodically, which is compact and fast to load but loses everything since the last one. AOF appends every write command to a log, which loses much less and produces a larger file and slower restarts. Running both is a common and reasonable choice.

How do I stop it consuming all the memory on the machine?

Set maxmemory to a specific limit and choose a maxmemory-policy. Without both, it grows until the kernel kills it, and with a limit but the default noeviction policy it starts rejecting writes instead. For a cache, allkeys-lru evicts the least recently used keys and is what you want.

Is it safe to expose the port to a network?

Not without authentication and a firewall. Historically it bound to all interfaces with no password, and unprotected instances are found and abused constantly, since some commands let an attacker write files. Bind to localhost, require a password, rename or disable dangerous commands, and use TLS across any network you do not control.

Should I use it as a message queue?

Streams are a reasonable queue with consumer groups and acknowledgements. The older list based pattern with push and pop has no acknowledgement, so a consumer that crashes mid-job loses it. For anything where losing a job matters, use streams or a purpose built broker.