Caching¶
How the platform caches, what each tier is for, and how stale each cache can be
across pods. The deployment this is written for: 2+ badger-web replicas behind
Cloudflare with no session affinity, badger-rest, badger-worker and
sync-worker, one shared Valkey (Redis), RabbitMQ.
The three cache managers¶
All three are defined in common/.../cache/CacheConfig.java.
| Manager | Where the data lives | Use it for |
|---|---|---|
redisCacheManager (@Primary, the default) |
Shared Valkey | Anything every pod must agree on the moment it changes. One eviction is seen by all pods. |
tieredCacheManager |
5 s per-pod Caffeine L1 in front of the Redis L2 | Read-on-every-request data that tolerates a few seconds of staleness on other pods. |
inMemoryCacheManager |
Per-pod Caffeine only | Hot data with its own cross-pod invalidation (message source, stereotypes). |
Pick with @Cacheable(cacheNames = "...", cacheManager = "tieredCacheManager");
omit cacheManager for Redis.
Redis key prefixes and tenants¶
Caches configured in CacheConfig.redisCacheConfigurations() with the
siteSpecific prefix — and every cache not listed there (the 5-minute
default) — are prefixed with the current tenant's siteId
(TenantContext), falling back to GENERIC_CACHE when there is no tenant.
Caches whose key already identifies the data globally (siteContexts-*,
siteConfiguration, siteEntitlement, ...) are not prefixed.
A cache of something every app can write must not be tenant-prefixed if the
writer may run without a TenantContext (the worker) — its eviction would land
in GENERIC_CACHE and miss the storefront's entry.
Redis failure is non-fatal¶
ResilientCacheErrorHandler (installed by CacheResilienceConfig, a
CachingConfigurer) applies to every annotation-driven cache:
- a failed get is a miss — the annotated method runs and reads Mongo;
- a failed put / evict / clear is logged and skipped.
Warnings are rate limited to one per operation and cache per minute, with a
count of what was suppressed; stack traces are at DEBUG. The Redis command
timeout defaults to 1 s (REDIS_DB_TIMEOUT). The connect timeout stays at
Lettuce's 10 s (REDIS_DB_CONNECT_TIMEOUT): Lettuce opens one shared connection
and reconnects in the background, so a short connect timeout protects no request
thread. It did make a CPU-starved pod's first connection fail right after
startup (seen on the dev worker at 2 s).
Circuit breaker¶
A 1 s timeout alone still makes a hung Valkey expensive: every cached lookup
waits the full second, runs the method, then waits again on the put, so a page
with a dozen cache calls takes tens of seconds. CircuitBreakingRedisCacheWriter
wraps the RedisCacheWriter underneath every Redis cache (including the L2 of
the tiered caches) with a breaker, RedisCacheCircuitBreaker:
| State | What cache calls do |
|---|---|
| Closed | Call Redis. failure-threshold consecutive failures open the circuit. |
Open for open-duration |
Don't touch Redis. Gets are misses (Mongo answers), puts are skipped, evictions and clears are queued. |
| Half-open | One caller is the trial: it replays the queued evictions, then makes its own call. Success closes the circuit; failure re-opens it for another open-duration. |
badger:
cache:
redis:
circuit-breaker:
failure-threshold: 3 # consecutive Redis failures before opening
open-duration: 10s # how long to bypass Redis before a trial call
Only a reply counts as success. Spring Data Redis routes put, evict and
clear through its asynchronous writer and discards the future, so those three
return in microseconds — and report success — against a Redis that is hung or
unreachable. Two consequences the writer has to work around, both of which
silently disabled the breaker before they were fixed:
- a
putis reported as unconfirmed and must not clear the failure count. A@Cacheablemiss is a failing get followed by one of those "successful" puts, so counting the put as a success reset the count on every request and the circuit never reached the threshold — every read kept paying the full 1 s timeout for the whole outage. evictandclearare issued throughevictIfPresentandinvalidate, which run the sameDEL/SCAN+DELbut return a value and therefore wait for the reply. Without that, an eviction during an outage raised no exception, so nothing was ever queued for replay andpending_invalidationsstayed at 0.
If you add a writer method, check whether it waits for a reply before feeding its result to the breaker.
Evictions are never just dropped. An eviction that can't reach Redis
(short-circuited, or failed while the circuit was still closed) is queued and
replayed as soon as Redis answers, before anything is read back from it. An
admin edit made during an outage doesn't leave a stale entry for the cache's
whole TTL. The queue holds 10,000 entries; beyond that, the affected caches are
cleared entirely on recovery. Queued evictions live in the pod's memory, so a
pod restarted mid-outage loses its queue, and the TTL is the backstop. That's
why no Redis cache lives longer than a day (themeContext and taxonomy).
Metrics: badger_cache_redis_circuit_open (1 while open or half-open — alert on
it), badger_cache_redis_circuit_short_circuited_total,
badger_cache_redis_pending_invalidations.
The breaker covers Spring caches only. Sessions, OAuth authorizations, the performance monitor and idempotency markers use Redis directly and still pay the command timeout while it is down.
allEntries evictions clear by SCAN, not KEYS, so they don't block Valkey.
Serialisation¶
Every Redis cache (and the copy-on-read L1, below) uses
CacheSerializers.redisValueSerializer(): Spring's
GenericJackson2JsonRedisSerializer plus JavaTimeModule and
FAIL_ON_UNKNOWN_PROPERTIES = false. A type is only safe to cache if it
round-trips field for field through it — if the cached copy is later saved
back to Mongo, any field lost in JSON is wiped in the database. Write a
round-trip test when you cache a new domain type. The
usual culprits: @JsonIgnore on a getter drops the whole property, computed
getters without setters, setters with side effects, no usable constructor.
Tiered caches¶
TieredCacheManager wraps a Redis cache with a Caffeine L1 (5 s, 1,000 entries
per cache).
- L1 keys mirror L2 keys. The L1 key includes the Redis key prefix, evaluated per call, so a tenant-prefixed cache stays tenant-scoped in L1 too.
- Copy-on-read. Caches in
CacheConfig.TIERED_COPY_ON_READ_CACHESkeep the serialised bytes in L1 and hand every caller a fresh copy — the same isolation Redis gives. Use it for anything request code mutates in place. Other tiered caches return the same instance to every caller on the pod: never mutate what they return. - Writes evict L1 on the writing pod and L2 everywhere. Other pods' L1 is
cleared by a RabbitMQ broadcast (
CacheInvalidationPublisher→CacheInvalidationListener→TieredCacheManager.clearLocal) where the writer sends one; otherwise it expires within 5 s.
| Tiered cache | Read by | Invalidation | Worst-case staleness on another pod |
|---|---|---|---|
siteContexts-id, siteContexts-domain |
TenantResolutionFilter, every request (copy-on-read) |
SiteContextCacheEvictionListener on any SiteContext save/delete, in any app, + broadcast |
one broadcast hop; 5 s if the broadcast is lost |
siteConfiguration |
SuperSiteConfigService, many times per request |
@CacheEvict on SiteConfigurationRepository |
5 s |
siteEntitlement |
entitlement checks on gated requests | @CacheEvict on SiteEntitlementRepository |
5 s |
themeContext |
template resolution | @CacheEvict on ThemeRepository |
5 s |
oldMenuNodes, menuHash |
navigation | Redis: @CacheEvict on the menu repositories (oldMenuNodes only); L1: broadcast from MenuServiceImpl |
one broadcast hop for L1 — but see the note below on menuHash |
Known gap: menuHash in Redis
Nothing evicts menuHash from Redis — the repositories evict oldMenuNodes
only, and the broadcast clears L1s. After a menu edit the hash can stay stale
for its 30-minute Redis TTL, so browsers keep their cached menu. Not yet
fixed.
Deliberately not tiered: pages/pageCache/collectionService (one or
two lookups per request, larger values, no copy-on-read audit done) and anything
per-user or per-order (baskets) — mutable, read-modify-written data
must see an eviction on every pod immediately, which only a Redis-only cache
gives.
The invalidation broadcast¶
CacheInvalidationPublisher sends a CacheInvalidationEvent to a RabbitMQ
fanout exchange; each pod has its own auto-delete queue bound to it, drained by
CacheInvalidationListener. The publisher serialises with the shared
messageConverter bean from AMQPConfig, which writes the payload's class name
into the __TypeId__ header.
Trusted packages: why the listener names its payload type
On the consuming side spring-amqp refuses to instantiate any class named in
__TypeId__ unless its package is in the converter's trusted list
(AMQPConfig.messageConverter()). The match is exact package equality, not
a prefix, and * (trust all) is not acceptable on a broker shared between
environments.
That check is only reached when the converter has no inferred payload type,
i.e. for a listener that takes a raw Message (as CacheInvalidationListener
does, because its queue is created at runtime) or a bean with several
@RabbitHandler methods. A @RabbitListener method with one concrete bean
parameter supplies the type itself and never hits it.
CacheInvalidationEvent lives in ...settbuilder.cache, which is not
trusted, so until the listener set
messageProperties.setInferredArgumentType(CacheInvalidationEvent.class) every
broadcast was rejected and other pods kept stale caches (invisible on a
single-replica dev cluster, "half applied" in prod). With the type supplied,
__TypeId__ is ignored altogether.
When you add an AMQP-published event type: give the consumer a single
concrete bean parameter; or, if it must take a raw Message, set the inferred
type as above; or, if it needs @RabbitHandler dispatch, put the bean in a
package listed in AMQPConfig.messageConverter() (today
...settbuilder.events.notifications) and add that package explicitly.
AMQPConfigMessageConverterTest and CacheInvalidationListenerTest push real
messages through the configured converter — extend them for a new type.
Specific caches¶
Users are not cached¶
UserRepository.getUserById (read by UserResolutionFilter on every storefront
request with a BDGR_USR cookie) deliberately goes to Mongo every time. The
request's User is saved back whole from many places (checkout details,
remember-me rotation, registration), so a cached copy that missed an eviction
would overwrite an admin's change to roles, a ban or a password. It would also
put password hashes and remember-me tokens in Valkey. A Redis hit costs the same
round trip as the _id read it replaces, so there is nothing to win.
The filter's own per-request write (last-active date and location) is a
targeted $set (UserService.recordActivity), not a whole-document save.
@DBRef documents (per pod)¶
Product.images is an eager @DBRef List<Media>. MongoConfiguration installs
CachingDbRefResolver, which serves media documents from
DbRefDocumentCache for both fetch and bulkFetch (hits from cache, misses
in one $in). Only collections in DbRefDocumentCache.CACHEABLE_COLLECTIONS
are cached.
- Stored as
RawBsonDocumentbytes, decoded per read — callers can't mutate cached values; BSON types are preserved. - 32 MB cap, 5-minute TTL.
- Evicted on every pod:
MediaDbRefCacheEvictionListener(Mongo events) andMediaRepositoryImpl'supdateFirstwrites (the worker's AI tagging) go throughDbRefDocumentInvalidator, which evicts locally and broadcastsCacheInvalidationEvent.DBREF_DOCUMENTS. The 5-minute TTL only bounds a lost broadcast. - Metrics:
cache_gets_total{cache="dbRefDocuments"}.
Message source (per pod)¶
BadgerMessageSource uses three dedicated Caffeine caches in the
inMemoryCacheManager: messageCache (50k), messageFormatCache (20k) and
siteSpecificMessageMap (5k), 1-hour TTL. Editing a site's text calls
reloadSiteMessages, which drops that catalogue's entries locally and
broadcasts CacheInvalidationEvent.SITE_MESSAGES; other pods handle it as an
InvalidateMessagesEvent. The TTL is only a backstop for a lost broadcast.
Adding a cache — checklist¶
- Redis unless you have a reason; tiered only for read-hot, staleness-tolerant data; per-pod only with a broadcast or a TTL you can live with.
- Decide the key's tenant scope explicitly (prefix, or a key that includes the tenant, or a globally unique id).
- If the value is ever mutated by callers: Redis, or tiered with copy-on-read.
- Evict from every write path. If apps without
@EnableCaching(worker, sync-worker) can write it, evict from a Mongo mapping-event listener rather than@CacheEvict. - Round-trip test the value type through
CacheSerializers. - Never cache
nullfor something that is about to be created. - Don't cache an entity that callers load, modify and save back whole: a stale
copy becomes a lost update (this is why
Userisn't cached). - Keep the TTL at a day or less. It is the backstop for a lost eviction.
- If you broadcast a new payload type, mind the trusted-packages rule (see the invalidation broadcast).