Skip to content

Caching

How the platform caches, what each tier is for, and how stale each cache can be across pods. The deployment this is written for: 2+ badger-web replicas behind Cloudflare with no session affinity, badger-rest, badger-worker and sync-worker, one shared Valkey (Redis), RabbitMQ.

The three cache managers

All three are defined in common/.../cache/CacheConfig.java.

Manager Where the data lives Use it for
redisCacheManager (@Primary, the default) Shared Valkey Anything every pod must agree on the moment it changes. One eviction is seen by all pods.
tieredCacheManager 5 s per-pod Caffeine L1 in front of the Redis L2 Read-on-every-request data that tolerates a few seconds of staleness on other pods.
inMemoryCacheManager Per-pod Caffeine only Hot data with its own cross-pod invalidation (message source, stereotypes).

Pick with @Cacheable(cacheNames = "...", cacheManager = "tieredCacheManager"); omit cacheManager for Redis.

Redis key prefixes and tenants

Caches configured in CacheConfig.redisCacheConfigurations() with the siteSpecific prefix — and every cache not listed there (the 5-minute default) — are prefixed with the current tenant's siteId (TenantContext), falling back to GENERIC_CACHE when there is no tenant. Caches whose key already identifies the data globally (siteContexts-*, siteConfiguration, siteEntitlement, ...) are not prefixed.

A cache of something every app can write must not be tenant-prefixed if the writer may run without a TenantContext (the worker) — its eviction would land in GENERIC_CACHE and miss the storefront's entry.

Redis failure is non-fatal

ResilientCacheErrorHandler (installed by CacheResilienceConfig, a CachingConfigurer) applies to every annotation-driven cache:

  • a failed get is a miss — the annotated method runs and reads Mongo;
  • a failed put / evict / clear is logged and skipped.

Warnings are rate limited to one per operation and cache per minute, with a count of what was suppressed; stack traces are at DEBUG. The Redis command timeout defaults to 1 s (REDIS_DB_TIMEOUT). The connect timeout stays at Lettuce's 10 s (REDIS_DB_CONNECT_TIMEOUT): Lettuce opens one shared connection and reconnects in the background, so a short connect timeout protects no request thread. It did make a CPU-starved pod's first connection fail right after startup (seen on the dev worker at 2 s).

Circuit breaker

A 1 s timeout alone still makes a hung Valkey expensive: every cached lookup waits the full second, runs the method, then waits again on the put, so a page with a dozen cache calls takes tens of seconds. CircuitBreakingRedisCacheWriter wraps the RedisCacheWriter underneath every Redis cache (including the L2 of the tiered caches) with a breaker, RedisCacheCircuitBreaker:

State What cache calls do
Closed Call Redis. failure-threshold consecutive failures open the circuit.
Open for open-duration Don't touch Redis. Gets are misses (Mongo answers), puts are skipped, evictions and clears are queued.
Half-open One caller is the trial: it replays the queued evictions, then makes its own call. Success closes the circuit; failure re-opens it for another open-duration.
badger:
  cache:
    redis:
      circuit-breaker:
        failure-threshold: 3   # consecutive Redis failures before opening
        open-duration: 10s     # how long to bypass Redis before a trial call

Only a reply counts as success. Spring Data Redis routes put, evict and clear through its asynchronous writer and discards the future, so those three return in microseconds — and report success — against a Redis that is hung or unreachable. Two consequences the writer has to work around, both of which silently disabled the breaker before they were fixed:

  • a put is reported as unconfirmed and must not clear the failure count. A @Cacheable miss is a failing get followed by one of those "successful" puts, so counting the put as a success reset the count on every request and the circuit never reached the threshold — every read kept paying the full 1 s timeout for the whole outage.
  • evict and clear are issued through evictIfPresent and invalidate, which run the same DEL / SCAN+DEL but return a value and therefore wait for the reply. Without that, an eviction during an outage raised no exception, so nothing was ever queued for replay and pending_invalidations stayed at 0.

If you add a writer method, check whether it waits for a reply before feeding its result to the breaker.

Evictions are never just dropped. An eviction that can't reach Redis (short-circuited, or failed while the circuit was still closed) is queued and replayed as soon as Redis answers, before anything is read back from it. An admin edit made during an outage doesn't leave a stale entry for the cache's whole TTL. The queue holds 10,000 entries; beyond that, the affected caches are cleared entirely on recovery. Queued evictions live in the pod's memory, so a pod restarted mid-outage loses its queue, and the TTL is the backstop. That's why no Redis cache lives longer than a day (themeContext and taxonomy).

Metrics: badger_cache_redis_circuit_open (1 while open or half-open — alert on it), badger_cache_redis_circuit_short_circuited_total, badger_cache_redis_pending_invalidations.

The breaker covers Spring caches only. Sessions, OAuth authorizations, the performance monitor and idempotency markers use Redis directly and still pay the command timeout while it is down.

allEntries evictions clear by SCAN, not KEYS, so they don't block Valkey.

Serialisation

Every Redis cache (and the copy-on-read L1, below) uses CacheSerializers.redisValueSerializer(): Spring's GenericJackson2JsonRedisSerializer plus JavaTimeModule and FAIL_ON_UNKNOWN_PROPERTIES = false. A type is only safe to cache if it round-trips field for field through it — if the cached copy is later saved back to Mongo, any field lost in JSON is wiped in the database. Write a round-trip test when you cache a new domain type. The usual culprits: @JsonIgnore on a getter drops the whole property, computed getters without setters, setters with side effects, no usable constructor.

Tiered caches

TieredCacheManager wraps a Redis cache with a Caffeine L1 (5 s, 1,000 entries per cache).

  • L1 keys mirror L2 keys. The L1 key includes the Redis key prefix, evaluated per call, so a tenant-prefixed cache stays tenant-scoped in L1 too.
  • Copy-on-read. Caches in CacheConfig.TIERED_COPY_ON_READ_CACHES keep the serialised bytes in L1 and hand every caller a fresh copy — the same isolation Redis gives. Use it for anything request code mutates in place. Other tiered caches return the same instance to every caller on the pod: never mutate what they return.
  • Writes evict L1 on the writing pod and L2 everywhere. Other pods' L1 is cleared by a RabbitMQ broadcast (CacheInvalidationPublisher → CacheInvalidationListener → TieredCacheManager.clearLocal) where the writer sends one; otherwise it expires within 5 s.
Tiered cache Read by Invalidation Worst-case staleness on another pod
siteContexts-id, siteContexts-domain TenantResolutionFilter, every request (copy-on-read) SiteContextCacheEvictionListener on any SiteContext save/delete, in any app, + broadcast one broadcast hop; 5 s if the broadcast is lost
siteConfiguration SuperSiteConfigService, many times per request @CacheEvict on SiteConfigurationRepository 5 s
siteEntitlement entitlement checks on gated requests @CacheEvict on SiteEntitlementRepository 5 s
themeContext template resolution @CacheEvict on ThemeRepository 5 s
oldMenuNodes, menuHash navigation Redis: @CacheEvict on the menu repositories (oldMenuNodes only); L1: broadcast from MenuServiceImpl one broadcast hop for L1 — but see the note below on menuHash

Known gap: menuHash in Redis

Nothing evicts menuHash from Redis — the repositories evict oldMenuNodes only, and the broadcast clears L1s. After a menu edit the hash can stay stale for its 30-minute Redis TTL, so browsers keep their cached menu. Not yet fixed.

Deliberately not tiered: pages/pageCache/collectionService (one or two lookups per request, larger values, no copy-on-read audit done) and anything per-user or per-order (baskets) — mutable, read-modify-written data must see an eviction on every pod immediately, which only a Redis-only cache gives.

The invalidation broadcast

CacheInvalidationPublisher sends a CacheInvalidationEvent to a RabbitMQ fanout exchange; each pod has its own auto-delete queue bound to it, drained by CacheInvalidationListener. The publisher serialises with the shared messageConverter bean from AMQPConfig, which writes the payload's class name into the __TypeId__ header.

Trusted packages: why the listener names its payload type

On the consuming side spring-amqp refuses to instantiate any class named in __TypeId__ unless its package is in the converter's trusted list (AMQPConfig.messageConverter()). The match is exact package equality, not a prefix, and * (trust all) is not acceptable on a broker shared between environments.

That check is only reached when the converter has no inferred payload type, i.e. for a listener that takes a raw Message (as CacheInvalidationListener does, because its queue is created at runtime) or a bean with several @RabbitHandler methods. A @RabbitListener method with one concrete bean parameter supplies the type itself and never hits it.

CacheInvalidationEvent lives in ...settbuilder.cache, which is not trusted, so until the listener set messageProperties.setInferredArgumentType(CacheInvalidationEvent.class) every broadcast was rejected and other pods kept stale caches (invisible on a single-replica dev cluster, "half applied" in prod). With the type supplied, __TypeId__ is ignored altogether.

When you add an AMQP-published event type: give the consumer a single concrete bean parameter; or, if it must take a raw Message, set the inferred type as above; or, if it needs @RabbitHandler dispatch, put the bean in a package listed in AMQPConfig.messageConverter() (today ...settbuilder.events.notifications) and add that package explicitly. AMQPConfigMessageConverterTest and CacheInvalidationListenerTest push real messages through the configured converter — extend them for a new type.

Specific caches

Users are not cached

UserRepository.getUserById (read by UserResolutionFilter on every storefront request with a BDGR_USR cookie) deliberately goes to Mongo every time. The request's User is saved back whole from many places (checkout details, remember-me rotation, registration), so a cached copy that missed an eviction would overwrite an admin's change to roles, a ban or a password. It would also put password hashes and remember-me tokens in Valkey. A Redis hit costs the same round trip as the _id read it replaces, so there is nothing to win.

The filter's own per-request write (last-active date and location) is a targeted $set (UserService.recordActivity), not a whole-document save.

@DBRef documents (per pod)

Product.images is an eager @DBRef List<Media>. MongoConfiguration installs CachingDbRefResolver, which serves media documents from DbRefDocumentCache for both fetch and bulkFetch (hits from cache, misses in one $in). Only collections in DbRefDocumentCache.CACHEABLE_COLLECTIONS are cached.

  • Stored as RawBsonDocument bytes, decoded per read — callers can't mutate cached values; BSON types are preserved.
  • 32 MB cap, 5-minute TTL.
  • Evicted on every pod: MediaDbRefCacheEvictionListener (Mongo events) and MediaRepositoryImpl's updateFirst writes (the worker's AI tagging) go through DbRefDocumentInvalidator, which evicts locally and broadcasts CacheInvalidationEvent.DBREF_DOCUMENTS. The 5-minute TTL only bounds a lost broadcast.
  • Metrics: cache_gets_total{cache="dbRefDocuments"}.

Message source (per pod)

BadgerMessageSource uses three dedicated Caffeine caches in the inMemoryCacheManager: messageCache (50k), messageFormatCache (20k) and siteSpecificMessageMap (5k), 1-hour TTL. Editing a site's text calls reloadSiteMessages, which drops that catalogue's entries locally and broadcasts CacheInvalidationEvent.SITE_MESSAGES; other pods handle it as an InvalidateMessagesEvent. The TTL is only a backstop for a lost broadcast.

Adding a cache — checklist

  1. Redis unless you have a reason; tiered only for read-hot, staleness-tolerant data; per-pod only with a broadcast or a TTL you can live with.
  2. Decide the key's tenant scope explicitly (prefix, or a key that includes the tenant, or a globally unique id).
  3. If the value is ever mutated by callers: Redis, or tiered with copy-on-read.
  4. Evict from every write path. If apps without @EnableCaching (worker, sync-worker) can write it, evict from a Mongo mapping-event listener rather than @CacheEvict.
  5. Round-trip test the value type through CacheSerializers.
  6. Never cache null for something that is about to be created.
  7. Don't cache an entity that callers load, modify and save back whole: a stale copy becomes a lost update (this is why User isn't cached).
  8. Keep the TTL at a day or less. It is the backstop for a lost eviction.
  9. If you broadcast a new payload type, mind the trusted-packages rule (see the invalidation broadcast).