Skip to content

Build and Runtime Footprint

Operational knobs that keep images small and services inside their resource limits. Each section says what the default is, why, and how to change it.

Vendored front-end libraries

Third-party browser libraries (jQuery, Bootstrap, Monaco, Summernote, ...) are npm dependencies in web-mvc/package.json. During the Maven build web-mvc/scripts/copy-versioned-deps.js copies them into the web-mvc jar under resources/vendor/{package}/{version}/, served at /vendor/{package}/{version}/.... Templates reach them through the system.css.and.js.cdn.url message, extensions through @vendor/... paths:

<script th:src="#{system.css.and.js.cdn.url('/jquery/3.7.1/dist/jquery.min.js')}"></script>
requiresJS(model, "@vendor/sweetalert2/9.17.2/dist/sweetalert2.all.min.js");

Only the files named in the script's ALLOW_LIST are copied — not whole packages. An npm package also ships sources, dev and ESM builds, sourcemaps, docs and tests; copying everything put ~162 MB into the jar (and into the web and rest images) for about 3 MB of files anything asks for.

Referencing a new vendored file

  1. Add the file, and anything it loads relatively (fonts, images, lazily loaded modules), to that package's globs in ALLOW_LIST. * matches within a path segment and ** across segments. .map files are skipped unless a glob names them.
  2. Build. scripts/check-vendored-refs.js runs straight after the copy and fails the build if a reference is missing. It collects every cdn.url('...'), '/vendor/...' and '@vendor/...' string in every module's src/main, follows url(...) and @import in each referenced stylesheet, and requires a referenced directory to be copied whole.

Run it by hand, in a container, to see every URL it checks:

docker run --rm -v "$PWD":/r -w /r/web-mvc node:22-slim \
  sh -c "npm ci --ignore-scripts --legacy-peer-deps && npm run copy-deps && node scripts/check-vendored-refs.js --list"

Adding a dependency

Every package in package.json must have an ALLOW_LIST entry; the copy fails otherwise, and also fails if a glob matches nothing (typically because an upgrade moved the dist files). So adding a library means choosing what of it the browser needs.

Monaco is the special case: its AMD loader is pointed at min/vs (require.config({ paths: { vs: '/vendor/monaco-editor/<ver>/min/vs' } })) and lazily fetches workers, language modes, CSS and the codicon font from anywhere under it, so the whole min/vs tree is kept.

The checker only sees references in the source tree. A /vendor/... URL typed into site content in the database (custom HTML, page content) is not checked.

Storefront critical path

What a Nova page blocks on before first paint is kept to jQuery, the Bootstrap JS bundle, the Bootstrap CSS and the theme's bdgr.css. Keep it that way:

  • Fonts are self-hosted. Inter comes from @fontsource-variable/inter. Its @font-face rules are inline in nova/fragment/head.html: one variable file per unicode subset, and the Latin file is preloaded. Don't add Google Fonts <link>s: each one is a render-blocking third-party stylesheet plus two extra connections, and it sends visitors' IPs to Google.
  • Icons on every page are inline SVG (nova/fragment/icons.html, generated from Font Awesome 5.15.4's own SVGs). Font Awesome's CSS still loads for content that uses fa-* classes, but via rel="preload" + onload, so it doesn't block render. html.fa-loaded marks when it has applied. A page whose content uses no brand icons never downloads the 77 KB brands font.
  • Page-specific CSS/JS belongs with the template that needs it. The product gallery loads Splide/PhotoSwipe, the legacy bootstrap theme's galleries load Fotorama (bootstrap/fragment/fotorama), and the two animated bootstrap-theme alerts load animate.css. None of them are global.
  • Our own scripts carry data-cookieconsent="ignore". Tenants that run Cookiebot in auto-blocking mode otherwise have every script held back until the consent config loads. Each one is then re-inserted as async, so it's fetched twice and the jQuery → Bootstrap → inline-script order is lost. Our scripts set no cookies; tenant tracking scripts stay consent-gated. Add the attribute to any new first-party <script> in a head/footer fragment. Stripe.js carries it too: its cookies are strictly necessary for fraud prevention, and on live Cookiebot had classed it as "preferences", which blocked checkout for shoppers who declined those cookies.
  • View transitions are cross-document. A navigation loads a fresh page, so don't re-initialise anything on pagereveal. view-transitions.js used to replay every jQuery ready handler there, and every page's JS ran twice.

Collection listing index

The collection listing (ProductService.findBySiteAndCollectionIDs, used by CollectionController and RESTCollectionController) filters on catalogueId + collectionIDs and sorts by attributes.collectionOrdering descending, then _id. It is served by one compound index on product:

{ catalogueId: 1, collectionIDs: 1, "attributes.collectionOrdering": -1, _id: 1 }   name: catalogue_collection_listing_idx

With it a page is read straight off the index in order — no in-memory SORT stage, and only the page's documents are fetched. On 20,000 seeded products (15,000 in the collection), page 11 went from 15,000 keys + 15,000 documents examined to 264 keys + 24 documents.

ProductListingIndexInitializer (in common) creates it on a background thread once the application is ready. It is deliberately not a class-level @CompoundIndex: Variant extends Product and is embedded in Product.variants, and Spring Data would build a second, unused copy under variants.*. createIndex is a no-op when the index already exists, and a failure is logged, not fatal.

Sorting by price, rating, productName (the other validSortByValues) or by ascending collectionOrdering still uses the index for the filter, but sorts the matched set in memory.

Build it by hand first on a large catalogue

On a large product collection, create the index manually before deploying, with the same name and keys, so no pod builds it at startup:

// Must be 0: collectionIDs is already an array, and MongoDB cannot index two
// arrays in one document ("cannot index parallel arrays").
db.product.countDocuments({ "attributes.collectionOrdering": { $type: "array" } })

db.product.createIndex(
  { catalogueId: 1, collectionIDs: 1, "attributes.collectionOrdering": -1, _id: 1 },
  { name: "catalogue_collection_listing_idx" })

Once the index exists, a write that sets attributes.collectionOrdering to an array is rejected. The admin UI stores it as a single value.

Sentry performance tracing

Errors go to Sentry whenever SENTRY_DSN is set. Performance tracing is controlled per deployment by SENTRY_TRACES_SAMPLE_RATE, bound to sentry.traces-sample-rate in each service's application-production.yml:

SENTRY_TRACES_SAMPLE_RATE Effect
unset (default) Tracing off — no transactions are created or sent
0.1 10% of requests are traced
1.0 Every request is traced — avoid in production

common/application.yml is not loaded by the services

Spring Boot reads only the first classpath:application.yml, and every app module (web, rest, worker, auth, sync-worker) ships its own, which shadows the one in common. The sentry: block there (traces-sample-rate: 1.0, send-default-pii: true) has therefore never applied: production has run with tracing off and send-default-pii false. The same goes for the badger.stripe, badger.aws and server.compression settings in that file, which is why the app modules repeat them. Put service configuration in the app module's own files.

JVM heap sizing in containers

The web, rest, worker, auth and web-admin images size the heap as a share of the container memory limit:

-XX:MaxRAMPercentage=${JAVA_MAX_RAM_PERCENTAGE}   # default 60.0, set in the Dockerfile

The rest of the limit is not spare. Metaspace, compressed class space and code cache alone are ~330Mi for web in production (measured: 190 + 23 + 116Mi), and GC structures, thread stacks, direct buffers and malloc arenas come on top. At the previous 75% of a 1000Mi limit the heap could grow to 750Mi and leave ~250Mi for all of that — a heap that grows to its maximum then ends in an OOMKill (no heap dump, no GC back-pressure) instead of a GC.

Tune it per deployment without rebuilding the image:

Variable Use
JAVA_MAX_RAM_PERCENTAGE Heap share of the memory limit, e.g. 50 for a tight limit
JAVA_OPTS Any other JVM flag; appended last, so it also overrides the flags above (for duplicate -XX flags the last wins)

As a rule of thumb, heap = limit − ~350Mi for web (~250Mi for rest), and the limit should leave the heap at least 1.5× its observed peak use (jvm_memory_used_bytes{area="heap"} in Grafana). The local Docker Compose environment (dev-scripts/badger/docker-compose-dev.yml) starts the jars with its own command line and is not affected.

Startup time

A JVM start is a fixed amount of CPU work: class loading and linking, bean creation, and JIT compilation (about half of it). Under a Kubernetes CPU limit a start takes that work divided by the quota. Before these changes dev web took ~88s to start at a 700m limit and then ran at ~15m. Three things bring it down.

AOT cache

The web, rest, worker and auth Dockerfiles train a JDK AOT cache (JEP 483/514/515) during the image build and start with it:

RUN ... java $JAVA_AOT_FLAGS -XX:AOTCacheOutput=app.aot -Dspring.context.exit=onRefresh ... -jar app.jar
ENTRYPOINT ... java $JAVA_AOT_FLAGS -XX:AOTCache=app.aot ...

The training run stops once the Spring context has refreshed, before anything connects to Mongo, Rabbit or Redis, so it needs no services. It records the classes the app loads and links plus the JIT's method profiles, and every later start maps them in instead of redoing that work.

Measured locally (production profile, 750Mi limit, local deps; startup as reported by Spring):

App 0.7 CPU, before 0.7 CPU, with cache 2 CPU, before 2 CPU, with cache
web 44s 22–30s 11s 7s
rest 34s 19s
worker 27–35s 16s
auth 16s 11s

CPU used during startup roughly halves as well. Heap and other non-reclaimable memory are unchanged. The mapped cache shows up as page cache, which the kernel can reclaim under pressure. Each image grows by ~40–50MB compressed.

Things to know:

  • The cache only matches the exact JDK build, jars and GC flags it was trained with. That is why it is trained in the final image stage, and why JAVA_AOT_FLAGS is shared by training and the ENTRYPOINT. On a mismatch the JVM logs a warning and starts without the cache: slower, never broken.
  • auth only gets a partial cache. Its AuthorizationServerConfig.jwkSource reads Mongo while the context is created, so the training run fails at that point. The cache still covers everything loaded before it, which is worth ~30%. Making jwkSource lazy would give auth the full cache.
  • The training run gets a throwaway BADGER_USER_COOKIE_KEY so the storefront's fail-fast lets it get that far. That value exists only in the RUN step's environment.
  • CI's DinD sidecar gets 2 CPU / 2Gi for the training runs (KUBERNETES_SERVICE_*_LIMIT in .gitlab-ci.yml, allowed by the runner's service_*_limit_overwrite_max_allowed in k3s-stack).

CPU limits

The k3s deployments set CPU requests but no CPU limits. The request still decides each pod's share when a node is busy. A limit only stops a pod from using cores that are otherwise idle, which is exactly what a start needs. Memory keeps its limit.

JAVA_TOOL_OPTIONS=-XX:ActiveProcessorCount=2 stops the JVM from sizing its GC and JIT thread pools for every core on the node.

Mongo index creation

Only the worker creates the indexes declared with @Indexed/@CompoundIndex (badger.mongo.auto-index-creation, read by MongoConfiguration). With it on, an app sends one createIndex per declared index (~115) on every boot. web, rest and sync-worker turn it off. The worker deploys with every release and registers every @Document up front, so new indexes appear when it starts.

Because of this, a local database that only ever sees web gets no indexes. Unique indexes enforce correctness, not just speed, so run the worker at least once against a fresh database.