Build and Runtime Footprint¶
Operational knobs that keep images small and services inside their resource limits. Each section says what the default is, why, and how to change it.
Vendored front-end libraries¶
Third-party browser libraries (jQuery, Bootstrap, Monaco, Summernote, ...) are
npm dependencies in web-mvc/package.json. During the Maven build
web-mvc/scripts/copy-versioned-deps.js copies them into the web-mvc jar
under resources/vendor/{package}/{version}/, served at
/vendor/{package}/{version}/.... Templates reach them through the
system.css.and.js.cdn.url message, extensions through @vendor/... paths:
Only the files named in the script's ALLOW_LIST are copied — not whole
packages. An npm package also ships sources, dev and ESM builds, sourcemaps,
docs and tests; copying everything put ~162 MB into the jar (and into the
web and rest images) for about 3 MB of files anything asks for.
Referencing a new vendored file¶
- Add the file, and anything it loads relatively (fonts, images, lazily
loaded modules), to that package's globs in
ALLOW_LIST.*matches within a path segment and**across segments..mapfiles are skipped unless a glob names them. - Build.
scripts/check-vendored-refs.jsruns straight after the copy and fails the build if a reference is missing. It collects everycdn.url('...'),'/vendor/...'and'@vendor/...'string in every module'ssrc/main, followsurl(...)and@importin each referenced stylesheet, and requires a referenced directory to be copied whole.
Run it by hand, in a container, to see every URL it checks:
docker run --rm -v "$PWD":/r -w /r/web-mvc node:22-slim \
sh -c "npm ci --ignore-scripts --legacy-peer-deps && npm run copy-deps && node scripts/check-vendored-refs.js --list"
Adding a dependency¶
Every package in package.json must have an ALLOW_LIST entry; the copy
fails otherwise, and also fails if a glob matches nothing (typically because
an upgrade moved the dist files). So adding a library means choosing what of
it the browser needs.
Monaco is the special case: its AMD loader is pointed at min/vs
(require.config({ paths: { vs: '/vendor/monaco-editor/<ver>/min/vs' } })) and
lazily fetches workers, language modes, CSS and the codicon font from anywhere
under it, so the whole min/vs tree is kept.
The checker only sees references in the source tree. A /vendor/... URL typed
into site content in the database (custom HTML, page content) is not checked.
Storefront critical path¶
What a Nova page blocks on before first paint is kept to jQuery, the Bootstrap JS bundle, the
Bootstrap CSS and the theme's bdgr.css. Keep it that way:
- Fonts are self-hosted. Inter comes from
@fontsource-variable/inter. Its@font-facerules are inline innova/fragment/head.html: one variable file per unicode subset, and the Latin file is preloaded. Don't add Google Fonts<link>s: each one is a render-blocking third-party stylesheet plus two extra connections, and it sends visitors' IPs to Google. - Icons on every page are inline SVG (
nova/fragment/icons.html, generated from Font Awesome 5.15.4's own SVGs). Font Awesome's CSS still loads for content that usesfa-*classes, but viarel="preload"+onload, so it doesn't block render.html.fa-loadedmarks when it has applied. A page whose content uses no brand icons never downloads the 77 KB brands font. - Page-specific CSS/JS belongs with the template that needs it. The product gallery loads
Splide/PhotoSwipe, the legacy bootstrap theme's galleries load Fotorama
(
bootstrap/fragment/fotorama), and the two animated bootstrap-theme alerts load animate.css. None of them are global. - Our own scripts carry
data-cookieconsent="ignore". Tenants that run Cookiebot in auto-blocking mode otherwise have every script held back until the consent config loads. Each one is then re-inserted asasync, so it's fetched twice and the jQuery → Bootstrap → inline-script order is lost. Our scripts set no cookies; tenant tracking scripts stay consent-gated. Add the attribute to any new first-party<script>in a head/footer fragment. Stripe.js carries it too: its cookies are strictly necessary for fraud prevention, and on live Cookiebot had classed it as "preferences", which blocked checkout for shoppers who declined those cookies. - View transitions are cross-document. A navigation loads a fresh page, so don't
re-initialise anything on
pagereveal.view-transitions.jsused to replay every jQuery ready handler there, and every page's JS ran twice.
Collection listing index¶
The collection listing (ProductService.findBySiteAndCollectionIDs, used by
CollectionController and RESTCollectionController) filters on
catalogueId + collectionIDs and sorts by attributes.collectionOrdering
descending, then _id. It is served by one compound index on product:
{ catalogueId: 1, collectionIDs: 1, "attributes.collectionOrdering": -1, _id: 1 } name: catalogue_collection_listing_idx
With it a page is read straight off the index in order — no in-memory SORT
stage, and only the page's documents are fetched. On 20,000 seeded products
(15,000 in the collection), page 11 went from 15,000 keys + 15,000 documents
examined to 264 keys + 24 documents.
ProductListingIndexInitializer (in common) creates it on a background
thread once the application is ready. It is deliberately not a class-level
@CompoundIndex: Variant extends Product and is embedded in
Product.variants, and Spring Data would build a second, unused copy under
variants.*. createIndex is a no-op when the index already exists, and a
failure is logged, not fatal.
Sorting by price, rating, productName (the other validSortByValues) or
by ascending collectionOrdering still uses the index for the filter, but
sorts the matched set in memory.
Build it by hand first on a large catalogue
On a large product collection, create the index manually before deploying,
with the same name and keys, so no pod builds it at startup:
// Must be 0: collectionIDs is already an array, and MongoDB cannot index two
// arrays in one document ("cannot index parallel arrays").
db.product.countDocuments({ "attributes.collectionOrdering": { $type: "array" } })
db.product.createIndex(
{ catalogueId: 1, collectionIDs: 1, "attributes.collectionOrdering": -1, _id: 1 },
{ name: "catalogue_collection_listing_idx" })
Once the index exists, a write that sets attributes.collectionOrdering to
an array is rejected. The admin UI stores it as a single value.
Sentry performance tracing¶
Errors go to Sentry whenever SENTRY_DSN is set. Performance tracing is
controlled per deployment by SENTRY_TRACES_SAMPLE_RATE, bound to
sentry.traces-sample-rate in each service's application-production.yml:
SENTRY_TRACES_SAMPLE_RATE |
Effect |
|---|---|
| unset (default) | Tracing off — no transactions are created or sent |
0.1 |
10% of requests are traced |
1.0 |
Every request is traced — avoid in production |
common/application.yml is not loaded by the services
Spring Boot reads only the first classpath:application.yml, and every app
module (web, rest, worker, auth, sync-worker) ships its own, which
shadows the one in common. The sentry: block there
(traces-sample-rate: 1.0, send-default-pii: true) has therefore never
applied: production has run with tracing off and send-default-pii false.
The same goes for the badger.stripe, badger.aws and
server.compression settings in that file, which is why the app modules
repeat them. Put service configuration in the app module's own files.
JVM heap sizing in containers¶
The web, rest, worker, auth and web-admin images size the heap as a
share of the container memory limit:
The rest of the limit is not spare. Metaspace, compressed class space and code
cache alone are ~330Mi for web in production (measured: 190 + 23 + 116Mi),
and GC structures, thread stacks, direct buffers and malloc arenas come on top.
At the previous 75% of a 1000Mi limit the heap could grow to 750Mi and leave
~250Mi for all of that — a heap that grows to its maximum then ends in an
OOMKill (no heap dump, no GC back-pressure) instead of a GC.
Tune it per deployment without rebuilding the image:
| Variable | Use |
|---|---|
JAVA_MAX_RAM_PERCENTAGE |
Heap share of the memory limit, e.g. 50 for a tight limit |
JAVA_OPTS |
Any other JVM flag; appended last, so it also overrides the flags above (for duplicate -XX flags the last wins) |
As a rule of thumb, heap = limit − ~350Mi for web (~250Mi for rest), and
the limit should leave the heap at least 1.5× its observed peak use
(jvm_memory_used_bytes{area="heap"} in Grafana). The local Docker Compose
environment (dev-scripts/badger/docker-compose-dev.yml) starts the jars with
its own command line and is not affected.
Startup time¶
A JVM start is a fixed amount of CPU work: class loading and linking, bean
creation, and JIT compilation (about half of it). Under a Kubernetes CPU limit a
start takes that work divided by the quota. Before these changes dev web took
~88s to start at a 700m limit and then ran at ~15m. Three things bring it down.
AOT cache¶
The web, rest, worker and auth Dockerfiles train a JDK AOT cache
(JEP 483/514/515) during the image build and start with it:
RUN ... java $JAVA_AOT_FLAGS -XX:AOTCacheOutput=app.aot -Dspring.context.exit=onRefresh ... -jar app.jar
ENTRYPOINT ... java $JAVA_AOT_FLAGS -XX:AOTCache=app.aot ...
The training run stops once the Spring context has refreshed, before anything connects to Mongo, Rabbit or Redis, so it needs no services. It records the classes the app loads and links plus the JIT's method profiles, and every later start maps them in instead of redoing that work.
Measured locally (production profile, 750Mi limit, local deps; startup as reported by Spring):
| App | 0.7 CPU, before | 0.7 CPU, with cache | 2 CPU, before | 2 CPU, with cache |
|---|---|---|---|---|
| web | 44s | 22–30s | 11s | 7s |
| rest | 34s | 19s | ||
| worker | 27–35s | 16s | ||
| auth | 16s | 11s |
CPU used during startup roughly halves as well. Heap and other non-reclaimable memory are unchanged. The mapped cache shows up as page cache, which the kernel can reclaim under pressure. Each image grows by ~40–50MB compressed.
Things to know:
- The cache only matches the exact JDK build, jars and GC flags it was trained
with. That is why it is trained in the final image stage, and why
JAVA_AOT_FLAGSis shared by training and theENTRYPOINT. On a mismatch the JVM logs a warning and starts without the cache: slower, never broken. authonly gets a partial cache. ItsAuthorizationServerConfig.jwkSourcereads Mongo while the context is created, so the training run fails at that point. The cache still covers everything loaded before it, which is worth ~30%. MakingjwkSourcelazy would giveauththe full cache.- The training run gets a throwaway
BADGER_USER_COOKIE_KEYso the storefront's fail-fast lets it get that far. That value exists only in theRUNstep's environment. - CI's DinD sidecar gets 2 CPU / 2Gi for the training runs
(
KUBERNETES_SERVICE_*_LIMITin.gitlab-ci.yml, allowed by the runner'sservice_*_limit_overwrite_max_allowedin k3s-stack).
CPU limits¶
The k3s deployments set CPU requests but no CPU limits. The request still decides each pod's share when a node is busy. A limit only stops a pod from using cores that are otherwise idle, which is exactly what a start needs. Memory keeps its limit.
JAVA_TOOL_OPTIONS=-XX:ActiveProcessorCount=2 stops the JVM from sizing its GC
and JIT thread pools for every core on the node.
Mongo index creation¶
Only the worker creates the indexes declared with @Indexed/@CompoundIndex
(badger.mongo.auto-index-creation, read by MongoConfiguration). With it on,
an app sends one createIndex per declared index (~115) on every boot. web,
rest and sync-worker turn it off. The worker deploys with every release and
registers every @Document up front, so new indexes appear when it starts.
Because of this, a local database that only ever sees web gets no indexes.
Unique indexes enforce correctness, not just speed, so run the worker at least
once against a fresh database.