Observability¶
Every Badger JVM service exposes Micrometer metrics over Actuator. In the k3s cluster vmagent scrapes them into VictoriaMetrics, Promtail ships the logs into Loki, and Grafana reads both.
This page covers what the applications expose and how to add to it. The
dashboards and scrape configuration live in the k3s-stack repository under
flux/infrastructure/monitoring/.
How a service gets scraped¶
Three things have to line up. Miss any one and the service is invisible.
micrometer-registry-prometheuson the classpath.commerce-corealready declares it, so every module that depends oncommerce-coreinherits it.prometheusexposed over HTTP, viamanagement.endpoints.web.exposure.include.- A management port the scraper can reach, declared to Kubernetes with
prometheus.io/scrape,prometheus.io/portandprometheus.io/pathannotations on the pod template — annotations on the Service are not enough, because discovery is pod-level.
web, rest and auth all run their management server on MANAGEMENT_PORT
(9090 in the cluster), separate from the public application port, so actuator is
never reachable from the internet.
badger-worker exports no metrics
The worker sets spring.main.web-application-type: none, so no HTTP server
starts and Actuator has nowhere to publish. It shows up on the dashboards
with replica counts, restarts, CPU, memory and logs — all from Kubernetes and
Loki — and nothing JVM-level. Giving it metrics means giving it a management
server, which is a deliberate change to a process that currently opens no
ports.
What is exposed today¶
| Area | Metrics | Useful because |
|---|---|---|
| HTTP | http_server_requests_seconds_* tagged uri, method, status, outcome, exception |
Rate, errors and duration per endpoint |
| Logging | logback_events_total{level} |
Error counts at the real level, counted in-process — no log-scraping regex involved |
| JVM | jvm_memory_*, jvm_gc_*, jvm_threads_*, jvm_classes_*, process_cpu_usage |
Heap pressure, GC cost, thread saturation |
| Startup | application_started_time_seconds, application_ready_time_seconds |
The number to watch when tuning image layering or CDS |
| Caches | cache_gets_total{cache,result}, cache_size, cache_evictions_total |
Hit ratio per named cache |
| Redis | lettuce_seconds_* |
Command rate and latency |
| MongoDB | spring_data_repository_invocations_seconds_* tagged repository, method, state |
Per-repository latency and error rate |
| Messaging | rabbitmq_* |
Publish/consume flow, rejects, unrouted messages |
| Web server | jetty_threads_*, jetty_connections_* |
Thread-pool saturation |
| Scheduled work | tasks_scheduled_execution_seconds_* |
Whether scheduled jobs run and how long they take |
| Per-tenant | badger_performance_seconds* tagged method, siteId |
The performance monitor aspect |
The Mongo driver meters (mongodb_driver_pool_*,
mongodb_driver_commands_seconds_*) come from Spring Boot's
MongoMetricsAutoConfiguration, which contributes its command and pool
listeners as MongoClientSettingsBuilderCustomizers. MongoConfiguration
builds the client itself, so it applies those customizers explicitly (it used
to discard the settings object entirely, and the meters never appeared). They
sit below the Spring Data repository timer and include queries that don't go
through a repository — e.g. @DBRef resolution. See Caching for
the cache metrics that matter most.
Latency percentiles¶
Micrometer publishes timers as _count, _sum and _max only. Without
histogram buckets, histogram_quantile() returns nothing at all — so a p95 panel
or alert built on http_server_requests_seconds_bucket silently shows no data
rather than failing loudly.
web, rest and auth therefore configure explicit boundaries:
management:
metrics:
distribution:
slo:
http.server.requests: 25ms,50ms,100ms,200ms,400ms,800ms,1s,2s,5s,10s
minimum-expected-value:
http.server.requests: 10ms
maximum-expected-value:
http.server.requests: 10s
These are SLO boundaries rather than percentiles-histogram: true on purpose.
The full histogram generates roughly 40–70 buckets per series; across a couple of
hundred endpoint/status/method combinations that is tens of thousands of extra
series. Ten boundaries cost about 11 series per combination and give percentile
estimates that are accurate in the range anyone acts on.
If you add a timer that needs percentiles, add its name here rather than turning on percentile histograms globally.
Adding a metric¶
Inject MeterRegistry and register a meter. Keep tag values bounded — a tag that
can take user-supplied values (an order id, a URL, an email) multiplies series
count without limit and is the usual cause of a monitoring system falling over.
@Service
@RequiredArgsConstructor
public class RefundService {
private final MeterRegistry meterRegistry;
public void refund(Order order) {
// ...
meterRegistry.counter("badger.refunds", "reason", reason.name()).increment();
}
}
badger.refunds is exported as badger_refunds_total. Micrometer replaces dots
with underscores and appends a suffix per meter type, so name meters in dotted
form and read them back in underscored form.
The performance monitor aspect¶
PerformanceMonitor times every public method on *Service, *Repository,
*Renderer and *Configurator classes, plus anything annotated @Monitored,
and tags each timing with the method signature and the current tenant's
siteId. See the performance-monitoring section of CLAUDE.md for how to opt a
class in.
Those timings reach Grafana as badger_performance_seconds*, which is what
drives the per-tenant panels. The admin screen at
/system/performanceStatistics reads the same data out of Redis for a live
view; Grafana keeps the history.
Because the aspect tags by siteId, this is the one place where tenant count
drives cardinality. That is intentional and worth the cost, but it is a reason
not to add further high-cardinality tags to the same meter.
Logs¶
Logback writes to stdout; Promtail tails the container logs and ships them to
Loki with namespace, pod, container, node and app labels. Log lines are
stored as plain text — no JSON parsing — so LogQL filters use line matchers
(|~ "(?i)error") rather than label filters over parsed fields.
For counting errors, prefer logback_events_total{level="error"} over matching
log text: it is counted in-process at the real level, so it cannot be fooled by
the word "error" appearing in a stack trace or a URL.