Skip to content

Observability

Every Badger JVM service exposes Micrometer metrics over Actuator. In the k3s cluster vmagent scrapes them into VictoriaMetrics, Promtail ships the logs into Loki, and Grafana reads both.

This page covers what the applications expose and how to add to it. The dashboards and scrape configuration live in the k3s-stack repository under flux/infrastructure/monitoring/.

How a service gets scraped

Three things have to line up. Miss any one and the service is invisible.

  1. micrometer-registry-prometheus on the classpath. commerce-core already declares it, so every module that depends on commerce-core inherits it.
  2. prometheus exposed over HTTP, via management.endpoints.web.exposure.include.
  3. A management port the scraper can reach, declared to Kubernetes with prometheus.io/scrape, prometheus.io/port and prometheus.io/path annotations on the pod template — annotations on the Service are not enough, because discovery is pod-level.

web, rest and auth all run their management server on MANAGEMENT_PORT (9090 in the cluster), separate from the public application port, so actuator is never reachable from the internet.

badger-worker exports no metrics

The worker sets spring.main.web-application-type: none, so no HTTP server starts and Actuator has nowhere to publish. It shows up on the dashboards with replica counts, restarts, CPU, memory and logs — all from Kubernetes and Loki — and nothing JVM-level. Giving it metrics means giving it a management server, which is a deliberate change to a process that currently opens no ports.

What is exposed today

Area Metrics Useful because
HTTP http_server_requests_seconds_* tagged uri, method, status, outcome, exception Rate, errors and duration per endpoint
Logging logback_events_total{level} Error counts at the real level, counted in-process — no log-scraping regex involved
JVM jvm_memory_*, jvm_gc_*, jvm_threads_*, jvm_classes_*, process_cpu_usage Heap pressure, GC cost, thread saturation
Startup application_started_time_seconds, application_ready_time_seconds The number to watch when tuning image layering or CDS
Caches cache_gets_total{cache,result}, cache_size, cache_evictions_total Hit ratio per named cache
Redis lettuce_seconds_* Command rate and latency
MongoDB spring_data_repository_invocations_seconds_* tagged repository, method, state Per-repository latency and error rate
Messaging rabbitmq_* Publish/consume flow, rejects, unrouted messages
Web server jetty_threads_*, jetty_connections_* Thread-pool saturation
Scheduled work tasks_scheduled_execution_seconds_* Whether scheduled jobs run and how long they take
Per-tenant badger_performance_seconds* tagged method, siteId The performance monitor aspect

The Mongo driver meters (mongodb_driver_pool_*, mongodb_driver_commands_seconds_*) come from Spring Boot's MongoMetricsAutoConfiguration, which contributes its command and pool listeners as MongoClientSettingsBuilderCustomizers. MongoConfiguration builds the client itself, so it applies those customizers explicitly (it used to discard the settings object entirely, and the meters never appeared). They sit below the Spring Data repository timer and include queries that don't go through a repository — e.g. @DBRef resolution. See Caching for the cache metrics that matter most.

Latency percentiles

Micrometer publishes timers as _count, _sum and _max only. Without histogram buckets, histogram_quantile() returns nothing at all — so a p95 panel or alert built on http_server_requests_seconds_bucket silently shows no data rather than failing loudly.

web, rest and auth therefore configure explicit boundaries:

management:
  metrics:
    distribution:
      slo:
        http.server.requests: 25ms,50ms,100ms,200ms,400ms,800ms,1s,2s,5s,10s
      minimum-expected-value:
        http.server.requests: 10ms
      maximum-expected-value:
        http.server.requests: 10s

These are SLO boundaries rather than percentiles-histogram: true on purpose. The full histogram generates roughly 40–70 buckets per series; across a couple of hundred endpoint/status/method combinations that is tens of thousands of extra series. Ten boundaries cost about 11 series per combination and give percentile estimates that are accurate in the range anyone acts on.

If you add a timer that needs percentiles, add its name here rather than turning on percentile histograms globally.

Adding a metric

Inject MeterRegistry and register a meter. Keep tag values bounded — a tag that can take user-supplied values (an order id, a URL, an email) multiplies series count without limit and is the usual cause of a monitoring system falling over.

@Service
@RequiredArgsConstructor
public class RefundService {

  private final MeterRegistry meterRegistry;

  public void refund(Order order) {
    // ...
    meterRegistry.counter("badger.refunds", "reason", reason.name()).increment();
  }
}

badger.refunds is exported as badger_refunds_total. Micrometer replaces dots with underscores and appends a suffix per meter type, so name meters in dotted form and read them back in underscored form.

The performance monitor aspect

PerformanceMonitor times every public method on *Service, *Repository, *Renderer and *Configurator classes, plus anything annotated @Monitored, and tags each timing with the method signature and the current tenant's siteId. See the performance-monitoring section of CLAUDE.md for how to opt a class in.

Those timings reach Grafana as badger_performance_seconds*, which is what drives the per-tenant panels. The admin screen at /system/performanceStatistics reads the same data out of Redis for a live view; Grafana keeps the history.

Because the aspect tags by siteId, this is the one place where tenant count drives cardinality. That is intentional and worth the cost, but it is a reason not to add further high-cardinality tags to the same meter.

Logs

Logback writes to stdout; Promtail tails the container logs and ships them to Loki with namespace, pod, container, node and app labels. Log lines are stored as plain text — no JSON parsing — so LogQL filters use line matchers (|~ "(?i)error") rather than label filters over parsed fields.

For counting errors, prefer logback_events_total{level="error"} over matching log text: it is counted in-process at the real level, so it cannot be fooled by the word "error" appearing in a stack trace or a URL.