Skip to content

Search and Discovery

Badger Commerce's search and discovery layer is built around a single principle: search is the substrate for browsing, not a feature bolted on top of it. Category pages, search results, virtual categories, megamenu queries, and curated landing pages all resolve through the same Typesense-backed search platform over the same per-tenant index.

For the architectural principles and domain model behind this, see Search Architecture. This document covers the customer-facing surface and what operators can author.

Operator Surface

Operators author discovery through four primitives, all of which are unified into a single merchandising tree model:

Manual Collections

Hand-curated groupings of products by seoName. Suit small catalogues, hero collections, and any case where editorial control beats classification. Existing customers' collections continue to work without migration — collections are now one node-type within the merchandising tree.

Taxonomy Projections

A node in the merchandising tree backed by a Taxonomy level. Products assigned to that level (or any descendant) appear automatically. Optional attribute filters narrow the projection further. The right tool when a customer's taxonomy already organises the catalogue the way customers should browse it.

Virtual Categories

A node backed by a saved query — "orange clothes under £100", "in-stock vegan footwear", "manufacturer = Apple AND rating > 4". Authored once in admin and indexed; behaves to the customer like any other category, with its own URL, SEO, and breadcrumb. Powerful for seasonal pages, themed landings, and merchandising experiments without code.

Manual Lists

A node backed by an explicit, ordered list of products. Useful for landing pages, gift guides, "Featured Today", and curated experiences where ordering and exact membership matter.

These can be mixed freely within a tree. A primary tree can use a taxonomy projection at the top, virtual queries on seasonal nodes, manual collections at the leaves, and a manual list for the homepage hero.

Multiple Trees per Site

A site has one primary tree (which drives canonical URLs and SEO) and may define additional named trees for:

  • Megamenus and navigation — different shapes for header, footer, mobile drawer.
  • Gift guides and seasonal landings — built as full trees, reused as needed.
  • A/B testing alternative information architectures (planned for round 2).

The same product can appear in many nodes across many trees without duplication; node membership is computed once at index time.

Search Engine

The platform uses Typesense as its search engine. A shared cluster runs one Typesense collection per tenant (products__{catalogueId}), giving each site:

  • Sub-50ms typeahead and full-text search at typical catalogue sizes.
  • Typo tolerance out of the box.
  • Per-tenant schema evolution — facetable attributes derived from the tenant's active taxonomy, not a lowest-common-denominator schema.
  • Strong tenant isolation — cross-tenant queries are impossible by construction.

Index freshness is best-effort within seconds (event-driven via RabbitMQ) with nightly reconciliation against Mongo to catch drift.

Faceted Navigation

Facets are derived from the active taxonomy and surfaced according to a three-level fallback (node → tree default → taxonomy default). Most sites configure a tree-level default and rarely touch individual nodes.

Supported facet types:

  • Categorical — colour, size, manufacturer, material (any taxonomy attribute).
  • Boolean — in stock, on sale, featured.
  • Range — price, rating, custom numeric attributes.
  • Variant roll-up — variant colours/sizes appear as parent-product facets, so a "red shirts" facet shows each shirt once.

Variants in Listings

A search returns one document per product. Variant attributes (colour, size, SKU codes) roll up onto the parent so:

  • Listings dedupe naturally — the red shirt appears once, not once per variant.
  • Facet counts are by product, matching shopper intuition.
  • SKU-level lookups remain available via dedicated index fields.

SEO

Each merchandising node owns its SEO surface:

  • seoName and a denormalised canonicalPath (mirrors the existing Collection.canonicalPath model).
  • Optional override of meta title, meta description, and canonical URL.
  • Sitemap inclusion controlled per-node.
  • Breadcrumb chain denormalised onto each node for cheap rendering.

The primary tree's structure determines canonical URLs; non-primary trees can reuse nodes without competing for canonicality.

Indexing Pipeline

Mongo write  ──►  ApplicationEvent  ──►  RabbitMQ  ──►  Worker  ──►  Typesense
                                                                          ▲
                                          Nightly Reconciliation ─────────┘
  • Event-driven updates: product saves, collection edits, taxonomy changes, and merchandising tree mutations all flow through RabbitMQ to a worker that updates Typesense.
  • Reconciliation: a nightly job re-asserts state against Mongo, catching any missed events and removing orphans.
  • Schema migration via alias swap: when the facetable attribute set changes, a new versioned collection is built in the background and the stable alias is repointed atomically. Zero downtime, instant rollback.

See Search Architecture for the full pipeline detail.

Administration

Operators interact with search through the admin UI:

  • Merchandising → Trees — create and manage merchandising trees, mark one primary.
  • Tree editor — author nodes, choose backing types, configure SEO and facets, preview matching products before indexing.
  • Taxonomy admin — manage the classification that drives projections and facets.
  • Collection admin — manual collections continue to be authored as before.

The platform handles index lifecycle, schema migrations, breadcrumb computation, alias swaps, and node resolution. None of these are operator concerns.

Search Term Analytics

The platform records what visitors search for on each site, so merchants and AI agents can see the most popular searches and the searches that found nothing (gaps in the range, missing synonyms, products that need better names). It is built into the platform rather than using Typesense analytics, which isn't configured and would mix tenants in one store.

What is recorded. A search is counted when a visitor lands on the results page, /search?q=…, with a fresh query: the first page, with no facet, sort, tab or content-tag refinement. Paging, sorting and filtering the same search don't count it again. The search-box dropdown (/search/suggest) is not recorded, because it fires as the visitor types; pressing Enter or choosing "see all results" takes them to /search, which is. Clicking a product or category straight from the dropdown isn't counted. The result count is products plus, when content search is on, matching pages and posts.

Privacy. Only the normalised term, a daily count, a zero-result count and the latest result count are stored. No user id, session id, IP address or user agent is kept. Terms are normalised (trimmed, lower-cased, whitespace collapsed, capped at 100 characters) and dropped entirely when they look like personal data: an email address, a run of 8 or more digits, or a phone-like sequence of 9 or more digits. These checks run on the whole entry, before it is truncated.

Storage and retention. SearchTermAnalyticsServiceImpl (commerce-core, search.analytics) buffers searches in memory, keyed by site, day and term, and writes them every 10 seconds (badger.search-analytics.flush-interval-ms) and at shutdown, as one unordered bulk of $inc upserts into the searchTermDailyStats Mongo collection. The search request itself only does a config lookup and a map update. Each document is one term on one site on one day (siteId, day as yyyy-MM-dd in the site's time zone, term, count, zeroResultCount, lastResultCount, lastSearchedAt), with a unique index on (siteId, day, term), so several web nodes add to the same bucket. A TTL index on expireAt deletes each bucket 90 days after its day. The buffer holds at most 20,000 distinct entries between flushes. Past that, new terms are dropped (repeats still count), so a flood of junk searches can't exhaust the heap. A failed write is retried on the next flush.

Switching it off. The search-term-tracking-enabled site config key (search-config, default on) stops new recording for a site. Searches already recorded stay until they expire.

Where to see it. The admin dashboard has a Searches tab next to the page-view tabs. It shows the top 20 searches and the top 20 zero-result searches for the last 30 days, each with the result count of the latest search. AI agents use the MCP siteTraffic tool's topSearches and zeroResultSearches actions, with optional from/to (yyyy-MM-dd, default the last 30 days, clamped to the 90-day retention) and limit (default 10, max 50). Every read goes through SearchTermAnalyticsService with the caller's SiteContext, and every query filters on its siteId, so a tenant only sees its own terms.

Current State and Roadmap

Round 1 (in progress) delivers the foundation: domain model, persistence, admin scaffolding, and the indexing pipeline. Customer-facing rendering still serves from Mongo until round 2.

Round 2 will land:

  • Customer-facing query API (/api/search, /api/merchandising/{path}).
  • Rendering category pages, search results, and autocomplete from Typesense.
  • URL strategy and redirects from legacy /collection/{seoName} paths.

Further out:

  • Personalisation and ranking signals (popularity, recency, user history).
  • A/B testing of merchandising trees.
  • Synonym and search-redirect management.
  • SaaS Sett authoring tooling — collaborative editing, drafts, scheduled publishing.