Skip to content

Content Search

Content search makes a site's CMS pages and blog posts searchable from the same storefront search box as its products. Shoppers typing "returns" or "care guide" see matching pages and articles in the as-you-type dropdown, and the /search results page gains a Pages & articles tab with highlighted snippets and tag filtering.

It is an optional, per-site service: off by default, switched on per site, and built on the same Typesense cluster as product search (see Search Architecture) but with its own index, so turning it on never changes product search or ranking.

Turning it on

  1. Make sure the site has Typesense configured (it already does if product search works).
  2. Set Enable Content Search (contentSearchEnabled, search-config group) to true in the site configuration admin (/admin/shopSettings/guided-config), or with the MCP manageSiteConfig tool.
  3. Run Reindex search (Merchandising > Trees) once to index the site's existing pages. From then on pages are indexed as they're saved, and the nightly reconciliation keeps the index honest.

Turning it off hides content results immediately (the storefront checks the flag on every query). The site's content index is left in place, so switching it back on needs only a reindex.

What shoppers see

  • Search box dropdown — a Pages & articles group below products and categories: the page title, a page/article icon, and the matching passage with the searched words in bold. Configured per search-bar instance (Show page & article suggestions, Max page & article suggestions, default on / 3); it only ever appears on sites with content search enabled. See Search Bar & Autocomplete.
  • /search tabs — Products (n) and Pages & articles (n). The content tab (/search?q=…&tab=content) lists results with type (Page / Article), publish date and author for articles, the best-matching snippet, tags, and the article's masthead image as a thumbnail.
  • Tags — chips above the results show the tags across the match set with counts; clicking one filters to it (contentTag=…), clicking it again clears it. Each result's own tags link the same way, so /search?tab=content&contentTag=Guides browses a topic with no query at all (newest first).
  • Nothing in the catalogue? If a search finds no products but does find content, /search opens on the Pages & articles tab rather than showing an empty products page.

Every customer-facing theme is covered: Nova and the themes layered on it (Brock, Depot, Pop, Atelier, Spec) share nova/fragment/content-results.html; the Bootstrap family uses bootstrap/fragment/content-results.html.

What gets indexed

One document per page, in a per-tenant Typesense collection pages__{catalogueId}.

Field From
title The page name (falling back to its first heading) — weighted highest
headings Markdown # headings, HTML <h1>–<h6>, JSON-component headings, accordion titles and tab labels, the hero's title slot, FAQ title and questions
tags The page's tags (also a facet)
description The page description, else its metaDescription attribute
body Everything else the page shows as text (capped at 50,000 characters)
type article for blog posts (blogPost stereotype), else page
url The page's canonical path, else /p/{seoName}
imageUrl The blog masthead image
author, publishedDate attributes.author; the page's creation date

Text comes from the extensions the page actually shows — its own enabled placements plus those inherited from its stereotype, merged exactly as rendering merges them (ExtensionService.findEffectivePlacements). It is read from each extension's stored configuration, not from a render: indexing runs in the worker, which has no templates or request, and a render would also pull in navigation, personalisation and analytics.

Extension Indexed
markdownFragment Markdown source, syntax stripped; # lines as headings
htmlFragment HTML with scripts/styles/comments removed; <h*> as headings
jsonComponent, hero Literal text, quote/author, image alt props; headings as above. Data-bound props ({"$ref": …}) and button/link labels are skipped
faqExtension Title and questions as headings, answers as body
blogMastheadExtension Its image becomes the result thumbnail

Product grids, forms and other functional blocks contribute nothing — they aren't content. To make a new content extension searchable, add a @Component implementing ContentTextExtractor (commerce-core, search/content).

Left out: disabled pages; pages ticked Exclude from Site Search (page editor, next to Exclude from Sitemap); global system pages (catalogue ALL); platform pages (search, basket, checkout, account, login); and pages with no text and no description.

How it stays fresh

Trigger Effect
Any page save (admin, MCP, import, extension edits — anything through MongoTemplate) PageSearchIndexListener → PageIndexEvent → the page is re-extracted and upserted, or removed if it no longer qualifies
PageService.delete The page's document is removed
Stereotype save StereotypeSearchIndexListener → ContentReindexRequestedEvent → the site's content index is rebuilt, since every page using the stereotype may have changed
Reindex search (Merchandising > Trees) Rebuilds products and content
Nightly reconciliation Rebuilds content for every enabled site and prunes documents whose page is gone (including deletes that bypassed the service)

All of these travel the existing search.index.* RabbitMQ route to the worker's ContentIndexer, and every one is a no-op for a site without content search — the decision lives in one place, ContentSearchSettings, read by both the worker and the storefront.

Ranking

Queries search title, headings, tags, description, body with weights 8, 5, 4, 2, 1, so a page titled "Returns" beats a page that mentions returns in passing. Ties, and tag browsing with no query, sort newest first. Snippets come from Typesense highlighting of the body (else the description) and are returned as plain-text runs. The storefront escapes them and wraps only the matched words in <mark>/<strong>, so nothing from the index is ever rendered as HTML.

Not yet

  • A REST endpoint (/v1/search/content) for headless storefronts. It will be added spec-first to api.yaml alongside /v1/search/products.
  • Indexing products' own long-form content, or collections/merchandising-node descriptions.
  • Synonyms and per-site ranking tweaks.