Skip to content

Site Data Transfer

BadgerCommerce provides built-in export and import functionality for transferring site configuration and catalogue data between environments. This feature is essential for copying data from production to development, creating site templates, or migrating between installations.

Key Features

Data Export

  • Background Export: Exports run asynchronously in the worker, not on the web request. The archive is streamed straight off a Mongo cursor to a temp file and uploaded to S3, so even very large sites export with a flat, bounded memory footprint (no out-of-memory failures).
  • Queue → poll → download: Requesting an export creates an ExportJob; the admin UI polls its status and downloads via a short-lived pre-signed S3 URL once it reports COMPLETED.
  • Selective Export: Only exports catalogue and configuration data - excludes PII and transactional data
  • Password Protected: Requires admin password confirmation for security
  • Portable Format: Standard ZIP file with JSON data, easily inspectable and version-controllable

Data Import

  • Validation Preview: See what will be imported before committing
  • Conflict Resolution: Automatically replaces existing data in the target site
  • ID Remapping: All MongoDB IDs and cross-references are automatically remapped
  • Tenant Isolation: Imported data is fully owned by the target site

Privacy by Design

The export explicitly excludes all personally identifiable information (PII) and transactional data:

Exported (Safe) Excluded (PII/Transactional)
Products & Variants Orders
Collections Users/Customers
Pages (CMS) Donations
Taxonomies Gift Aid Declarations
Menus Gift Cards
Media metadata Subscriptions/Invoices
Delivery Options Fundraiser Requests
Promotions (automatic and code-entered) Analytics/Page Views
Stereotypes AI Conversations
Variant Groups (Colour, Size…) Sensitive config values
Site Text overrides Sales counts
Stock on hand per SKU
Site Configuration*

*Site configuration excludes sensitive values like API keys, tokens, and secrets. It also leaves behind keys that describe the source site itself: its Badger Sett organisation id (which routes subscription webhooks and federation to every site carrying it) and the legal-entity sign-off (legal-entity-confirmed-at / -by), so the target site's owner confirms their own details.

Storefront theme: the manifest records the source site's theme (sourceThemeName) and the import switches the target site to it, provided that theme is installed on the target database. If it isn't (say, a bespoke theme), the target keeps its current theme and the validation preview warns about it.

Not carried: the rest of the site (tenant) document — logo, social accounts, product image path, domains — and merchandising trees. The target site keeps its own.

Export File Structure

site-export-{siteId}-{timestamp}.zip
├── manifest.json              # Export metadata (version, timestamp, source site and theme)
├── catalogue/
│   ├── products.json          # All products with variants
│   ├── collections.json       # Product collections/categories
│   ├── media.json             # Media metadata (URLs, not binaries)
│   ├── variant-groups.json    # Colour / Size / … definitions the variants refer to
│   └── stock.json             # Stock on hand per SKU and warehouse
├── content/
│   ├── pages.json             # CMS pages
│   └── menus.json             # Navigation menus
├── taxonomy/
│   └── taxonomies.json        # Product taxonomies
└── configuration/
    ├── site-configurations.json  # Site settings (non-sensitive)
    ├── stereotypes.json          # Product/page stereotypes
    ├── delivery-options.json     # Shipping methods
    ├── site-text.json            # Site Text overrides (Admin > Shop Settings > Site Text)
    └── promotions.json           # Promotions and discounts (Jackson default typing)

How ID Remapping Works

When importing data from Site A into Site B, all identifiers are automatically remapped:

Field On Export On Import
catalogueId Source site's ID Replaced with target site's ID
siteId Source site's ID Replaced with target site's ID
MongoDB _id Original IDs New ObjectIds generated
Cross-references Original IDs Remapped to new IDs

Some references are remapped specially: a product's images (media ids), any extension's mediaId setting (the blog masthead image, which blog cards and og:image also use) on products, collections, pages and stereotypes, and the site's active taxonomy (taxonomy-active-id, which names a taxonomy by id). Collections, pages and other items that arrive without an id each get their own new one.

This ensures: - No ID collisions between sites - Complete data ownership by the target site - No lingering references to the source site

Usage

Accessing the Feature

  1. Navigate to System > Tenants
  2. Click on the tenant you want to export from or import into
  3. Find the Data Transfer panel on the right side

Exporting Site Data

  1. Click Export Site Data
  2. Enter your admin password when prompted
  3. The export is queued to the worker and the dialog shows progress while it generates in the background. Large catalogues may take a little while.
  4. When it completes, the download starts automatically (via a pre-signed S3 link). The file is named site-export-{siteId}-{timestamp}.zip

The export job is processed by the worker, so you can navigate away — the generated archive is stored in S3 and the download link remains valid for 24 hours.

Importing Site Data

  1. Click Import Site Data and select a ZIP file
  2. Enter your admin password when prompted
  3. Review the validation summary showing:
  4. Source site information
  5. Entity counts to be imported
  6. Warnings about data replacement
  7. Click Import Now to proceed
  8. Wait for the import to complete

Once the import succeeds, a full search reindex is queued for the site (products, and pages and blog posts for content search), since the imported documents were never indexed one by one.

Use Cases

Development Environment Setup

Export production catalogue data to set up a realistic development environment without exposing customer data.

Site Templates

Create a "template site" with products, pages, and configuration, then export it for use as a starting point for new sites.

Staging/QA Testing

Copy catalogue data to a staging environment to test changes with real product data.

Disaster Recovery

Maintain offline backups of site configuration and catalogue data (note: for full backups including transactional data, use the S3 archive feature).

Multi-Environment Deployment

Promote catalogue changes from development through staging to production by exporting and importing.

Technical Details

Service Layer

The feature is implemented in the SiteTransferService interface with these operations:

public interface SiteTransferService {
    // Streams the ZIP straight into the provided stream off a Mongo cursor — driven by the worker,
    // which writes to a temp file and uploads it to S3. Memory-bounded regardless of catalogue size.
    Map<String, Long> writeSiteExport(SiteContext siteContext, OutputStream out) throws IOException;

    ImportValidationResult validateImport(InputStream zipFile);
    SiteImportResult importSiteData(SiteContext targetSite, InputStream zipFile);
}

Both directions are streaming and memory-bounded:

  • Export — the worker (ExportProcessingJob) picks up a pending SITE_TRANSFER_ZIP job, calls writeSiteExport to stream each entity collection (via a Mongo cursor + Jackson SequenceWriter) into a temp ZIP, then uploads it to S3 with RequestBody.fromFile.
  • Import — the uploaded ZIP is spooled to a temp file and each entry is deserialised with a Jackson MappingIterator, persisting in fixed-size batches. ID-remap tables hold only string pairs, so peak heap stays flat even for large catalogues.

Sensitive Data Filtering

Configuration keys are automatically filtered if they: - Are marked with ofSecretString() in ConfigKeyDefaults - Contain patterns like: apikey, apisecret, token, secret, password, credential, stripe, recaptcha

Import Order

Data is imported in dependency order to maintain referential integrity:

  1. Media (referenced by products)
  2. Taxonomies (referenced by products)
  3. Collections (referenced by products)
  4. Products
  5. Pages
  6. Menus
  7. Stereotypes
  8. Delivery Options
  9. Promotions
  10. Site Configuration

Limitations

  • Media files: Only metadata is exported, not the actual binary files. Media URLs are preserved but files must exist at those URLs.
  • Product images: @DBRef image references are cleared during import; re-associate images in admin if needed.
  • Global promotions: Only site-specific promotions are exported; global promotions are excluded.
  • User data: No user accounts are transferred; users must be created separately on the target site.

Security Considerations

  • Export/import operations require admin authentication
  • Password confirmation is required for both operations
  • Sensitive configuration values are automatically stripped
  • All PII and transactional data is excluded by design
  • Audit logging captures export/import operations