Skip to content
heapbyte - A name of excellence

Architecture · 9 October 2026

Large catalogue migrations break on options, not files

A commercial upholstery distributor exports twenty-eight thousand fabrics from a legacy ERP to migrate to Shopify Plus. The engineering team scripts an export, converts the data into JSONL, and loads it via the GraphQL Admin API. The network requests return HTTP 200, and the initial products appear in the Shopify admin. But when the merchandising team audits the imported upholstery lines, chaos ensues: premium velvet fabrics that should have fifty colorways display only three, flame-retardant commercial vinyls have had their backing specifications scrambled with their widths, and on incremental reruns, thousands of previously imported variants quietly disappear from the store.

8 min read
Written by the HeapByte engineering team

shopify large catalog migration

Why migrations stall on modeling, not bandwidth

External inventory systems—particularly wholesale databases, ERPs like SAP and Exact Online, and multi-brand supplier feeds—do not organize product data around ecommerce customer journeys. They store manufacturing specifications: roll widths, rub test ratings, weave densities, fabric compositions, and dye lots.

When non-specialist teams attempt a Shopify migration, the standard reflex is a direct 1:1 column mapping. If the supplier spreadsheet has seven attribute columns, the migration script attempts to create seven product options.

In Shopify, that script fails on line one.

The 3-option ceiling after the variant increase

In October 2025, Shopify raised the maximum variant limit from 100 to 2,048 per product. That change eliminated the variant wall for merchants selling extensive size and colour combinations, but it created a widespread misconception: developers assumed that because Shopify could hold 2,048 variants, its attribute flexibility had expanded proportionally.

It had not. As we established when analysing when variants are the wrong model, Shopify's fundamental architecture remained strictly bound to three option axes: Option1, Option2, and Option3. A product can represent 2,048 distinct combinations, but it can only ever combine them along three named dimensions.

When migrating a supplier feed containing five or six attribute columns—for example, Pattern, Colour, Width, Material, Flame Retardancy, and Rub Count—a naive migration pipeline hits an immediate platform wall. Shopify rejects any product creation payload defining more than three options with MAXIMUM_OPTIONS_EXCEEDED. You cannot force six options into Shopify's native variant matrix.

Triaging options: merchandising axes versus technical specifications

To migrate large supplier feeds cleanly, the data pipeline must perform deterministic attribute triage before transforming records into Shopify payloads.

An attribute belongs as a product option if and only if it represents a physical choice the customer selects directly in the buying flow, changing the attribute modifies inventory availability or SKU tracking, and the merchant wants Shopify's native option dropdowns or visual swatches to display it.

Every other attribute is a technical specification, not a variant option. In a 28,000-product textile migration, the customer selects the Colour and the Cut Length. They do not select the Flame Retardancy or the Abrasion Rating—those are intrinsic physical properties of the fabric that apply to all colorways or inform B2B specification sheets.

Those technical properties must be systematically stripped from the variant matrix and routed into Shopify's native structured storage: typed metafields. As we detailed when analysing Shopify metaobject migrations, structured metadata belongs in typed fields rather than flattened into product tags or free-text descriptions. By demoting non-merchandised attributes into single_line_text_field or list.single_line_text_field metafields, the product complies with the 3-option ceiling while preserving rich, filterable specification data on the storefront.

The productSet variant replacement trap

Once option triage is established, the next hazard appears when running migration reruns and incremental corrections.

In modern Shopify engineering, the primary mutation for catalogue ingestion is productSet. Unlike legacy REST endpoints that updated individual variants, productSet is a declarative, state-based mutation where you provide the target product and the desired array of variants.

What many developers learn the hard way is that productSet treats the variants array as the complete and authoritative record of truth.

If an existing product in Shopify has 40 colorway variants, and your incremental migration script attempts to update pricing for the 5 variants that changed supplier cost this week, sending only those 5 variants to productSet permanently deletes the remaining 35 variants from Shopify. Those 35 variants do not simply become unlisted; their historical inventory history is detached, existing customer bookmarks break, and active line references in open orders lose their variant links.

To make large catalog migration reruns safe, the migration pipeline must query existing Shopify products by SKU or handle, retrieve their active variant IDs, and pass the complete, unified set of variants with their matching IDs in every productSet execution.

ts
export interface RawVariantRow {
  sku: string;
  price: string | number;
  barcode?: string;
  attributes: Record<string, string>;
}

export interface RawSupplierProduct {
  title: string;
  handle: string;
  variants: RawVariantRow[];
  technicalSpecs?: Record<string, string>;
}

export interface ExistingShopifyVariant {
  id: string;
  sku: string;
}

export interface NormalizedProductPayload {
  handle: string;
  title: string;
  productOptions: { name: string; values: { name: string }[] }[];
  variants: { id?: string; price: string; optionValues: { optionName: string; name: string }[] }[];
  metafields: { namespace: string; key: string; value: string; type: string }[];
  demotedOptionKeys: string[];
}

const MAX_SHOPIFY_OPTIONS = 3;

/**
 * Normalizes raw supplier catalog records into an authoritative Shopify productSet payload,
 * enforcing the 3-option ceiling and preserving existing variant IDs during updates.
 */
export function normalizeProductMigrationRecord(
  raw: RawSupplierProduct,
  priorityOptionKeys: string[] = [],
  existingVariants: ExistingShopifyVariant[] = []
): NormalizedProductPayload {
  const attributeKeysFound = new Set<string>();
  for (const v of raw.variants) {
    for (const key of Object.keys(v.attributes)) {
      if (key.trim()) attributeKeysFound.add(key.trim());
    }
  }

  // Triage attributes: keep priority merchandising options (max 3), demote the rest to metafields
  const sortedKeys: string[] = [];
  for (const pKey of priorityOptionKeys) {
    for (const key of attributeKeysFound) {
      if (key.toLowerCase() === pKey.toLowerCase() && !sortedKeys.includes(key)) sortedKeys.push(key);
    }
  }
  for (const key of attributeKeysFound) {
    if (!sortedKeys.includes(key)) sortedKeys.push(key);
  }

  const chosenOptionKeys = sortedKeys.slice(0, MAX_SHOPIFY_OPTIONS);
  const demotedOptionKeys = sortedKeys.slice(MAX_SHOPIFY_OPTIONS);

  const productOptions = chosenOptionKeys.map((name) => {
    const values = Array.from(new Set(raw.variants.map((v) => (v.attributes[name] || "Default").trim())));
    return { name, values: values.map((val) => ({ name: val })) };
  });

  const existingBySku = new Map(existingVariants.map((ev) => [ev.sku.trim().toLowerCase(), ev]));

  // Build complete variant list, preserving existing IDs to prevent productSet variant deletion
  const variants = raw.variants.map((v) => {
    const existing = existingBySku.get(v.sku.trim().toLowerCase());
    const optionValues = chosenOptionKeys.map((optKey) => ({
      optionName: optKey,
      name: (v.attributes[optKey] || "Default").trim(),
    }));

    return {
      ...(existing?.id ? { id: existing.id } : {}),
      price: parseFloat(String(v.price) || "0").toFixed(2),
      optionValues,
    };
  });

  // Convert demoted options into typed structured metafields
  const metafields = demotedOptionKeys.map((demotedKey) => {
    const uniqueValues = Array.from(new Set(raw.variants.map((v) => v.attributes[demotedKey]?.trim()).filter(Boolean)));
    const isSingle = uniqueValues.length <= 1;
    return {
      namespace: "custom",
      key: demotedKey.toLowerCase().replace(/[^a-z0-9_]/g, "_"),
      value: isSingle ? (uniqueValues[0] || "") : JSON.stringify(uniqueValues),
      type: isSingle ? "single_line_text_field" : "list.single_line_text_field",
    };
  });

  return { handle: raw.handle, title: raw.title, productOptions, variants, metafields, demotedOptionKeys };
}
The normaliser enforces the three-option limit before GraphQL serialization and routes excess attributes into typed metafields. Crucially, it maps existing variant IDs by SKU: because productSet is declarative, omitting existing IDs causes Shopify to destroy unmentioned variants on every incremental update.

Positional drift across batch supplier feeds

The final trap in large supplier feeds is positional and casing drift across disparate supplier exports.

Supplier A formats options as Colour and Size. Supplier B formats the same product category as Color and Dimensions. When an import pipeline batches rows into Shopify without an explicit schema dictionary, Shopify creates different option definitions on different products, fracturing the store's faceted search filters and collection navigation.

Furthermore, within a single product's variant rows, if row 1 specifies [Colour: Navy, Width: 54in] and row 14 specifies [Width: 54in, Colour: Navy], naive positional parsers will map the value 54in to the option Colour. The product imports with zero validation errors, but the storefront swatch displays 54in as a color choice.

A robust catalog migration pipeline must enforce canonical option dictionaries: standardizing names, enforcing identical option indexing across every variant row, and ensuring that casing differences are resolved deterministically before ingestion.

What this architecture does not handle

This approach provides a reliable data pipeline for large catalogue restructuring, but it has specific boundaries.

A migration pipeline structures and imports products; it is not an ongoing ERP synchronization service. Once the catalog is imported, operational inventory updates should be handled by dedicated webhook- or batch-driven sync routines rather than full catalogue rebuilds.

Normalizing options into Shopify ensures the variant model is sound, but displaying rich fabric swatches or material previews requires theme-level Liquid or storefront components mapped to appropriate image assets.

Finally, if a product genuinely requires customer configuration across four or more simultaneous axes that affect price dynamically, variants are the wrong model altogether; that requires a custom product configurator.

When this needs an engineer

If you are migrating a few hundred products from a standard Shopify store or exporting simple CSVs from WooCommerce, commercial tools like Matrixify or standard Shopify CSV import handles the job smoothly without custom development.

It requires data engineering when the dataset reaches wholesale enterprise scale: when catalogues exceed 20,000 products where manual verification and post-import cleanup are mathematically impossible; when supplier data contains arbitrary, sprawling attribute columns that exceed Shopify's 3-option ceiling; or when repeat import runs must update prices and inventory in place without triggering declarative variant deletion.

In our work on large catalogue migrations and catalogue data architecture, we build automated data pipelines that treat catalogue structure as a data-engineering discipline.

Send us the store and the symptom.

Insights

Apply this to your store.

An audit turns the general principle into a specific list of changes, ordered by what actually pays back.