What actually arrives
Catalogue exports are almost always denormalised, and for a good reason: the system that produced them was storing a value against a product, not modelling a domain. An audiobook catalogue carries authors, narrators, ISBNs and chapter structures; the Lantern migration involved all four, and the recorded summary of the job is that it required more than moving product titles and prices. A supplier-driven catalogue carries brands, materials, care instructions and compliance data. In both cases the same fact is repeated across every product it applies to, with no identity of its own.
The Fabric Co.'s engagement put the general version of this plainly: source datasets are rarely normalised the way Shopify expects. That is not a complaint about suppliers. It is what a flat export is.
Shopify can model this, comfortably
The platform side of this is genuinely solved, and it is worth stating the numbers because they close off the question rather than opening it.
A shop gets 128 metaobject definitions on Basic, Shopify and Advanced, and 256 on Plus and Enterprise. Each installed app gets its own 128 on top. Every definition holds up to a million entries — a limit raised from the old 64,000 for non-Plus and 128,000 for Plus, with the plan distinction removed entirely. Standard definitions provided by Shopify do not count against any of it.
So: you can have an Author entity. You can have four hundred of them, or four hundred thousand. You can reference them from products through a metafield on the product whose value points at the metaobject, and edit a biography once rather than on every title. None of that is the hard part, and a migration plan that spends its time there is spending it in the wrong place.
So the question is how many you have
The hard part is that promoting a column to an entity requires you to know how many entities are in it, and a flat text column does not tell you.
Four hundred distinct strings is not four hundred authors. It is four hundred strings. Some are the same person typed differently, some are the same person with a middle initial on half their titles, and some are genuinely two different people with the same name. Counting the distinct values gives you a number that looks authoritative and is not, and the moment you create one metaobject per distinct string, that wrong number becomes the structure of the catalogue.
This is the same shape as the identity problems elsewhere in this work — matching a product across two stores, tying a product to the media that belongs to it, or keeping a record stable across an import that will certainly run again — except that in each of those there was an identifier to work from. Here there is not. The source system never issued one, because it never thought of an author as a thing.
A rule for what to promote
What you can automate is the triage: which columns are obviously per-product, which are obviously shared, and which cannot be judged until a person looks at them.
A column where nearly every value is unique is a metafield — it is an attribute of the product, and giving each value its own entity creates a thousand records that are referenced once each. A column where values repeat substantially is a metaobject candidate. And a column whose distinct count moves depending on how you normalise it is neither, yet, because the number the decision depends on is not knowable.
That third case is the one worth building for, and the useful behaviour is refusal. A function that reports “four hundred entries” when it means “somewhere between three hundred and forty and four hundred, depending on how you treat capitals” has done something worse than nothing.
/** A metaobject earns its place when the average value is used more than once.
* Below this, you have a per-product attribute wearing an entity's clothes. */
const MIN_REUSE = 2;
type Verdict =
| { kind: "skip"; reason: string }
| { kind: "metafield"; distinct: number; reuse: number }
| { kind: "metaobject"; entries: number; reuse: number }
| { kind: "needs-review"; collisions: { normalised: string; spellings: string[] }[] };
const norm = (s: string) => s.trim().toLowerCase().replace(/\s+/g, " ");
/**
* Decides what a denormalised source column should become in Shopify.
*
* Note what is NOT checked: the per-definition entry ceiling. A metaobject
* definition holds a million entries, so no real catalogue reaches it, and a
* check that never fires would imply the ceiling is the risk. It is not.
*/
export function shouldPromote(values: readonly string[]): Verdict {
const present = values.map((v) => v.trim()).filter((v) => v !== "");
if (present.length === 0) return { kind: "skip", reason: "column is empty" };
const groups = new Map<string, Set<string>>();
for (const v of present) {
const key = norm(v);
const spellings = groups.get(key) ?? new Set<string>();
spellings.add(v);
groups.set(key, spellings);
}
// Spelling collisions come first. While they exist the distinct count is a
// guess, so every number downstream of it would be a guess too.
const collisions = [...groups.entries()]
.filter(([, spellings]) => spellings.size > 1)
.map(([normalised, spellings]) => ({ normalised, spellings: [...spellings].sort() }))
.sort((a, b) => a.normalised.localeCompare(b.normalised));
if (collisions.length > 0) return { kind: "needs-review", collisions };
const distinct = groups.size;
const reuse = Number((present.length / distinct).toFixed(2));
return reuse >= MIN_REUSE
? { kind: "metaobject", entries: distinct, reuse }
: { kind: "metafield", distinct, reuse };
}The part that cannot be automated
Some normalisation is safe and some is a guess, and the difference is sharper than it looks.
Trimming whitespace is safe: a name with a trailing space and the same name without one are the same string with no information between them, and nothing is lost by collapsing them on the way in. Lowercasing for comparison is already a judgement — you can match on it, but you cannot tell which casing is the one to keep, and picking one silently means the catalogue will display somebody's name the way the import happened to see it first.
Past that it stops being automatable at all. A reordered name, an initial on some titles and not others, a translator credited as a narrator on one record, an imprint that changed its name halfway through the catalogue — all of these are the same entity to a human and different strings to any rule you can write. The honest architecture surfaces them for review rather than resolving them, and accepts that someone who knows the catalogue has to spend an afternoon on it.
That afternoon is the actual cost of the decision, and it is why the decision gets deferred.
Deciding late costs more than deciding wrong
Which is the trap, because deferring is the expensive option.
Import the column as plain text and everything works. The product pages render, the catalogue is live, and the structural question is still open — but it is now open across the live catalogue rather than across a spreadsheet. Promoting a text field to an entity afterwards means creating the entities, resolving the duplicates you avoided resolving the first time, rewriting every product's reference, and doing it on data that customers and staff have since edited.
The Fabric Co. engagement recorded the general lesson as deciding the Shopify catalogue model before transforming anything. This is the strongest case of it. Mapping a column to the wrong structure is recoverable; mapping it before you have asked what is in it means the recovery happens later, with more rows and an audience.
When this needs an engineer
It does not need one when the data is genuinely per-product. Most columns in most exports are attributes, not entities, and a metafield is the correct and cheap answer. The many guides telling you to use a metaobject for your size chart and a metafield for your SKU are right, and nothing here contradicts them — for data you are about to create, the reuse question is easy because you control the answer.
It needs engineering when the data already exists in a shape somebody else chose, at a volume nobody can read. Tens of thousands of products, a column that is probably an entity, and no identifier anywhere in the source to resolve it by — that is a data-modelling job with an audit in front of it, and the audit has to happen before the import rather than after. That work sits with our other migrations and catalogue data engagements and with the 28,000-product catalogue that taught us to run it first. It is part of what Shopify migration services means when the data model is the project.
Send us the store and the symptom.
