Product data
How to Unify Supplier Product Data Without Losing Product Identity
To unify supplier product data safely, preserve each product's stable identifiers, map every source into one governed attribute model, retain the evidence behind each value and route conflicts for review. Do not merge records merely because their titles look similar. The useful outcome is one trustworthy product record that can support ecommerce, search and automation without losing the source identities needed for updates, stock, pricing or audit.
For a repeatable version of this process, explore M.I.A.I Product Intelligence.
Why supplier catalogues become difficult to trust
Two suppliers can describe the same kind of product in completely different ways. One sends “bottom roller,” another sends “lower roller,” and a third puts the machine model in a free-text note. Measurements may arrive in millimetres, centimetres or inches. Brand names acquire punctuation changes, colours use local names, and important fields are buried in titles because the spreadsheet has nowhere else to put them.
The problem becomes more serious when those files are imported repeatedly. A changed title may create a second product. A supplier SKU may be mistaken for the merchant's SKU. A blank column may erase an approved value. Stock and price may update correctly while the description, variant or compatibility relationship is attached to the wrong item.
Product intelligence begins by separating identity, attributes, relationships and commercial values. Those categories can then follow different ownership and review rules instead of being treated as one undifferentiated row.
Start with product identity, not product wording
A title is meant for people and can legitimately change. It is a poor primary key. Keep stable identifiers for the source record, the merchant record and every destination platform. These may include a supplier item number, internal product ID, SKU, GTIN, Shopify product ID, Shopify variant ID, NetSuite item ID or Sage stock-item reference.
Do not assume that all identifiers mean the same thing. A GTIN identifies a trade item according to its issuing rules; an internal SKU is controlled by the business; a Shopify product ID identifies the product container; and each purchasable variant has its own platform identity. Store the identifier type, value and issuing system rather than putting every code into a field labelled “part number.”
Google's Merchant Center specification requires a unique product ID, advises keeping it unchanged during updates and says the same product should retain the same ID across countries or languages. That principle is valuable beyond a feed: a stable identity allows descriptions, prices and attributes to change without breaking the connection to the underlying item.
- Source identity: which supplier record produced the value
- Business identity: the merchant's canonical product and SKU
- Trade identity: a GTIN or other recognised identifier when applicable
- Platform identity: the destination product and variant IDs
- Relationship identity: the reviewed link between a product, model, category or application
Create one governed attribute model
Before merging files, define the fields the business actually needs. Give each attribute a clear name, data type, unit, allowed values, ownership rule and validation rule. A diameter should be numeric with an explicit unit. A brand should reference a canonical brand entity. A yes-or-no property should not accept six spellings of “yes.”
Keep the model practical. Start with the attributes that affect purchasing, fit, discovery, compliance, stock or fulfilment. Supplier-only notes can remain as source evidence without becoming customer-facing fields. The purpose is not to create the largest schema; it is to make the important facts consistent enough to use.
Shopify models a product as a container with options and variants, where a variant represents a specific purchasable combination and carries values such as price, inventory and barcode. A supplier's flat spreadsheet therefore needs to be mapped deliberately into product-level and variant-level fields rather than copied column for column.
- List the decisions customers and staff need the data to support.
- Define canonical fields, units and controlled values for those decisions.
- Map each supplier column to a canonical field or an evidence-only field.
- Preserve the original value beside the normalised value.
- Validate required fields and identifier uniqueness before merging.
- Send unresolved conflicts to review instead of choosing silently.
Normalise values without destroying the source
Normalisation makes equivalent values comparable. It can standardise whitespace, case, punctuation, units, date formats and known vocabulary. “Stainless steel,” “SS” and a supplier's approved material code may map to one canonical material value, provided the mapping is documented and genuinely equivalent.
Never overwrite the raw source. Store the original value, normalised value, transformation rule, source, import time and confidence or review state. This makes mistakes reversible and gives a reviewer enough context to decide whether a proposed mapping is safe.
Unit conversion needs the same discipline. Keep the supplied measurement and unit, record the converted value and apply appropriate precision. Rounding a technical dimension for display must not alter the exact value used for fitment, manufacturing or purchasing.
Resolve duplicates with evidence, not title similarity
Potential duplicates should be scored from multiple signals: stable identifiers, manufacturer part numbers, brand, dimensions, variant structure, supplier relationships and other approved attributes. A shared title or similar description can identify candidates, but should not authorise a merge.
Define what a merge means. Sometimes two supplier rows represent the same trade item from different sources and can link to one canonical record. Sometimes they are equivalent alternatives that must remain separate products. Sometimes one row represents a parent product while another represents a purchasable variant. These are different relationships and should not be collapsed into a single assumption.
When evidence conflicts, keep both assertions with their sources and mark the field for review. A reviewer should see exactly which records disagree, the values involved, the evidence date and the downstream destinations affected by a decision.
Keep facts, relationships and commercial data separate
Product facts describe the item: material, dimensions, brand and technical attributes. Relationships connect it to categories, models, applications, accessories or alternatives. Commercial data covers price, availability, tax, supplier cost and fulfilment. Each group may have a different source of truth and update frequency.
For example, an ERP may own stock and price while an approved manufacturer source owns dimensions. A product-information workflow may own normalised titles and categories. Shopify may remain the storefront destination. Separating these responsibilities prevents a descriptive supplier file from overwriting live inventory or an inventory feed from removing approved product content.
Schema.org's Product vocabulary reflects this distinction by providing properties for product identifiers, brand, category, material, model and offers. A structured model does not prove a claim is correct, but it helps systems carry different kinds of product information without reducing everything to prose.
A concrete example: combining three undercarriage catalogues
Imagine a merchant receives three files of excavator undercarriage parts. The first uses manufacturer numbers, the second uses supplier SKUs, and the third describes machine applications in a notes column. All three contain rollers, idlers and sprockets, but their category names and dimensions differ.
The workflow imports each file into a staging area and assigns a source identity to every row. It maps category synonyms into reviewed canonical categories, converts measurements into a common unit while preserving the originals, and separates machine models from product titles. Exact identifier matches link records automatically; probable matches become review candidates.
A roller with the same manufacturer number and dimensions in two sources can be linked to one canonical product while retaining both supplier offers. A visually similar roller with a different bore measurement remains separate. A claimed model application with no supporting identifier or reviewed relationship is stored as unverified evidence and is not published as a fitment claim.
The approved record can then send storefront content to Shopify while keeping the Shopify product and variant IDs, operational values from NetSuite or Sage 200, and the evidence trail behind each enriched attribute. Later supplier updates match the correct source record rather than relying on whichever title happens to be present.
Publish changes through a controlled review queue
Group proposed changes by risk. Formatting and approved vocabulary mappings may be low risk. Identity changes, merged records, compatibility claims, dimensions, price and availability deserve stronger checks. A batch should show how many records will change, which fields are affected and which destinations will receive the update.
The reviewer needs the current value, proposed value, source evidence and reason for the change. Approval should apply to a defined record and destination, not grant a blank cheque for future imports. Failed or rejected records stay visible with a clear reason so the same error is not repeated on the next file.
Where an external system is the source of truth, Shopify documents a complete-state synchronisation workflow for ERP or PIM data and targeted mutations when Shopify owns the record. Choosing the correct direction matters because a full replacement and a field-level update have very different consequences.
Measure whether the product record became more useful
Count data-quality outcomes rather than the number of values generated. Useful measures include records with stable identity, required-attribute completion, duplicate candidates resolved, conflicts awaiting review, relationships with evidence and destination updates confirmed.
Then connect data improvements to real journeys. Can a customer filter on the attribute? Can search distinguish variants? Can staff reconcile a supplier update? Does the landing page match the product feed? Google warns that inaccurate, missing or conflicting product information can cause disapprovals, limited eligibility or incorrect displays, which makes feed diagnostics a useful quality signal rather than a separate marketing problem.
M.I.A.I Product Intelligence is built for this work: attribute normalisation, entity relationships, evidence-backed enrichment and data-quality review. The objective is consistent product knowledge that can support commerce, search and automation without disconnecting the answer from its source.
A practical product-data quality checklist
- Every record has a stable business identity and its source identities.
- Product-level and variant-level fields are mapped deliberately.
- Original values remain available beside normalised values.
- Units, controlled vocabulary and transformation rules are explicit.
- Duplicate candidates require evidence beyond similar titles.
- Conflicting claims remain visible until reviewed.
- Each field has an owner and an authorised update direction.
- Destination writes retain Shopify, NetSuite or Sage record identifiers.
- Published facts, feeds and landing pages agree.
- Every import produces a reviewable audit record.
AUTHORITATIVE SOURCES
Guidance used in this article
FREQUENTLY ASKED QUESTIONS
Questions about ecommerce integrations and AI search content
What is the difference between a SKU and a GTIN?
A SKU is an identifier controlled by a merchant or supplier. A GTIN is a trade-item identifier assigned under GS1 rules. Store the identifier type and issuing system so the values are not treated as interchangeable.
Can supplier products be merged when their titles match?
No. Matching titles can create a review candidate, but a safe merge needs stronger evidence such as recognised identifiers, manufacturer numbers, dimensions and reviewed relationships.
Should normalisation replace the supplier's original value?
No. Preserve the raw value and record the normalised value, rule, source and review state. That keeps the change explainable and reversible.
How should variants be handled when unifying data?
Map product-level facts separately from purchasable variant combinations. Retain each destination variant ID and ensure option values, SKU, barcode, price and inventory stay attached to the correct variant.
Can Product Intelligence work with Shopify, NetSuite and Sage 200?
Yes. Approved integrations can connect governed product information with Shopify, NetSuite and Sage 200 while preserving system ownership, destination identifiers and review controls.
