M.I.A.I

Data discovery

How Do You Find Missing Product Specifications Without Copying Unreliable Supplier Data?

Find missing product specifications by starting with a field-level gap list, identifying each product with stable manufacturer identifiers, and searching only approved sources in a defined order. Capture the source page, retrieval date, exact evidence and proposed normalized value separately. A value should reach the catalogue only after its identity, unit, scope and source have been reviewed. If the manufacturer does not publish the value, or approved sources conflict, record that outcome instead of guessing or copying a convenient answer.

Voor een herhaalbare versie van dit proces, verken M.I.A.I Data Discovery.

Start with the catalogue gap, not a general web search

A request to complete the product data is too broad to control. First produce a gap table that names the affected product, the missing field, the market and language, the operational use of the value, and the consequence of getting it wrong. A missing marketing bullet is not the same risk as an incorrect bore size, electrical rating or compatible-machine range.

Prioritise gaps that block customer decisions, feed approval, filtering, fitment or fulfilment. Defer fields that are merely nice to have. This keeps the research workload bounded and prevents a discovery process from collecting large quantities of information that no approved workflow will use.

Define completion before research begins. For example, a dimensions task might require length, width and height in millimetres, the manufacturer's measurement convention, a source URL and a reviewer. A material task might allow a precise grade, a broad manufacturer term or an explicit not published outcome. The rule should be written before anyone sees the search results.

Identify the exact product before accepting a specification

A specification is useful only when it belongs to the correct product. Match the catalogue record to a manufacturer model, part number, GTIN or other approved identifier before researching its attributes. Titles and photographs can help a reviewer, but they are not reliable identity keys when suppliers reuse wording or imagery across sizes and generations.

Google Merchant Center's product-data specification says that brand, GTIN and manufacturer part number help identify products. It also warns merchants not to guess or invent GTINs or MPNs. The same discipline applies to research: never borrow a specification from a similar-looking item because its title happens to match.

Keep the identifier namespace with the value. The number 10025 may be a manufacturer part number, a supplier SKU or an internal ERP item. Record which organisation assigned it and, where relevant, the market, range or model generation in which it is valid.

  • Internal catalogue ID and exact record under review
  • Manufacturer or brand
  • Manufacturer part number, GTIN or approved model code
  • Supplier code as a secondary lookup key
  • Variant, pack size, market and model-generation qualifiers

Define an approved source order for every field family

Do not let the first search result become the authority. Define a source order that reflects who can legitimately state the fact. For a manufacturer's physical specification, the current manufacturer datasheet or official product page normally has greater authority than a distributor listing. For a supplier-specific lead time, the supplier may be authoritative. For an internal selling price, the commerce or ERP system remains the source even if a third-party page displays another amount.

An approved source list can include manufacturer websites, authenticated supplier portals, standards bodies, regulatory records and internally controlled files. Record permitted uses and licensing constraints. W3C data-on-the-web guidance recommends providing provenance, quality and version information and preserving identifiers; those practices make later review possible even when a page changes.

Search engines, marketplaces and comparison sites may help locate an official source, but their snippets and copied descriptions should not silently become evidence. If a secondary source is allowed, label it as secondary and define the extra corroboration or review it needs.

Search approved sources in deliberate layers

Begin with the exact manufacturer identifier on the approved manufacturer's domain. Then try the identifier with the missing field name, known document terms such as datasheet or technical specification, and approved model-family qualifiers. Only after exact searches should the process widen to a supplier code, a model range or a carefully reviewed title.

Use each layer to find a source document, not to manufacture confidence through repetition. Ten distributor pages may all copy the same incorrect table. Agreement among copied pages is not independent corroboration. Prefer the document closest to the organisation responsible for the product and keep its version or publication date when available.

Stop widening when identity becomes ambiguous. Several plausible products with different specifications are an exception to resolve, not an invitation to average the values or select the most common result.

  1. Search the exact manufacturer part number on the approved manufacturer domain.
  2. Add the specific missing field and any known model or market qualifier.
  3. Inspect official product pages, manuals, datasheets and revision notices.
  4. Use approved supplier sources only when the field belongs to the supplier or the source policy permits corroboration.
  5. Send ambiguous identities, inaccessible evidence and conflicting current sources to review.

Use sitemaps and structured pages as discovery aids, not proof

Large manufacturer sites can hide useful documents behind category navigation or JavaScript search. A published sitemap can help locate known product pages and files. Google describes a sitemap as a file that lists important pages or files and their relationships, while also noting that submission does not guarantee crawling. It is a discovery map, not a statement that every listed value is current or correct.

Likewise, structured data, page metadata and search indexes can expose identifiers and candidate URLs. They can shorten the route to the source, but the evidence must still be read in context. Confirm that the page describes the exact product and that the field is not inherited from a broader family, a different variant or an obsolete revision.

Preserve the final evidence URL rather than only the search query or sitemap URL. A reviewer must be able to return to the actual page or document that supports the proposed catalogue value.

Respect robots rules, terms, access controls and rate limits

Automated discovery must operate within the permissions granted by the source. RFC 9309 defines the Robots Exclusion Protocol through rules that crawlers are requested to honour. It also makes clear that robots rules are not access authorization. A page being crawlable does not grant permission to bypass a login, copy licensed content or ignore contractual restrictions.

Use an identifiable client where appropriate, respect allowed and disallowed paths, keep request rates bounded and stop on explicit denial or repeated service errors. Do not evade bot controls, rotate identities to defeat limits or reuse credentials outside the approved connection.

For authenticated supplier portals, store only the evidence and fields the agreement permits. Keep secrets out of research logs. If automated access is not approved, create a human research task rather than converting a technical obstacle into an unsafe workaround.

Capture raw evidence separately from the proposed value

A reviewable finding has at least two layers. The evidence layer records what the source actually showed: source organisation, URL or file, retrieved date, document version, exact product identity, original label, original value, unit, nearby qualification and an evidence excerpt or reference. The proposal layer contains the normalized field, proposed value, normalized unit, transformation rule and confidence state.

Do not overwrite the evidence when converting 2.5 inches to 63.5 millimetres or when mapping stainless steel 304 to an internal material vocabulary. Keep both. A reviewer can then check the calculation and decide whether the normalized term is precise enough for its intended use.

If a page changes later, the record should still explain why the catalogue value was accepted. W3C's provenance and versioning guidance is especially useful here: a URL alone is not a complete audit trail when the content at that URL can be revised.

Normalize units without changing the meaning

Unit conversion is not merely formatting. Confirm whether a dimension is nominal, maximum, packaged, assembled or measured in a particular orientation. Keep tolerances and qualifiers. Converting 2 inches to 50.8 millimetres is mathematically correct, but it may still be wrong for a catalogue field that expects a rounded nominal size of 50 millimetres.

Use deterministic conversion rules and approved vocabularies. Store the original and normalized values, conversion factor, rounding rule and field definition. Never infer a missing width by subtracting other dimensions, or derive an electrical rating from a related model, unless the business has a specific reviewed rule that permits that inference and labels it clearly.

Where the source supplies a range, retain the range. Where it supplies approximate wording, do not remove the approximation. A more precise-looking value can be less truthful than the source.

Treat absence, conflict and conditional values as useful outcomes

Research does not always produce a publishable value. Distinguish not found, not published, inaccessible, conflicting and not applicable. Those states tell the catalogue team what happened and prevent the same unproductive search from being repeated as if it had never occurred.

When approved sources conflict, keep both claims with their source and scope. Check revision dates, market, model generation, pack size and measurement definitions before deciding that the values truly disagree. The newer document should not automatically win if it describes a different variant or market.

A conditional value belongs with its condition. A maximum load may depend on mounting method; compatibility may depend on a serial range; a material may differ by colour or manufacturing period. Publishing the value without the condition creates a confident error.

A concrete example: filling gaps in an industrial pump catalogue

Imagine a catalogue of 4,000 pumps and service parts. Eight hundred records lack port size, seal material or maximum liquid temperature. Supplier spreadsheets contain some values, but their titles mix current and discontinued models, and several rows reuse a family photograph.

The team first creates a gap table by exact internal ID and manufacturer part number. It prioritises port size because customers use it to select fittings, then seal material and temperature because they affect suitability. For each field it defines an accepted unit, approved manufacturer sources and a review rule. Records without an exact manufacturer identity do not enter automated research.

The discovery process searches approved manufacturer domains by part number, follows current product pages and datasheets, and captures the document revision, raw field label and value. A sitemap helps find an archived PDF for one current model, but the PDF itself becomes the evidence. Supplier pages are used only to locate a manufacturer code or to state supplier-owned availability, not to override manufacturer specifications.

One datasheet lists a 1-inch port; another lists 25 millimetres for what appears to be the same family. The records reveal that the first is an NPT version and the second is a metric-thread variant. They remain separate. Another model has no published seal material, so its outcome is not published and it goes to a manufacturer-enquiry queue rather than receiving a copied family value.

After review, approved proposals are published in a small batch. The team reads the records back by immutable product ID, checks filters and product pages, and keeps the evidence link with each change. The catalogue becomes more complete without pretending that every gap had a trustworthy answer.

Prioritise exceptions by customer and operational risk

A review queue should not be a flat list. Rank findings by the business consequence of error, customer demand, number of affected products and evidence quality. A disputed safety limit or fitment condition should appear before a missing secondary marketing phrase.

Group repeatable exceptions. If 120 products lack a manufacturer identifier, the real task is identity repair before specification research. If one supplier repeatedly omits units, create a source-specific import rule and validation report. Discovery should expose systematic catalogue problems, not disguise them through one-off manual fixes.

Give every exception an owner and a next action: request manufacturer evidence, correct identity, approve a transformation rule, mark not applicable or leave the field unpublished. An unresolved item should not silently age into accepted truth.

Review and publish controlled batches

Preview proposed changes with the product ID, current value, proposed value, source, transformation and confidence. Require explicit approval for high-impact fields and broad rules. Keep the write set bounded so a mistaken mapping can be reversed without reconstructing the entire catalogue.

Publish by immutable system IDs, not title searches. Read each result back from the destination and compare the stored value, unit and product identity. Then inspect the customer surfaces that consume the data, including filters, product pages, feeds and fitment tools where relevant.

M.I.A.I Data Discovery is designed for approved-source discovery, evidence capture, structured extraction and review queues. Those capabilities help turn a defined research scope into structured evidence for human review. The business still decides which sources are approved, which transformations are safe and which values may be published.

Missing-specification research checklist

  • Create a field-level gap table with exact product IDs.
  • Prioritise fields by customer need and consequence of error.
  • Confirm manufacturer identity before accepting attributes.
  • Define accepted units, vocabularies and completion states.
  • Set an approved source order for each field family.
  • Use search, sitemaps and structured pages only to locate evidence.
  • Respect robots rules, terms, access controls and rate limits.
  • Capture source, version, date, raw value and qualifications.
  • Store normalization separately from raw evidence.
  • Keep absent, conflicting and conditional outcomes explicit.
  • Route material exceptions to an owned review queue.
  • Publish small approved batches by immutable ID and read them back.

AUTHORITAIRE BRONNEN

In dit artikel gebruikte richtsnoeren

VRAAGSTUKKEN

Vragen over ecommerce integraties en AI zoekinhoud

Which source should be checked first for a missing product specification?

Start with the current manufacturer page or datasheet for the exact identified product. Use a supplier source first only when the field is supplier-owned, such as that supplier's lead time, or when an approved source policy explicitly allows it.

Can a search-result snippet be used as evidence?

No. A snippet can help locate a source, but it may be truncated, stale or taken out of context. Open the actual approved page or document and capture the product identity, value, unit and qualifications there.

What should happen when the manufacturer does not publish the value?

Record not published, attach the research evidence and route the item to the appropriate enquiry or review queue. Do not copy a family value or invent a plausible specification.

Can an inferred value ever be published?

Only under a documented business rule that permits the inference, records the method and requires the appropriate review. Safety, compliance, fitment and other high-impact fields should normally remain unpublished without direct approved evidence.

How does M.I.A.I Data Discovery help?

M.I.A.I Data Discovery helps teams search approved sources, capture evidence, structure proposed values and organise review queues so missing product data can be investigated without hiding provenance or uncertainty.