Product reference data: the invisible project that is changing the entire business analysis
- Claire Brunaud

- 2 days ago
- 5 min read

The dashboard is ready. Volumes are consolidated, trends are clearly visible, and teams can finally compare the performance of several distributors.
At least, that's how it appears.
Because behind a perfectly legible curve, a much less visible problem may be lurking: are the compared products truly the same? Has a reference named differently in the files been correctly matched? Have the packaging types been differentiated? Do the product families follow a common classification?
Without a reliable product database, a business analysis can appear accurate while relying on incomplete comparisons. The subject may seem technical, but its consequences are very real: poorly consolidated volumes, underestimated performance, unreliable comparisons, and business decisions made based on a partial view.
One product, multiple identities
Each distributor has its own way of describing and classifying products.
Based on the data provided, a product reference can be identified by a label, a distributor code, an EAN, a family, a sub-family, or a sales unit. The level of detail available also varies from one provider to another.
Therefore, the same reference can appear in several forms within the files:
with a shortened description at a distributor;
with a full trade name at another;
under an internal code that does not correspond to the code used by the manufacturer;
associated with a different product family;
expressed in units in one file and in kilograms in another.
Taken separately, each of these files can remain understandable. The problem arises when they need to be combined.
How can we tell if two lines correspond to the same product? How can we compare volumes if the units differ? How can we analyze a product range when the categories are not structured in the same way?
This is precisely the role of the product repository: to create a common structure to identify, describe and classify each reference in a consistent manner.
Product repositories are not limited to a list of codes
Sometimes the reference system is reduced to a table containing a product code and a label. In reality, it forms the backbone of the entire analysis.
It may include, in particular:
the internal identifier of the reference;
the different codes used by distributors;
the EAN when available;
the product's trade name;
its brand;
his family and sub-family;
its format or packaging;
its unit of measurement;
its status in the range.
The goal is not to accumulate information. It is to create a stable identity for each product, regardless of the source in which it appears.
This stability then allows for the accurate reconciliation of sell-out data from multiple distributors. Without it, teams risk working with duplicates, excluding certain lines, or grouping references that shouldn't be together.
Why do reference frame errors distort business analysis?
A mismatch may seem insignificant when it concerns a single line in a file. However, when repeated across multiple distributors, over several months, and with hundreds of references, it ultimately alters the interpretation of performance data.
Incomplete volumes or volumes distributed across several products
When a single item has multiple labels, its sales can be broken down into different lines.
The product then appears to perform less well than it actually does. Part of its volume may appear under its usual name, another part under a distributor code, and a third part under an older name.
For a Key Account Manager or regional manager, this fragmentation can lead to underestimating the importance of a reference before a sales meeting.
Comparisons between distributors are impossible.
Cross-analyses assume that the data use a common basis.
If one distributor classifies a product in one family and another in a different category, comparing performance by product range becomes unreliable. Teams no longer know whether they are observing a genuine sales gap or simply a difference in classification.
However, multi-distributor analysis is a core use of a sell-out data management solution. It requires prior harmonization of information with the manufacturer's product repositories.
Packaging types combined
Two similar references are not necessarily interchangeable.
Differences in format, paper weight, or packaging can correspond to distinct uses and customer types. Artificially grouping them together masks these differences. Conversely, separating them when they represent the same product creates duplicates.
In both cases, the analysis loses precision.
Decisions made based on a distorted photograph
A product may appear to have limited distribution when some of its sales have not been properly attributed. A product category may seem to decline because a product has been reclassified. An innovation may fall short of expectations if its new code has not been aligned with the old one.
These situations can influence the preparation for negotiations, the defense of a product range, the prioritization of field actions, or even the evaluation of a promotional operation.
The reference framework remains invisible in the final presentation. Yet, it determines the reliability of each displayed indicator.
Building a usable reference framework: where to begin?
The first step is to start from the manufacturer's internal reference framework.
This must become the common reference point. Each product thus has a reference identifier to which the codes and labels used by the different distributors can be linked.
Next, the discrepancies between the sources must be identified:
variations in wording;
different distributor codes;
reference changes;
heterogeneous sales units;
non-aligned families and subfamilies;
formats or packaging that are difficult to distinguish.
This work also requires defining rules. What should be done when a code changes? How should a discontinued reference be handled? At what level should a product range be analyzed? Should the distributor's classification be retained or should the manufacturer's classification be applied?
There is no one-size-fits-all answer. The right level of structuring depends on the expected business uses.
A sales department might want to track performance by brand and product category. A category manager will need a more detailed analysis of product assortments. A regional manager will primarily focus on identifying active or unavailable SKUs in each warehouse.
The framework must therefore be built based on the decisions that the teams wish to make.
A construction project that never completely stops
A product repository is not a file that is built once and then archived.
Product ranges evolve. New items are launched, some change their packaging, others are discontinued. Distributors may also modify their own codes or classification methods.
Without regular updates, the discrepancies gradually reappear.
Therefore, a simple governance structure is needed: who validates new correspondences? Who reports a code change? How do sales teams report anomalies? How frequently is the data checked?
This monitoring does not necessarily require a cumbersome process. It mainly requires that responsibility be clearly assigned and that the rules be shared.
What sales teams actually earn
The harmonization of reference frameworks is not an end in itself. It allows teams to work with more reliable indicators.
Volumes can be consolidated without duplication. Performance becomes comparable between distributors. Analyses by family, sub-family, or SKU are based on a common classification. Field teams, key account managers, and sales management use the same interpretation of results.
This foundation also facilitates meeting preparation. When a distributor and a supplier discuss a product reference, they know precisely which product, format, and scope are involved.
The discussion can then focus on the real issues: volume development, warehouse coverage, assortment performance, or end-user needs.
Make this job easier to maintain - Product reference data
Sell-out data is often transmitted in Excel files whose formats and structure vary depending on the distributor. Centralizing this data therefore requires prior harmonization work.
A solution like KaryonFood allows these different sources to be linked with the manufacturer's reference data to provide a shared understanding of performance. Teams can then analyze the results by product reference, category, warehouse, or customer segment without manually reconstructing the mappings for each new file.
The product repository will probably remain one of the least visible areas of commercial analysis.
Yet it is this very factor that determines whether two figures can truly be compared, added together, or used to make a decision. Therefore, before attempting to produce more indicators, it is worthwhile to verify that the products behind these indicators all speak the same language.




Comments