Reference Data Management vs Master Data Management

Reference Data Management vs Master Data Management

By Jacob Wilson

Reference Data Management (RDM) and Master Data Management (MDM) are routinely treated as the same discipline, or as one being a sub-topic of the other. They are related but different. The confusion starts with the terms reference data and master data, so that is where this article starts.

Master data

Master data is the core non-transactional data entities like customers, employees, products, suppliers, accounts etc. that participate in the activities and events of your business processes. It is the key "people", "places", "things", or "nouns" (e.g. account, policy etc.) that are involved in a business process. Master data tends to be small in volume but high in complexity, with many descriptive attributes.

Master data is among the most important data in your business, yet its core role means it is often duplicated across multiple systems, creating a management nightmare. For example, which version of a customer email address is the correct one? The record in the sales system, the manufacturing system, or finance system?

Reference data

Reference data is harder to pin down, because two competing definitions of the term tend to confuse the subject.

The narrow definition comes from the Data Management Association's DMBOK (Data Management Body of Knowledge, the profession's standard reference text). In this case, reference data is the lists of valid values behind the attributes of master data entities. It's discussed as a sub-topic of Master Data Management. It is the lists of valid country codes, currency codes, and status codes that sit behind the attributes of these master data entities. For example, the allowable country codes on a customer address (e.g. AU) is reference data. Operational systems will generally maintain their own lists of reference data.

The broader definition also includes the mappings, hierarchies, categorizations, bandings, and augmentation of data that determine how data from different systems is translated, grouped, rolled up, and assigned for the purposes of data integration, migration, warehousing, reporting, and analysis.

The difference between the two is ownership. For the narrow definition, an operational system is the system of record. For the broader definition, in most organizations, there is no system of record. It survives only in a sprawl of poorly governed spreadsheets, CSV files, and hard-coded transformation logic.

The broader definition is the one this article uses. It contains the narrow definition rather than replacing it: the valid value lists are still reference data, they are simply the part of it that an operational system already owns.

Reference data vs master data

Master data is the core entities. Reference data is the list of valid values behind the attributes of that entity, plus everything that maps, groups, and rolls those entities and values up across systems.

Master dataReference data
What it isThe core non-transactional entities: the "people", "places", "things", or "nouns" of a business processThe lists of valid values behind the attributes of those entities, plus the mappings, hierarchies, categorizations, bandings, and augmentation applied across systems
ExamplesCustomer, Employee, Product, Supplier, AccountCountry codes, currency codes, status codes, a legacy-to-target chart of accounts mapping, a GL account hierarchy, age bands
ShapeSmall in volume, high in complexity, many descriptive attributesSmall lists, mappings, categorizations etc, whose value comes from being agreed, versioned, and consistently applied
System of recordUsually several operational system holds the record, and that is the problemAn operational system for the valid value lists; for the mappings, hierarchies, categorizations and bandings, usually nothing
Where it lives between systemsDuplicated across sales, manufacturing, financeSpreadsheets, CSV files, and hard-coded rules

Master Data Management

Master Data Management is a discipline focused on matching records (across systems), deduplicating, and applying survivorship rules to generate a canonical "golden record" for each core master data record in your business. A "golden record" is a single authoritative trustworthy version of each master data record.

The two main functions of MDM are:

  • Entity/Identity Resolution. The practice of identifying multiple records from different systems as the same person, place or thing. It involves matching records across multiple attributes with matching rules and confidence scoring. In some cases, a partial match is forwarded to a human for adjudication.
  • Survivorship. If multiple records have conflicting values for an attribute (e.g. address, email address, phone number), which does the MDM system use in its golden record? This is where survivorship rules come into play. They determine which value "wins". Again, in some cases a conflict is forwarded to a human for adjudication.

When is MDM really critical? Some examples include post-merger identity or legal exposure if you pick the wrong surviving record (e.g. health or aged care). There can be real financial implications of not managing master data correctly.

Reference Data Management

Reference Data Management is the practice of authoring, owning, approving, storing, and distributing the code sets, mappings, categorizations, hierarchies, and bandings that every reporting, analytics, integration, and migration initiative depends on.

RDM encompasses several distinct practices:

  • List management. Maintaining an authoritative set of valid values for an attribute of a master data entity, for example, the valid position codes for an Employee.
  • Mapping. Relating one code set to another, for example, mapping a legacy chart of accounts to its replacement during an ERP migration.
  • Categorization. Assigning each member of a code set to one of a discrete, business-defined set of categories, for example grouping cost centers into functional areas for management reporting.
  • Augmentation. Enriching a code set with attributes the business needs but no source system holds, for example, a reporting label, an owner, or a regulatory flag.
  • Hierarchy management. Defining the multi-level structures that roll detail levels up to summary levels, for example, a GL account hierarchy or a product taxonomy.
  • Banding. Defining how a continuous value maps onto discrete categories through a set of contiguous ranges, the continuous-value counterpart to categorization, for example age bands or revenue tiers.

Across all six, RDM adds what spreadsheets cannot: versioning, approval workflow, distribution and a full audit history.

The first practice, list management, is the exception to the ownership rule above, but it sits with RDM in a few circumstances. Where operational systems have no adequate way to maintain a list, or where a common definition is required across several systems, RDM becomes the authoritative source. Separately, in fields like healthcare, banking, and stockbroking, external bodies maintain industry reference data. In these cases, the system of record sits outside the organization altogether, and RDM's job is to ingest, version, and distribute it to every system that needs it.

The two disciplines solve different problems

MDM works on the entities themselves. It takes the same customer, held five times in five systems, and resolves it into one golden record.

RDM works in the spaces between systems. It is concerned with the mapping, categorization, hierarchies etc. that augment and map key business entities across systems, and it earns its place in the data architecture as an independent capability by becoming the system of record for the reference data that no operational system holds.

An organization can have a mature MDM program and still have no idea which spreadsheet holds the current version of its cost centre to functional area mapping.

Why MDM tools don't handle reference data well

Follow the narrow definition to its conclusion and you get the MDM tool market. If reference data is the lists of valid values behind the attributes of master data entities, and it is a sub-topic of Master Data Management, then the reference data capability of an MDM tool only needs to do one of the six practices above: list management, scoped to the attributes of the entities that tool already masters.

That is what these MDM tools generally do. The mappings, categorizations, hierarchies, bandings, and augmentation that data integration, migration, warehousing, reporting, and analysis actually depend on sit outside that view of the world, so they sit outside the MDM tool.

The result is familiar. Reference Data Management in many organizations is sporadically implemented as 100s of spreadsheets, csv files, and hard coded rules, including in organizations that have already bought an MDM platform.

Conclusion

Master data and reference data are different things, and Master Data Management and Reference Data Management are different disciplines with different systems of record. MDM resolves duplicate representations of your core entities into a golden record. RDM owns the code sets, mappings, categorizations, hierarchies, and bandings that no operational system owns.

Buying an MDM tool does not give you the second one. For that, a better option is a dedicated Reference Data Management system like TitanRDM.

Get the full picture download the whitepaper - RDM: The Missing Capability in Enterprise Data Architecture.