Practical Data: AI Doesn't Need More Data. It Needs an Identity!
I think one of the most overlooked parts of any data strategy is identity. We spend enormous amounts of time discussing platforms, pipelines, and AI, yet rarely stop to decide what makes something unique to our particular business consciously. I believe that decision is fundamental: if we can't agree on what something is, we can't reliably share, connect, or reason about it.
The Identity Problem Hiding Inside Your Data Strategy
Organizations spend an extraordinary amount of time discussing how to give AI more data: more sources, more documents, more history, and more context. Yet in many organizations, there is a much more basic problem underlying all of this. The business still cannot reliably identify a particular thing.
Take something as apparently simple as a customer. Sales may identify a customer using a CRM account number. Finance may have a billing account. Marketing may use an email address. The website uses one customer ID; the booking platform uses another; and the legacy system that's being retired may contain the identifier that half the downstream processes still depend on.
This is often treated as a data integration problem because that provides something technical to fix. Usually, it isn't. The underlying problem is that the business has never explicitly agreed on what makes one customer different from another.
This is not an abstract data management concern. It is an operating necessity. Identity is dealt with constantly outside the walls of businesses and is barely noticed because, for the most part, it works. People have Social Security numbers, driver's license numbers, and passport numbers. Cars have VINs. Books have ISBNs. Products have barcodes. These identifiers allow independent organizations and processes to refer to something with a reasonable expectation that everyone understands exactly which thing they mean.
Businesses need the same capability for the subjects that matter to them.
Defining a Subject Means Defining Uniqueness.
There is a fundamental consequence of subject-based data modeling that is often overlooked.
By choosing a subject, a business inherently defines its uniqueness.
The two decisions cannot really be separated. If Person exists as a subject within a business, there must be a way of determining whether two records represent the same Person or two different People. If Product is a subject, the business needs to know what makes one Product different from another. The same applies to an Organization, Account, Asset, Contract, Security, Trip, or any other subject the business recognizes.
This isn't a modeling preference. It is an axiom that results from defining the subject in the first place. A definition of Person that cannot determine whether two representations refer to the same Person is not yet a complete definition of Person. Attributes may have been described, but the subject itself has not been completely defined.
Once those rules have been agreed, every instance of that subject can receive its own identifier. That identifier becomes the practical expression of the business decision that this particular thing exists independently of all the other things like it.
Identity Becomes an Operating Contract.
Establishing that identifier also creates an implied contract across the organization.
Different functions do not need to agree about every attribute they maintain about something. They do not even need to interact with it in the same way. They do, however, need to agree that when they refer to a particular instance, they are talking about the same thing.
This distinction is important because business collaboration increasingly depends upon information moving between functions. A common identifier provides a stable point of reference through those interactions. Different groups can maintain distinct perspectives, processes, and attributes without losing agreement on the underlying subject.
It sounds obvious. In practice, it is something organizations repeatedly fail to establish.
The Primary Identifier Should Belong to the Business
Once a business has established what makes an instance of a subject unique, it should generate its own identifier for it. This can be considered the Primary Identifier.
The term is deliberate. This isn't a database primary key, and it isn't simply whichever identifier happens to come from the system currently considered authoritative. It is primary because it represents the organization's own definition of that unique subject.
That distinction matters.
If a customer identifier is a CRM account ID, the CRM effectively owns part of the organization's definition of Customer. If it is the customer number generated by a billing platform, the billing process owns part of that definition. If it is inherited from a twenty-year-old mainframe, because that is where customer records originated, a historical technology decision has quietly become part of the business's definition of identity.
None of these is a particularly durable foundation.
Business Identity Needs to Outlive Technology.
Systems get replaced. Companies merge. Operating models change. Processes are redesigned. Data moves from one platform to another. Today's system of record eventually becomes tomorrow's migration program.
The subjects themselves are far more persistent.
A Person remains a Person when the CRM changes. A Product remains a Product when the order management platform is replaced. An Asset does not become a different Asset because Finance moves to a new ERP.
The identifiers established for those subjects should have the same characteristics. They should be ubiquitous across the business and perpetual in nature, deliberately independent of the applications, processes, and organizational structures that exist today.
This is one of the important differences between identifying a business subject and simply assigning a technical key to a record. The former is intended to survive change. The latter usually exists within the lifecycle of the technology that created it.
Secondary Identifiers are Inevitable.
Creating a Primary Identifier does not remove all existing identifiers. Nor should it.
A single subject may have identifiers from operational platforms, payment providers, marketing systems, legacy applications, and external partners. Every new platform or business relationship can introduce another.
These can be considered Secondary Identifiers. They matter because they allow the organization to recognize the identifiers used by the systems and organizations with which it interacts. Still, none of them needs to become the organization's definition of the subject itself.
Instead, the relationship between each Secondary Identifier and the Primary Identifier can be maintained explicitly.
That crosswalk becomes increasingly valuable over time. A record arriving from a partner can contain the partner's identifier. An event arriving from a legacy system can contain its old identifier. A new platform can create another identifier tomorrow. As long as those identifiers can be resolved against the Primary Identifier, the organization has not lost its understanding of the underlying subject.
This is also why allowing source-system keys to spread too far through a data estate creates long-term problems. They look convenient when a system is first implemented. Ten years later, they have become invisible dependencies on systems and definitions that nobody intended to preserve.
Uniqueness is Contextual to the Business Model.
None of this means there is a universal definition of a Person, Product, Security, or any other subject that every organization should adopt.
Identity is contextual to the business model.
A passport provides a useful everyday example. A country can issue a passport according to its own rules for identifying people. Within that context, the passport identifier is legitimate. Move across national boundaries, however, and the model changes. One Person can legitimately hold passports issued by multiple countries.
Nobody necessarily modeled the Person incorrectly. The jurisdictions are applying their own definitions within their respective operating contexts.
Financial markets provide another example. An equities trading business needs to distinguish securities at the level required for trading. Two tradable instruments associated with the same issuing company may need to exist as distinct subjects because, for that business, the distinction materially changes what can be done with them.
A retail wealth advisory firm can have a different concern. It may primarily need to understand an investor's exposure to the issuing company. Its definition of the relevant subject, and therefore its definition of uniqueness, can legitimately be different.
Neither model is inherently wrong. They represent different businesses asking different questions of the same underlying world.
Crosswalk Complexity is not Necessarily Bad Data.
Different definitions of uniqueness also explain why mappings between organizations are not necessarily one-to-one.
One organization's Primary Identifier may map to several identifiers used by another organization, or several identifiers in the first organization may map to a single identifier in the second. In more complex cases, the relationship can legitimately be many-to-many.
That isn't automatically a data-quality problem that needs to be cleaned up. It may be revealing something important about how the two businesses define their subjects differently.
The mistake is not having different definitions. The mistake is failing to understand that the definitions are allowed to be different.
A well-designed crosswalk therefore does more than translate identifiers. It captures the relationship between different definitions of identity.
Identity Makes Relationships Understandable.
The importance of identity becomes even clearer when moving beyond individual subjects.
Much of the useful information within a business comes from understanding relationships. A Person owns an Account. An Organization employs a Person. A Customer purchases a Product. A Company issues a Security. A Supplier provides a Product. A Contract covers an Asset.
But a relationship cannot be represented reliably until the things at either end of it can be identified reliably.
This is where apparently small identity problems become much larger information problems. If two customer records cannot be confidently established as representing the same Person, the organization cannot confidently determine which Accounts belong to that Person. If Products cannot be consistently identified, it becomes difficult to reliably determine which Customers bought them. If Companies and Securities cannot be identified consistently, any resulting view of exposure is immediately compromised.
The relationships still exist in the real world. Failure to establish identity prevents the business from representing those relationships reliably in its data.
Guesswork is a poor foundation for any successful working business, let alone AI.
Subject-Based Modeling Provides the Durable Foundation
This is one of the reasons subject-based data modeling matters.
Applications are temporary. Business processes change. Organizational structures change. Roles change constantly. The underlying subjects are considerably more durable.
Consider a Person. That Person might initially appear to an organization as a Prospect. Later they become a Customer. They may subsequently become a former Customer and, at some point, perhaps even an Employee. Those are different roles the Person plays, and different relationships the Person has with the organization. They do not require the organization to continually create a new Person.
This distinction becomes increasingly important when constructing a coherent representation of a business from its data. If identity is organized around today's actors, processes, and applications, the information model inherits their lifecycle. If identity is established around the underlying subjects, the model can survive them.
The subject tells the organization what it believes exists. The uniqueness rules establish when one instance is different from another. The Primary Identifier allows the rest of the organization to refer to that decision consistently.
Only then can relationships between those subjects and the events that happen between them be reliably represented.
AI Should not have to Determine What the Business Already Knows.
This brings the discussion back to AI. Organizations increasingly talk about giving AI context. They build knowledge graphs, vector stores, semantic layers, ontologies, and increasingly elaborate retrieval architectures to help models understand their businesses. There is value in all of them, but they cannot compensate for ambiguity about the fundamental aspects of the business.
Before an AI system can reliably determine which products a customer owns, it needs to know which customer and which products are being discussed. Before it can reason about household exposure, it needs to understand which People belong to the household and which Accounts belong to those People. Before it can understand the relationship among a supplier, a product, and a contract, each needs an identity of its own.
Otherwise, AI is being asked to infer something the business should already know.
This is where much of the current conversation about AI and data starts in the wrong place. The challenge isn't simply getting more information into a model. It is about establishing enough structure around that information, so the model knows what it is actually about.
Identity has to Come Before AI.
That work starts well before AI. The business needs to define the subjects it recognizes, determine what makes each instance unique, and assign each one an identifier that will persist across the systems and processes surrounding it. External and system-generated identifiers can then be mapped back to those business identities, allowing relationships between subjects to be established with confidence.
When done well, this creates something considerably more valuable than another data source for AI. It creates a durable representation of the things the organization believes exist and a common language for referring to them across functional, process, and technology boundaries.
AI doesn't need more data. It needs an identity.
Is your data delivering value?
If your data is not yet delivering the value your organization expects, the problem may not be the technology. Contact us. We help organizations establish the data foundations that turn fragmented information into something the business can understand, trust, and use.
By