Most KYC discussions start at the point of verification. They should start earlier, at the point where the identity data was first created.
I have spent years working where consumer data is created: lead generation, verification, routing and real-time data delivery. We built Databowl in 2015, and it now processes hundreds of millions of leads per year for brands, agencies and networks. Provero, the newer business, is a low-latency API layer focused on stopping fraud and proving that data is valid at the point it enters a system. Because of that work, I see consumer data before it reaches CRMs, identity graphs, KYC, Anti-Money Laundering (AML) or fraud systems.
The core question I keep coming back to is simple. Is this data real, and did the human behind it know it happened? That question matters more than it sounds, because the answer increasingly decides whether a commercial KYC match is meaningful or whether it is just two copies of the same compromised record bumping into each other.
To be clear, I am not talking about every form of KYC.
I am not talking about passport scans, biometric liveness, sanctions screening, PEP checks, bank account checks, utility data or electoral roll verification.
I am talking about the commercial identity data layer used in B2C KYC and IDV API matching: the databases and identity graphs that help decide whether a name, address, date of birth, email or phone number looks like a real and trusted person.
That distinction matters. The strongest primary identity sources may be banks, utilities, government systems, credit reference agencies and electoral roll data. The question in this article is different: when commercial KYC providers use broader commercial identity datasets to improve match rate, fill gaps, corroborate identity spines or support risk signals, where did that data originate?
Regulated use is not proven origin
KYC, AML and fraud prevention operate inside regulated and audited frameworks. That creates real confidence in how data is handled, but it can also hide a gap.
This is not a claim that all KYC checks rely on murky marketing data. They do not. It is a claim about the commercial data layer that often sits alongside stronger primary sources and is used to increase coverage, match rates, corroboration and risk intelligence.
Data suppliers warrant their sources contractually. Each link in the chain signs off on the link above. On paper, provenance appears to be covered all the way down.
In practice, contractual provenance only works if the party giving the warranty actually knows what they are warranting. In affiliate, co-reg and brokered data chains, that is often the problem. The publisher warrants to the network. The network warrants to the supplier. The supplier warrants to the KYC provider. Each link is relying on assurances from the link above.
If fraud or compromised data enters the chain several hops upstream, the downstream parties may never see the original collection event. They see a contract, a warranty, a supplier name and a data feed. They do not necessarily see the person, the form, the traffic source, the device, the consent event or the point where the record was first created.
That is the compliance gap. The regulatory wrapper proves how data is processed. The contracts prove that everyone agreed it was clean. Neither proves where the underlying data originally came from.
Contracts can transfer responsibility. They cannot manufacture visibility.
Provenance does not start at the last reputable-looking vendor in the chain. It starts at origin.
When a commercial KYC provider says a record matched, the natural next questions should be: what exactly did it match against, where did that data originate, why do we believe that source, is the source independent, and is there real evidence of consent and human collection?
- KYC
- AML
- Fraud prevention
- Audit
- Compliance wrapper
- Where the data was first collected
- Whether a human submitted it
- Consent evidence
- Source class
- Freshness
- Independence
"Regulated use is not the same thing as proven origin."
The identity spine commercial KYC matches
The commercial identity spine that most B2C KYC APIs check against is fairly consistent. It usually combines first name, last name, address, postcode and date of birth, with phone and email layered on top. If a consumer types the same spine into an onboarding journey and that combination appears in a commercial identity database, it can look like a strong match. That looks reassuring, but the provenance question sits underneath every one of those fields.
Where did that spine come from? How was it created? Did the person actually submit it? Was it collected directly, brokered, inferred, recycled, breached or re-permissioned? Two records with the same spine can have very different origin stories, and a commercial KYC API often cannot tell you which it is matching.
In many commercial KYC workflows, the end client decides the level of corroboration required. A lighter "one-source" match may treat one qualifying data source as enough to support an identity result. A stricter "two-source" or "two-by-two" approach requires the identity spine to be corroborated across more than one source, or across multiple matched attributes. But stricter matching only helps if the sources are genuinely independent. If two sources contain the same record because it has been redistributed through the same ecosystem, the match may look stronger than it really is.
This matters for one-source matching as much as it matters for two-source matching.
If a one-source match can rely on a commercial database, then a polluted commercial database can still produce a positive result. In that case, the problem is not only false corroboration across multiple sources. The problem is that a single compromised source may be enough to create confidence.
That creates a perverse possibility. A person whose identity has been breached, sold, duplicated, brokered and redistributed may be easier to match than a legitimate person whose data has never been breached, never been widely brokered, or has a thinner commercial footprint.
In other words, a fraudster using stale but widely circulated stolen identity data may sometimes look more "verifiable" than a genuine person with cleaner data.
That is the opposite of what a KYC system should reward.
Commercial identity graphs have history
Before GDPR, UK and European direct marketing databases were drowning in consumer data. It was common to see datasets claiming tens of millions of people with full name, address, telephone, email, date of birth and demographics. Consent standards were weaker, broker chains were longer, and plenty of companies claimed to own data they did not really own. Mediabowl was an agency we started in 2011. Our own early Mediabowl positioning talked about access to large numbers of emails and mobiles, which reflects the old market dynamic.

GDPR changed the economics of large-scale consumer data. It raised risk and reset consent expectations. But data did not disappear because the law changed. Some of it was reclassified, repurposed, laundered through distance, or moved into contexts where processing could be more easily defended, such as fraud prevention, AML and identity verification.
I am not claiming this is true of every commercial identity graph. I am saying it is a serious provenance question that buyers and regulators should be asking. Does old, poorly provenanced, pre-GDPR or legacy direct marketing data still form part of the commercial identity infrastructure used today? In many cases, nobody seems to be checking.
An identity graph is not a warehouse
Identity data decays constantly. People move. People die. People turn 18. People enter the country. People leave the country. People change email addresses. People change mobile numbers. National-scale identity graphs require constant fresh inflow just to stand still.
If you add together address movement, mortality, new adults, migration and other lifecycle changes, a meaningful percentage of the adult identity spine needs to be created, updated or removed each year. Email churn and mobile churn add further pressure. So a commercial identity graph is not a static warehouse. It is a machine that requires constant replenishment. The question is where the fresh inflow actually comes from.
of the UK adult identity spine needs to be created, updated or removed every year. Just to stand still.
- People move house.
- People die.
- People turn 18.
- People enter the country.
- People leave the country.
- Email and mobile numbers churn.
"A commercial identity graph is not a warehouse. It is a machine that requires constant fresh inflow."
The obvious primary sources are restricted
The strongest primary identity sources are banks, utilities, telecoms companies, government systems, credit reference agencies, large consumer platforms, finance companies and electoral roll data. Much of that data is regulated, restricted or contractually controlled, and not generally available as a commercial supply of fresh identity spines. Many of these organisations already work with tier-one KYC and IDV providers; they are not open commercial inflow sources for the wider market.
So the question this article asks is different. What other commercial B2C engines can lawfully generate enough fresh, downstream-usable identity records to support broad commercial identity coverage and high match rates? When you look honestly at the available answers, the list is much shorter than the volume of commercial inflow suggests.
Lead generation has the same incentive as KYC: volume
Lead generation, and specifically co-registration, a type of lead generation, is a major source of high-volume consumer data. Privacy policies often reveal where KYC, IDV, AML or fraud prevention companies are listed as downstream recipients or partners. Co-reg works because it uses incentives and high-volume traffic to collect data at scale.
Lead generation rewards volume. More sign-ups, more leads sold, more records routed downstream. Commercial KYC has a related incentive: match rate. Match rate sells, because it reduces manual review, smooths onboarding and increases pass rates. A KYC provider with more data can return more matches. When the market rewards volume on both sides, provenance becomes a cost unless someone forces it to matter.
What high-volume inflow looks like
One of the largest UK co-reg campaigns currently runs through Databowl. It is for a household brand, buying roughly 800,000 to 1 million co-sponsor leads per month, opted into IDV, KYC and AML checks for various downstream companies. Roughly 40% of records are duplicate for that client. After duplicates and other rejections, the client accepts roughly 250,000 to 350,000 records per month. The client asked us to analyse the data for this campaign.
* None of the biggest suppliers of co-registration data for the UK are UK companies, nor are they based in the UK.
It is important to note that the gross will go into IDV, KYC and AML databases. Not the net. So the whole 1 million. This is because corroboration, or duplicates, are seen as a positive in IDV. It is corroboration that a person exists in multiple datasets. But as we shall see, should this really be trusted?
I analysed 100,000 random accepted records from March. These were not old rejected records. They were accepted, paid-for, real-time inflow. At the field level the data looked valid. Underneath, the picture was different. The largest supplier had almost a 100% breach match rate. The second largest had around 70%, with more evidence of synthetic construction. The third sat between the two at around 90%, with the rest looking synthetic. Across more than 80% of the data, breach exposure was extremely high. This is why volume without provenance is dangerous: the more widely a compromised record circulates, the more "real" it can appear to a system optimised for matching.
Breach matching is not a fraud verdict at record level. It becomes useful when the pattern is extreme at supplier, cohort or source level.
It is important to note that I am not saying there is anything wrong with co-registration, or co-registration companies; I am pointing out the stats as they were analysed. And some co-registration companies are entirely legitimate, it is just they do not have the volume to interest IDV/KYC/AML databases that fixate on volume.
When I analysed first breach date against date of birth, around 13% of people in the accepted data appeared not to have been born when their data was first breached. That should not be possible for a real human record.
Important caveat: breach presence is not proof that any individual record is fraudulent. Plenty of legitimate people have old emails and old breaches. The issue is pattern-level evidence across suppliers, cohorts and inflow streams.
Clean data exists. It just doesn't scale.
I also analysed data from sources believed to be high quality. Across those sources, breach match rates averaged around 40%. High-intent form-fill data feeding directly into contact centres showed similar exposure. The difference was not really the percentage. It was the age of the breaches.
In breach-heavy co-sponsor data, the average breach age was around 9 years. In higher-quality data, the average was around 3 years. Normal digital life produces breach exposure. Around 40% breach presence may simply reflect real people using real emails in a world where breaches happen. But extremely old breach exposure at huge scale, combined with impossible date-of-birth chronology and synthetic-looking records, tells a different story. The market can produce scale. The market can produce clean, well-provenanced data. It struggles to produce both at the same time.
The commercial contradiction of a million sign-ups a month
Think about what it would actually take to generate one million genuine co-reg sign-ups every month from engaged UK adults. That kind of supplier would not be a small lead vendor. They would effectively own a national-scale direct-response media channel, with direct access to the UK adult population, repeat engagement at massive scale, measurable clicks and conversions, first-party data and performance-based acquisition that could move huge volumes of consumers every month.
To deliver 1M opted-in co-reg sign-ups a month at industry-typical funnel rates, you need to send on the order of 833 million emails every month to UK adults. The UK has roughly 54 million adults. The arithmetic does not work without either an enormous, repeatedly-emailed list or a very different inflow source.
If that audience truly existed, it would be monetised at premium rates by major brands. It would not usually be sold as low-cost co-reg traffic at rock-bottom CPLs. That does not prove fraud by itself. It is a strong commercial reason to interrogate the source.
One origin, many paths: why source independence collapses
In 2022 I was introduced to a journalist investigating illicit personal data sharing. They had used subject access requests to trace the path of a single record from an apparent breach event into a network of intermediaries and downstream data companies. Many of those end recipients were businesses selling data into the wider ecosystem.
This is why we see so many duplicates. And it is why corroboration between two different commercial KYC databases cannot be treated as two separate original events. It may be one original event, fraud or not, redistributed many times back into the ecosystem.

Once a record enters the consumer data ecosystem, source independence can collapse very quickly. The same record can be shared, sold, routed, enriched, re-permissioned, brokered and redistributed. Later, when the same identity appears across multiple providers, the KYC market may treat that as corroboration. But it may not be true independent corroboration. It may be one origin distributed through many paths.
"Duplicate distribution is not the same as independent evidence."
The affiliate fraudster: when fraudulent inflow passes verification
About six months before the talk, one of the world's largest agencies, working with one of the world's largest brands, asked my team to help with a sign-up data problem. The brand did not want a normal lead-gen landing page, so we put a Cloudflare-style pre-page in front of their site to verify traffic before it reached the brand domain. A long-standing publisher and sub-network appeared to be driving fraudulent traffic.
Before any confrontation, the leads from that sub-source were nearly 100% breach matched. After confrontation, the breach rate dropped to around 40%, but the remaining data looked synthetic and carried similar script-injection signals. The records still passed basic email verification. Deliverability does not mean human. A record can be deliverable and still be synthetic, breach-derived, fraudulently injected or commercially polluted.
By the time the pre-check page exposed the issue, the sub-source had already been paid a large amount. That is the financial gravity behind the behaviour. The pattern also mirrors KYC: the end client often has no idea what is happening several hops upstream.
This is further evidence that traffic used to generate sign-ups on co-registration sites should not be trusted without fraud detection and verification at the point of collection. That is not to say it is always the affiliate. It is to say that affiliate marketing has the means and capability to drive fraudulent leads that later surface in KYC and IDV identity graphs, and that the original provenance of such a record is effectively lost by the time a bot injects it onto a co-registration form.
Where does the pollution start and stop?
A separate client had expensive leads with missing dates of birth. The traffic looked human, not obviously bot-generated. Name demographics looked plausible. Email domains skewed older: Hotmail, Yahoo, Googlemail and similar. Breach matching came back at around 85%. Referrer analysis pointed to a company I knew, with a large database that feeds many KYC databases and also drives traffic itself.
The plausible theory was a loop. That company may hold a large amount of old, poor-quality data. It may receive feeds from co-reg sites being hit by bots injecting breach data. It may send real emails to breach-heavy data. And it may generate its own traffic back into co-reg and other campaigns. In that loop, stolen or polluted data can re-enter legitimate-looking campaigns and then walk into KYC databases as fresh consumer activity. The breach data becomes intertwined with the ecosystem so deeply that even legitimate traffic generation may be powered by compromised data.
Who actually creates the data?
The KYC provider is often not creating the data. The data supplier may not be creating it either, we don't know. Real origin may frequently sit with affiliates, publishers, networks, sub-sources and arbitrage partners. The problem is that the system encourages more and more data, so the incentive to inject breached, or synthetic data anywhere along this chain is massive.
And the incentive chain runs in one direction: advertisers want leads, KYC providers want match rate, data suppliers get paid for volume, affiliates get paid for sign-ups, and sub-sources optimise for the cheapest traffic that still gets paid.
The more distance between the consumer and the final buyer, the weaker the provenance. The real origin of the record is often not the reputable company at the end of the chain. It is the traffic source several layers upstream that nobody at the end of the chain has ever met.
The measured system still leaks
UK Finance's Annual Fraud Report 2025, covering 2024 data, reports 24,407 successful third-party card application fraud cases in 2024, costing £19.3m. That is one product category. KYC is used across many other onboarding journeys: bank accounts, loans, BNPL, telecoms, gambling, crypto, insurance and marketplaces. Comparable totals are not always published across those sectors.
Cards are mature, regulated and heavily monitored. They have established application checks, credit bureau checks, fraud controls, payment monitoring and loss reporting. If that environment still produces confirmed third-party application fraud, what is happening in sectors with weaker reporting, faster onboarding, thinner margins, less mature fraud controls or less consistent data sharing? 24,407 is not the market size. It is one measured leak.
UK Finance also notes that compromise of personal data continues to drive Card ID theft, and that data stolen from breaches can be used for months or years after the incident. If compromised personal data drives identity abuse, and breach data has a long tail, what exactly is feeding the commercial identity graphs that are supposed to detect that abuse?
The synthetic identity smokescreen
There is a lot of attention on synthetic identity fraud, and some of that attention is justified. LexisNexis Risk Solutions has reported that almost three million synthetic “Frankenstein” identities may already be in circulation in the UK, based on an analysis of more than 72 million consumer profiles.
That sounds alarming, and it is. But in the context of commercial KYC data provenance, it may also be a partial red herring.
Based on what we are seeing in high-volume commercial data inflows, three million may be a pitifully small number compared with the volume of real, breached, recycled or redistributed identity records moving through the wider data ecosystem.
That does not mean the LexisNexis figure is wrong. It means it may be measuring a different and more visible category of risk.
Synthetic identities are important because they are constructed. They look like fraud. They fit the fraud narrative. They give the market a clear enemy: fake people.
But the bigger commercial KYC problem may not be fake people. It may be real people’s data being used in ways those people never intended.
That makes much more sense operationally. Why invent an identity from scratch if you can steal, buy or recycle one that already has a name, address, date of birth, email, phone number and some history attached to it?
Why invent an identity from scratch if you can steal one?
A synthetic identity has to be built and nurtured. A breached identity already comes with plausibility.
And when that breached identity is repeatedly sold, routed, injected, matched, duplicated and redistributed, it can start to look less like stolen data and more like corroboration.
That is the real danger.
A fake identity may be easier to identify as fake. A real identity, stolen or recycled through enough commercial pathways, can start to look like independent confirmation.
Fraud can also be monetised at more than one point in the chain. A fraudster can monetise the back end by using compromised identity data to pass onboarding, open accounts, access credit, or commit fraud after verification. But the front end can also be monetised if breached or recycled records are injected into lead generation, co-registration, affiliate traffic or data-supply flows that ultimately feed commercial KYC and IDV databases.
In the worst version of the incentive model, the same polluted identity data can be monetised twice: once when it is sold or injected into the data ecosystem, and again when it helps someone pass a downstream identity check.
This is why the distinction matters. If the industry frames the problem mainly as synthetic identities, the proposed answer becomes synthetic identity detection. That is useful, but incomplete.
If the deeper issue is breach-derived real identity data being recycled into commercial identity graphs, then the answer has to include provenance, breach exposure, source independence, collection context and freshness.
The question is not only: is this identity synthetic?
The question is also: is this identity real but stolen, recycled, redistributed, or falsely corroborated?
AI scales the problem. Governance raises the bar.
A lot of people frame AI as the new identity threat because criminals use AI to mimic legitimate behaviour. That is true but incomplete. AI does not create identity data. AI uses identity data. If the underlying data layer is breach-derived, recycled, duplicated and widely circulated, AI is not simply mimicking real people. It is mimicking a polluted system. The system may struggle to tell the difference, because polluted records are already part of what it treats as real.
"AI is not the thing that stops identity verification working. AI is the thing that makes it harder to keep pretending provenance does not matter."
The EU AI Act increases expectations around data governance, logging, documentation, transparency, human oversight, accuracy, robustness and cybersecurity for high-risk AI systems (Articles 10, 12, 13, 14 and 15). I am not claiming every KYC API is automatically a high-risk AI system. The point is the direction of travel. As AI becomes embedded in fraud scoring, onboarding, risk ranking and identity decisioning, the governance bar is moving towards systems that have to explain not just their outputs, but the data foundations beneath those outputs. The honest question for any KYC stack is: why do you trust the data under the score?
This affects everything built on top of the data
This is not only a problem for traditional KYC onboarding checks.
If the underlying commercial identity graph is polluted, then anything built on top of that graph inherits the same problem. That includes age verification, fraud scoring, identity intelligence, device and identifier history, risk signals, affordability signals, behavioural trust scores, and any product that treats historical identity data as evidence of legitimacy.
Age verification is an obvious example. If a system is matching against commercial identity data that contains inconsistent, recycled or breach-derived records, then the date of birth becomes a weak point. In one dataset we analysed, when duplicate records appeared from one year to the next, around half of those duplicate identities had a different date of birth. That is not a small data-quality issue. It is a whack-a-mole problem with one of the core fields used to establish age and identity.
The same problem applies to the newer layer of identity intelligence products. These tools often rely on signals around identifiers: how long an email has existed, where a phone number has appeared, whether an address has history, whether a device or identifier looks established, whether a person appears across multiple sources, and whether the pattern looks consistent.
Those signals can be valuable. But they are only valuable if the underlying history is well provenanced.
If the history is made up of breach-derived records, affiliate-injected records, co-reg duplicates, synthetic-looking records, or the same original identity redistributed through multiple commercial paths, then the intelligence layer is just the same problem wearing a different hat.
the same problem wearing a different hat
A risk signal built on poorly provenanced data may still look sophisticated. It may have a score, a dashboard, a model and an API response. But if the underlying data is polluted, the product has not escaped the provenance problem. It has abstracted it.
That is why provenance has to travel with the match and with the signal. It is not enough to know that an identifier has history. We need to know what kind of history it has, where that history came from, whether the sources are independent, whether the chronology is plausible, and whether the record appears because a human created it or because it has been recycled through the data economy.
"Anything using this data inherits the provenance problem."
Do not run KYC blind
The fix does not require throwing away commercial identity data. It starts with something simpler: match plus provenance. Stop returning naked matches. A match should not just say "this person matched". It should say "this person matched against this class of data, from this type of source, with this level of provenance, freshness, independence and confidence".
"Stop returning naked matches."
- Strong identity evidence
- Weak commercial corroboration
- Risk signal only
- Should not count as proof
A match against directly collected, recent, well-permissioned data is not the same as a match against brokered, legacy, co-reg, affiliate, duplicated or breach-heavy data. KYC providers do not need perfect knowledge of every record, but they do need source-level testing. Useful questions per source include:
- When did this record first enter the graph?
- What source class did it come from?
- Is the source genuinely independent?
- Does the record appear across providers because it is corroborated, or because it has been redistributed?
- Is there evidence of human collection?
- Is there evidence of consent?
- Is there abnormal breach exposure?
- Is there abnormal breach age?
- Is there impossible chronology?
- Is there evidence of synthetic construction?
- Is the source fresh enough to be meaningful?
With that context, a match can be weighted: strong identity evidence, weak commercial corroboration, risk signal only, or something that should not count as positive proof at all. That is a much more honest output than a binary green tick.
The Tier 1 KYC and IDV pre-check
Any Tier 1 KYC or IDV provider relying on commercial identity data should be able to run a provenance and plausibility pre-check before giving full evidential weight to a match.
The pre-check is simple in principle: does this look like a real human record, and is it plausible that this person carried out this action? Applied consistently, this kind of check reduces the rate at which breach-derived or synthetic data is matched against itself and treated as corroboration.
A few practical examples. Does the name pattern fit the demographic distribution implied by the date of birth? For example, does the financial profile associated with an address look plausible for the product being applied for? Are there signals that the person's declared identity, address context, product choice and behavioural pattern fit together, or do they look artificially assembled? If the email is known, does its breach history align with the declared date of birth, or does the chronology imply the person was matched in a breach before they were old enough to have an account?
The same logic applies to age verification and identity intelligence. If a duplicate record reappears a year later with a different date of birth, that should not be treated as a minor inconsistency. It should be treated as a signal that the underlying identity spine may not be stable enough to carry evidential weight. A system that cannot tell whether the date of birth belongs to the person, the source, the supplier, or the fraudster should not treat the resulting match as clean proof.
The point is that a number of provenance and plausibility checks can be run before treating a match as positive identity evidence. Done well, they reduce the chance of one polluted record being matched against another polluted record and called confirmation.
A KYC system should not reward the identities that have circulated most widely. It should reward the identities whose provenance, freshness and source independence make them trustworthy.
Match plus provenance
The shift is straightforward to state and harder to do. Not match or no match. Match plus provenance. In a world of breach data, synthetic identities, affiliate traffic, AI-generated behaviour and recycled consumer records, the question can no longer simply be "does this data match?". It has to be "should this match be trusted?".
The question is no longer: does this data match?
The question is: should this match be trusted?
Commercial KYC has spent years optimising for coverage and match rate. The next phase has to optimise for provenance, independence and trust. Because the real risk is not that KYC systems fail to find a match. The real risk is that they find a match against data that should never have counted as proof in the first place.
Questions buyers should ask their KYC provider
If you buy commercial KYC, IDV or AML data services, the practical questions are straightforward:
- What proportion of your match data comes from directly collected sources?
- What proportion comes from co-reg, affiliate, brokered or partner-supplied data?
- Do you classify source types in the match response?
- Do you test suppliers for breach exposure, breach age and impossible chronology?
- Do you test whether sources are genuinely independent?
- Can you distinguish between corroboration and redistribution?
- Do you return provenance, freshness and source confidence alongside the match?
- Can a customer choose to downweight or exclude weaker source classes?
- How do you audit upstream data suppliers?
- What happens when a source produces high match rates but poor provenance signals?
- Are you measuring only synthetic identity risk, or also the risk of real identities being breached, recycled, redistributed and falsely corroborated?
- Do you test for date-of-birth instability across duplicate records?
- Do you distinguish between real-but-compromised identities and synthetic identities?
- Do your identity intelligence products inherit the same source provenance metadata as your KYC match products?
- Can customers see whether a risk signal is based on directly collected data, brokered data, co-reg data, affiliate data, or redistributed commercial data?
- Do you treat age verification differently when the underlying DOB history is inconsistent?
The answer does not need to be perfect. But if the provider cannot answer these questions at all, the buyer is running KYC blind.
FAQ
What is KYC data provenance?
KYC data provenance is the record of where the underlying identity data came from, how it was collected, whether the person consented, and how independent and fresh it is. It sits beneath the match itself and explains why the match should or should not be trusted.
What is the difference between a KYC match and a trusted KYC match?
A KYC match says two pieces of data line up. A trusted match adds provenance, freshness, source independence and a confidence weighting. Two records that match each other can both come from the same compromised source, which is not real corroboration.
What is an identity spine in commercial KYC?
The identity spine is the core combination of fields used to identify a person commercially: first name, last name, address, postcode and date of birth, with phone and email layered on top.
Why does match rate create a provenance problem?
Match rate is a commercial selling point, so providers are incentivised to maximise data volume. The more data a system holds, the more matches it can return, even if some of that data is poorly provenanced, breach-derived or recycled.
How can breach data affect KYC systems?
Breach data has a long tail and can re-enter the ecosystem through brokers, co-reg flows and affiliate traffic. If a commercial identity graph absorbs breach-derived data, KYC matches may be partly matching against records that originated from a compromise rather than from the person themselves.
Is breach presence proof that an individual record is fraudulent?
No. Many legitimate people appear in breaches because they have used the same email or phone number for years. Breach presence is a pattern-level signal, not individual proof of fraud.
Does breach data mean a person is fraudulent?
No. Many legitimate people appear in data breaches simply because they have used the same email address or phone number for years, and breaches happen across the wider digital economy. Breach presence on its own is not evidence that any individual record is fraudulent. It becomes meaningful as a pattern-level signal across suppliers, cohorts and inflow streams, particularly when combined with extreme breach age, impossible chronology or other signs of synthetic or recycled construction.
Why does lead generation matter to KYC data?
Lead generation, especially co-registration, is one of the few engines capable of producing high-volume, fresh-looking consumer identity records. Some of that data feeds, directly or indirectly, into commercial identity graphs used downstream by KYC, AML and fraud systems.
What is the difference between corroboration and redistribution in KYC data?
Corroboration means an identity has been independently confirmed by two or more separate original sources. Redistribution means the same original record has been sold, shared or routed through multiple intermediaries and now appears in several places. The two can look identical inside a commercial KYC database, but they are very different in evidential terms. If two sources contain the same record because it has been redistributed through the same ecosystem, the match may look stronger than it really is. True corroboration requires source independence. Redistribution does not.
What does AI change about KYC data provenance?
AI accelerates the ability to mimic legitimate behaviour, but it does not create identity data. If the underlying data layer is polluted, AI scales matching against polluted data. EU AI Act expectations around governance, logging, transparency, oversight and robustness push systems towards explaining the data beneath their decisions.
What should KYC providers return alongside a match?
At a minimum, the source class of the data, its freshness, evidence of independence, signals of provenance and a confidence weighting. A match should be readable as strong identity evidence, weak commercial corroboration, risk signal only, or not positive proof.
What is match plus provenance?
Match plus provenance means a KYC or IDV response should not simply say that a person matched. It should explain the class of data matched against, the source type, freshness, independence, breach exposure, collection context and confidence. The goal is to show whether the match should be treated as strong identity evidence, weak commercial corroboration, a risk signal, or not positive proof.
How do you know if a KYC match is trustworthy?
A trustworthy KYC match is one where the underlying data has known provenance, recent freshness, independent sources and evidence of human collection and consent. A naked match that simply says 'this person matched' does not tell you whether the underlying data was directly collected, brokered, sourced without knowing the origination, breach-derived or redistributed. A more honest match response classifies the source type, indicates freshness and independence, and weights the result as strong identity evidence, weak commercial corroboration, risk signal only, or not positive proof. Without that context, buyers cannot tell whether a match deserves the weight it is being given.
What should buyers ask their KYC or IDV provider?
Buyers should ask where the underlying data was first collected, whether sources are genuinely independent, how fresh the records are, what proportion of inflow comes from co-reg or affiliate channels, and how the provider tests for breach exposure, synthetic construction and impossible chronology.
Are breached identities a bigger KYC problem than synthetic identities?
Possibly yes, and this is one of the most underdiscussed risks in commercial KYC. Synthetic identities are a real fraud problem, but they may not be the main provenance problem inside commercial KYC data. Synthetic identities get a lot of industry attention because they fit a clear fraud narrative: fake people, constructed from scratch. But synthetic identities have to be built and nurtured. A breached real identity already comes with a full spine: name, address, date of birth, email, phone number and history. Based on patterns observed in high-volume commercial data inflows, the volume of breached, recycled or redistributed real identity records moving through the wider data ecosystem may be far larger than the headline synthetic identity numbers currently discussed in the market. That means commercial KYC systems need to test not only whether an identity is synthetic, but whether a real-looking identity has been stolen, recycled, redistributed or falsely corroborated. Synthetic identity detection is useful but incomplete.
Why might a fraudster look more verifiable than a real person?
Because commercial KYC matching rewards records that appear across multiple sources, and breached identity data tends to appear across many sources by the time it has been sold, brokered and redistributed. A person whose identity has been breached, sold, duplicated and circulated through the data ecosystem may be easier to match than a legitimate person whose data has never been widely brokered or has a thinner commercial footprint. In practice, this means a fraudster using stale but widely circulated stolen identity data can sometimes look more 'verifiable' than a genuine person with cleaner data. That is the opposite of what a KYC system should reward, and it is a direct consequence of optimising for match rate without testing for source independence and provenance.
Does this only affect KYC onboarding?
No. Any product built on top of commercial identity data can inherit the same provenance problem. That includes age verification, fraud scores, identity intelligence, identifier history, device intelligence, risk signals and onboarding decisioning. If the underlying data history is poorly provenanced, then the products built on top of it may simply turn weak provenance into a more sophisticated-looking score.
Why does age verification have the same provenance problem?
Age verification often relies on the same commercial identity spine, especially date of birth, name, address, email and phone. If the underlying data contains recycled, breached or inconsistent records, the date of birth may not be stable enough to carry evidential weight. In one dataset we analysed, around half of duplicate identities appearing from one year to the next had a different date of birth. That makes DOB instability a provenance issue, not just a formatting or data-quality issue.
In summary...
The evidence suggests that large parts of the commercial KYC data ecosystem may be turning recycled breach data into apparent independent verification, creating a pipeline where compromised records can be sold, matched, redistributed and monetised again.
Are we funding the very thing we’re supposed to be fighting?
About the author
Simon Delaney is the founder of Databowl and Provero. Databowl processes hundreds of millions of consumer leads per year for brands, agencies and networks. Provero provides low-latency data verification and fraud-prevention APIs focused on proving whether consumer data is valid, human and trustworthy at the point it is created.
Sources and further reading
- UK Finance, Annual Fraud Report 2025 (covering 2024 data).
- EU AI Act, Articles 10, 12, 13, 14 and 15: data governance, record keeping, transparency, human oversight, and accuracy, robustness and cybersecurity.
- GDPR and ICO guidance on lawful basis, legitimate interests and fraud prevention.
- LexisNexis Risk Solutions, “Three million ‘Frankenstein’ identities pose a multi-billion pound fraud threat”, 2024.
Internal links: see also The Lead Gen Lessons Behind The Verification Dividend, Do Data Breaches Drive Fraud?, Does a Bigger KYC Database Cut False Matches?, and Affiliate models, fraud, and why CPL needs a tougher spine.
