Back to Thoughts

How Affiliate Lead Generation Fraud Feeds KYC Data Fraud

Two industries that think they're unrelated are running on the same data, through an identically shaped supply chain. I sit in both, and the connection is hiding in plain sight.

By Simon Delaney • Published 2 June 2026

I sit in two worlds that most people think are unrelated.

In one, I watch affiliate lead generation fraud happen every single day through Databowl: bots, spoofed sessions, fake sign-ups. It isn't a once-a-month case study. It's constant. Knowing how it happens and how to stop it is the job.

In the other, through Provero (where we screen records in real time for validity, fraud, breach and provenance signals before they're trusted), I can see where breached data goes once it leaves the lead generation ecosystem. And it goes somewhere most people would never imagine: straight into the commercial databases that power KYC and IDV identity checks.

So here is the claim, stated plainly so it's hard to misread:

KYC data fraud and affiliate lead generation fraud are not similar problems. They are the same problem, running on the same data, through an identical supply chain. The lead generation fraud feeds the KYC fraud.

I laid out the full evidence for the KYC half of this in my deeper piece on the data provenance problem in commercial KYC. This article is about the join: why the two ecosystems are structurally identical, and why that's exactly what lets the fraud persist.

What affiliate lead generation fraud actually looks like

Strip away the jargon and it's simple. A fraudulent affiliate is paid for every sign-up they deliver to a form. So they manufacture sign-ups.

It's a bot running a script of data. It hits the form, spoofs the device, browser and operating system on the initial tracking so it looks like a real human session, and drops in a record. Most often that record is breached data, a real person's details, harvested from an old leak.

There's a second flavour that looks identical from the outside but has a different villain: a company that owns a form points a bot at its own form, manufacturing the impression of sign-ups it can wave around as "evidence" if anyone ever questions the numbers. Same fingerprints, different hand.

Either way, a fake sign-up is born, and a real person's data has just entered the machine without that person doing anything at all.

The laundering step: how fake data becomes "legitimate"

This is the part that should worry everyone, and it's the bit almost nobody traces.

That single sign-up doesn't stay in one place. At the point of sign-up it's opted into a list of co-sponsors. Each of those destinations then opts it into their own network of partners, and so on. One sign-up can cascade into hundreds of companies that each hold a legitimate-looking opt-in for that record.

By the time it's settled across the ecosystem, the original source is gone. Every company down the line can point to "an opt-in", so the data looks clean from every angle. That's data laundering: the fraud happens once at the top, and the legitimacy is manufactured by the chain beneath it. The kicker is that the consumer never actually opted in to any of it.

From there, the same records flow into commercial KYC databases, and the company feeding them in gets paid every time their data is used to make a match.

Why "our data is unique" is a myth

Ask a commercial data broker where their identity records come from and you'll often hear a version of the same pitch: tens of millions of records (say 40 million) drawn "from a range of sources", refreshed every month. Press on the refresh and the truth is usually mundane. The monthly updates are co-registration data, a couple of months old, bought in bulk. The "range of sources" is co-reg sites. In practice, a large share of the commercial identity data feeding KYC is repackaged co-reg.

That matters because of how co-reg data actually moves. Once you understand the sharing ecosystem, the idea that any broker holds genuinely unique data falls apart. A record collected at one sign-up is opted into a chain of co-sponsors. Each of those re-shares it under its own privacy policy, which contains opt-ins to the next set of companies, and so on. The data isn't held. It's circulated. Two "independent" commercial databases can easily contain the same record because it originated from the same co-reg event and fanned out to both.

This isn't only true of breached data. It's true of all data in this ecosystem. Everything is massively shared, and every company it touches shares it onward, because every company has an opt-in and a privacy policy of its own pointing at hundreds more.

I've seen the map of this. A journalist filed subject access requests against the companies that held an opt-in to his data, then followed each one to the company above it that had passed it on, and the trail kept climbing, company after company, all the way up to a single point where his breached record first appeared. Hundreds of businesses, one origin.

Network diagram showing one breached personal record spreading through hundreds of intermediary companies and data brokers, every trail leading back to a single origin point.
One journalist's subject access requests, mapped. Hundreds of companies each held a “legitimate” opt-in to his data, and every trail climbed back to a single origin where the record first appeared. Source: independent journalist investigation shared with Simon Delaney.

Now picture data laundering running through that structure. It's completely opaque and effectively impossible to trace. Go to the company that sent you a record and ask for the opt-in, and they hand you one, from the company above them in the chain. They all have one. The consumer, or the bot, "agreed" to the privacy policies of hundreds of companies in a single sign-up. Provenance doesn't survive the first hop.

This is the data going into commercial KYC databases. Not as an edge case, as the bulk of it.

This is not a theory. Here's what the data shows.

I don't ask anyone to take this on faith, because I've measured it.

One of the largest UK co-registration campaigns runs through Databowl, a household brand buying roughly 800,000 to 1 million co-sponsor leads a month, opted in to IDV, KYC and AML checks for downstream companies. I analysed 100,000 random accepted records from a single month. Not old rejected data, accepted, paid-for, real-time inflow that looked valid at the field level.

Underneath, the picture was different. The largest supplier had close to a 100% breach match rate. The second sat around 70%, with more evidence of synthetic construction. The third was around 90%. Across more than 80% of the data, breach exposure was extreme. And when I checked first breach date against date of birth, around 13% of people appeared not to have been born when their data was first breached, which is not possible for a real human record.

Breach presence on its own isn't proof that any one record is fraudulent. Plenty of legitimate people sit in old breaches. The signal is the pattern at supplier and cohort level, and the pattern here is not normal digital life.

One important caveat, because fairness matters here: this is not an indictment of co-registration as a channel, or of co-reg companies as a whole. Plenty of co-reg operators are entirely legitimate. The point is that the legitimate ones don't produce the impossible volumes, and so they don't interest the IDV, KYC and AML databases that fixate on volume. The problem isn't co-reg. It's mass-volume co-reg, which is exactly where this kind of breach-heavy, recycled inflow comes from. I set out the supplier-level detail behind that distinction in the data provenance problem in commercial KYC.

Compare it to cleaner sources. In high-quality data I analysed, breach exposure averaged around 40%, but the breaches were recent, averaging around three years old. In the breach-heavy co-sponsor data, the average breach age was around nine years. The market can produce scale. It can produce clean, well-provenanced data. It struggles to produce both at once.

Then there's the arithmetic, which is the fastest way to see this for yourself. To deliver one million opted-in co-reg sign-ups a month at industry-typical funnel rates, you'd need to send on the order of 833 million emails a month to UK adults. The UK has roughly 54 million adults. The numbers don't work without either an enormous, repeatedly hammered list or a very different inflow source. If an audience that genuinely engaged at that scale existed, it would be sold at premium rates by major brands, not as rock-bottom co-reg traffic.

What 1M sign-ups per month takes
1,000,000
sign-ups / month
→ 15% of clicks sign up — TARGET
6,700,000
clicks / month
→ 20% of opens click
33,300,000
email opens / month
→ 4% open rate
833,000,000
emails sent / month
→ implied send volume
833 million emails a month would, on its own, make this the UK's biggest advertising channel. To a population of roughly 54 million adults, that's around 15 emails every month for every single adult in the country, just to feed one buyer's million-a-month target.

Take your own lead volumes, or your own match volumes, and work backwards. Ask how many real, new, accurate people could plausibly sit behind those numbers. The gap is the fraud.

The full supplier-level analysis, the date-of-birth chronology work and the corroboration-collapse evidence are in the data provenance problem in commercial KYC.

Why nobody catches it at the top: the click layer versus the execution layer

It's a layer problem. Fraud detection watches one layer of the data. The fraud only shows on a different one.

If the breach exposure is this extreme, why does none of it get caught before it enters a CRM, let alone a KYC database? Because the checks at the top are pointed at the layer the bot is built to fake.

At the top of the chain, a fraudulent affiliate or a dodgy co-reg operator points bot traffic at a form. The bot spoofs the browser, the device and the operating system. It runs at specific times of day. Often there's no fraud detection on the form at all, and here's the part most people miss: even when there is, it usually doesn't help, because it's watching the wrong layer.

Form-level and tracking-level fraud checks sit at what I'd call the click layer, the metadata of the visit: device, browser, OS, IP, referrer, timing. That's precisely the layer the bot is built to fake, and it fakes it convincingly. Meanwhile the PII it drops in is real breached data belonging to real people, so it sails through field validation: valid email, valid phone, plausible name and address. On every signal the click layer can see, it looks human and it looks real.

Databowl sees a different layer, because it's a separate system from the tracking the fraudster is busy deceiving. We see the execution layer, the actual mechanical behaviour of the system running the bot. And bots are mechanical, so they always leave a pattern. Identical device-and-browser fingerprints repeated across record after record. The same referrers. Tightly clustered or rhythmically timed submissions. The same record appearing twice. IPs cycling in ways no organic audience produces. None of it is visible if you're only reading the click-layer metadata. All of it is obvious once you're watching execution.

Tracking view a form or site owner sees: clean, accepted sign-ups with varied IPs, devices, browsers and UK locations, with no obvious signs of fraud.Execution-layer view of the same sign-ups, revealing identical device fingerprints, repeating referrers and duplicate consumer records that expose automated bot traffic.
Above: what the form or site owner sees, clean, accepted, unremarkable sign-ups. Below: what we see at the execution layer, the mechanical signature of automated traffic. Same sign-ups, two completely different stories.

In the execution-layer view above, the tells are mechanical and repeat row after row:

  • Identical device, OS and browser on every record: the same mobile, Android and Chrome fingerprint across submissions that should belong to different people.
  • Repeating referrers: the same handful of source URLs feeding record after record.
  • Clustered, rhythmic timestamps: submissions arriving in tight, machine-paced bursts rather than the scatter of real human behaviour.
  • Duplicate consumer records: the same person ID appearing more than once, seconds apart.
  • Every row marked "Accepted", because the site owner's click-layer view saw nothing wrong with any of it.

Individually, any one of those could be coincidence. Together, repeating at volume, they're the fingerprint of a single automated system running a script, the kind of trace you can only see if you're watching execution, not the click.

That's the whole trick, in one frame. The metadata is spoofed, so it passes. The PII is breached but genuine, so it passes. The record is accepted, paid for, and pushed downstream wearing every marker of a real human action.

By the time it reaches KYC, it has enormous legitimacy. It is a real person. It passes verification with flying colours. It checks out against other databases, which, as we've seen, often hold the very same record by a different route. The only two things wrong with it are the two things no one downstream can see: the data is old, and the human didn't do it. A bot did.

The identical machine

The two supply chains have the same shape: client, tier-one provider, an opaque middle layer where the chain disappears, then the consumer. That shared structure is why the fraud is so hard to catch and so persistent: it has the same blind spot in the same place on both sides.

Affiliate lead generation

  1. Client.
  2. Tier-one affiliate network.
  3. Sub-networks. This is where the chain starts to disappear. The sub-networks chase the cheapest leads they can find. Fraud enters here, and provenance dies at the last reputable-sounding vendor, where there could be ten more companies below, who knows?
  4. Consumer (real, breached, or synthetic).

Commercial KYC data

  1. Client.
  2. Tier-one KYC provider.
  3. Commercial KYC data providers. This is where the chain starts to disappear. The commercial data providers chase the cheapest records at the highest possible volume. Fraud enters here, and provenance dies at the last reputable-sounding vendor, where there could be ten more companies below, who knows?
  4. Consumer (real, breached, or synthetic).

Same set-up. Same blind spot at step three. Same death of provenance.

One honest caveat: in affiliate lead gen we have direct evidence that the sub-network itself is sometimes the party committing the fraud. I can't claim that about commercial KYC data providers. I don't have that evidence. But that's exactly what makes a dark supply chain so durable: the top layer doesn't have to be malicious for the bottom layer to be corrupt. The structure is identical, which means the structural vulnerability is identical, and it exists because nobody can see down the pipe, not because of who's standing at the top of it.

Why nobody stops it

Because everyone in both chains is optimising for the same thing: volume, at the lowest cost. Every party is incentivised to scale as far as the client's budget allows. The assumption that all these consumers are real gets quietly taken for granted, and a blind eye gets turned to impossible volumes because of the money moving through the pipe.

The tell is in what each industry celebrates as success. In affiliate lead gen, a lead is treated as a win. In commercial KYC, a match rate is treated as a win. Both are vanity metrics if the underlying data is fraudulent. A lead that's a bot is still counted. A match against a nine-year-old breached record is still counted.

Nobody is asking the one question that actually matters: where does all this data come from?

In affiliate lead gen we catch this through Databowl for the clients we work with. But for every client catching it, there are probably a hundred who aren't, and left unchecked, the fraud rate through sub-networks can run upwards of 60%. In KYC, as far as I can see, almost no one is catching it at all.

Why corroboration isn't what KYC thinks it is

This is where the supply problem turns into a verification problem. When the same identity turns up across multiple commercial databases, the KYC market treats it as corroboration, independent confirmation that the person is real.

But we already know this data is distributed, not independently collected. So that "corroboration" is often the same laundered record arriving by two routes and being counted twice, one origin, not two witnesses. A breached identity that's been circulated widely enough can end up looking more verifiable than a genuine person with a thinner commercial footprint, which is the exact opposite of what an identity system should reward. The full mechanics are in the data provenance problem in commercial KYC.

Compliance can't catch this. That's the whole problem.

There's a comforting assumption underneath all of this: that even if the technology misses it, compliance will catch it. It won't, and not because compliance officers are bad at their jobs. Because they aren't equipped to see this kind of evidence at all.

Picture taking a local bobby to a crime scene and asking what the robber left behind. He's competent, he's thorough, he checks everything he's trained and equipped to check. But the only trace in the room is a fingerprint that needs a forensic specialist with a UV light to find. The bobby isn't going to catch it. Not because he's careless, because he doesn't have the lamp, and nobody told him the only evidence is the kind you can't see without one.

That's compliance, here. A compliance review reads the paper: the last reputable-looking vendor in the chain, how the PII was apparently collected, the terms and warranties everyone signed. Laundered data passes every one of those, because the entire point of the laundering is that the paper is immaculate. There's an opt-in. There's a privacy policy. There's a supplier who warranted the data was clean, relying on the warranty of the supplier above them, who relied on the one above them. On paper it is 100% legitimate.

What would actually expose it, the execution-layer behaviour, the breach exposure, the impossible volume, is precisely the evidence compliance is not equipped to see. It isn't in the remit, and it isn't in the data they're handed. Nobody hands a compliance officer the 833-million-emails arithmetic and asks them to think there is no way this many real people exist. They're handed a contract and a feed.

It's the same blind spot as the click layer, one level up. The click layer can't see the bot, because it only reads spoofable metadata. Compliance can't see the laundering, because it only reads the paper. In both cases the thing doing the checking is looking exactly where the fraud has already been made to look clean, and has no instrument pointed at where the record was actually born. That's not a failure of effort. It's a failure of equipment, and it's why the fraud is so durable: it survives the one check everyone assumes will stop it, precisely because that check was never built to look in the right place.

And KYC has made itself worse

Affiliate lead gen at least has the instinct that more sources means less exposure. Commercial KYC has built an incentive that points the opposite way.

KYC providers typically pay per API call, not per match. So the rational move is to reduce the number of databases they call, because each call costs money. The whole industry trend is consolidation, funnelling identity checks into a single commercial KYC database to keep call costs down.

Think about what that means in lead generation terms. It's the equivalent of running all your affiliate lead generation through one sub-network, with zero knowledge of where the data comes from. And we already know, from years of evidence, that the sub-network layer is exactly where the fraud lives. KYC is consolidating toward the danger, not away from it, and calling it efficiency.

The fix is the same in both worlds

If the machine is identical, so is the repair. It comes down to three moves that reinforce each other.

Spread across more sources, not fewer. Diversification is a defence, not a cost to be minimised. Multiple sources give you something to benchmark against, and if one source is bad, it's limited exposure rather than total exposure.

Shorten the chain to each source. More sources sideways, fewer reputable-sounding vendors to hide behind between you and the consumer. The gap between the end client and the real person should be as close to direct as humanly possible.

Pay for results, not requests. Cost-per-match instead of cost-per-API-call removes the financial penalty for doing the first two. It's the change that makes the right architecture affordable instead of aspirational. (This mainly bites in mature, high-volume markets. In low-volume countries, per-call pricing can be perfectly reasonable.)

Underneath all three: transparency of data provenance, with proof. Not "we have an opt-in". Proof of where the record came from, when, and proof it's real. That's what turns a wide net of sources from extra exposure into extra protection, because now you can see which source is the problem, and cut it.

And the timing is non-negotiable. Once a record is through the door and laundered across the chain, it's too late. You can't un-trust data that hundreds of companies now vouch for. The checkpoint belongs at creation, before the record becomes evidence, not after the ecosystem has already legitimised it.

The KYC version of this has a name: match plus provenance. Stop returning naked matches, and start returning the source class, freshness, independence and confidence behind every match. I've set that argument out in full in the data provenance problem in commercial KYC.

Until that happens, KYC data fraud and affiliate lead generation fraud are the same thing they've always been: a free hall pass for fraudsters, dressed up as a verified customer.

If you're celebrating your lead volume or your match rate, it's worth asking your vendors the uncomfortable question: where does this data come from, when did it arrive, and can you prove it?

FAQ

What is KYC data fraud?

KYC data fraud is the use of false, stolen, synthetic, outdated or poorly sourced personal data to pass identity checks or create the appearance that an identity has been verified. In this article, the focus is specifically on commercial data fraud: where a record appears trustworthy because it exists in one or more databases, even though its original source cannot be proven.

How does affiliate lead generation fraud connect to KYC fraud?

The same breached consumer records used to manufacture fake lead-generation sign-ups are redistributed through co-registration and broker chains until they reach commercial identity databases used in KYC and IDV. Because the original source is lost in the chain, the record arrives looking legitimate, and a KYC system can match against data that was fraudulent at origin.

Why are the two supply chains described as identical?

Both run from client, to tier-one provider, to an opaque middle layer, to consumer. In both, the chain disappears at the layer below the last reputable-sounding vendor, where parties chase the cheapest data at the highest volume. The structure is the same, so the structural vulnerability is the same.

Is breach data proof that a record is fraudulent?

No. Many legitimate people appear in old breaches because they've used the same email or phone for years. Breach presence is a pattern-level signal across suppliers and cohorts, not individual proof, and it becomes meaningful when breach rates and breach age are extreme at source level, or when chronology is impossible.

Why does paying per API call make KYC fraud worse?

Per-call pricing penalises checking more sources, so providers consolidate onto fewer databases to control cost. That recreates a single point of failure, the equivalent of running all lead generation through one unverified sub-network, which is exactly where fraud concentrates.

Why isn't bot-driven lead fraud caught at the form?

Most form and tracking fraud checks watch the click layer (device, browser, OS, IP, referrer, timing), which is exactly what a bot spoofs. The submitted PII is usually real breached data, so it also passes field validation. The mechanical signature of automation only shows at the execution layer: repeated fingerprints, duplicate records, clustered timing. If you're not watching that layer, the traffic looks human.

Does a broker holding millions of "unique" records mean the data is independent?

Usually not. Much commercial identity data is repackaged co-registration data that has been shared across hundreds of companies through opt-in chains. The same original record routinely ends up in several "independent" databases, so volume and cross-database presence are not evidence of independence or genuine collection.

Can compliance catch this kind of data fraud?

Generally no, and that's a capability gap, not a failing of compliance officers. Compliance verifies the paper trail: the last named vendor, the apparent collection method, the signed warranties and opt-ins. Laundered data has a flawless paper trail by design, so it passes. The evidence that would expose it (execution-layer behaviour, breach exposure, impossible volume) is exactly the kind of evidence compliance isn't equipped to see, much as a general officer can't lift a print that needs forensic equipment. The check everyone assumes will stop the fraud was never built to look where the fraud lives.

What is the fix?

More sources rather than fewer, shorter chains to each source, payment per match rather than per call, and provenance with proof applied at the point a record is created, before it becomes trusted data.

Should I trust co-registration sites?

Yes, but not blindly. Co-registration is not inherently fraudulent, and plenty of co-registration sites do a good job. The issue is a specific subset of high-volume suppliers that appear able to generate unrealistic volumes of opted-in consumer data at very low cost. The tell is volume. If a supplier can deliver numbers that do not make sense against the size of the audience, the conversion rate, the channel economics or the freshness of the data, that is the warning sign. Good co-registration should be able to show where the record came from, when it was collected, what the consumer saw, what they agreed to, and why the volume is plausible. If the answer is only "we have an opt-in", that is not enough. An opt-in without provenance is just paperwork. So yes, co-registration can be trusted when it is transparent, proportionate and provable. But if someone can suddenly produce vast volumes of cheap records that nobody else can find, there is usually a reason.

Read next

This argument first ran as a LinkedIn article, where it's also open for discussion. Join the conversation there.

About the author

Simon Delaney is the founder of Databowl and Provero. Databowl processes hundreds of millions of consumer leads per year for brands, agencies and networks. Provero provides low-latency data verification and fraud-prevention APIs focused on proving whether consumer data is valid, human and trustworthy at the point it is created. Working across both ends of the consumer-data chain is the reason this connection is visible at all.