Procurement Data Engineering · 01

Coverage vs. Accuracy: What Procurement Data Quality Actually Depends On

Forty countries. Sixty-five portals. Millions of notices. Coverage is the number every procurement intelligence platform leads with — and it's the least useful measure of whether the data can be trusted.

About this series

This article opens ProcureDataLab's Procurement Data Engineering series, where we document the practical engineering problems behind procurement intelligence platforms: portal coverage, duplicate detection, schema changes, validation, and long-term scraper maintenance.

Coverage · Deduplication · Validation · Schema drift · Monitoring · Maintenance

Contents The Metric Every Platform Leads WithWhy Coverage Is the Easy Number to SellWhat a Coverage Number Doesn't Tell YouTED Isn't EuropeThe Same Tender, Counted TwiceThe Timestamp That Was Technically CorrectWhen the Scraper Returns ZeroWhy "Just Scrape It" BreaksWhere the Engineering Effort Should GoWeighing Coverage Against AccuracySix Questions, Not One NumberProof That Standardization Can WorkFrom Coverage to Decision ConfidenceFAQReferences

The Metric Every Platform Leads With

Visit almost any procurement intelligence platform's homepage and the first number you'll see isn't a client name or a case study. It's coverage.

40 countries. 65 portals. Millions of notices tracked. The number changes, but the pitch stays the same: bigger coverage, better platform.

Coverage is useful. It just isn't a measure of procurement data quality. After spending time building the pipelines that collect this data, rather than just consuming it, that's the distinction we keep coming back to: reach and reliability are different things, and reliability depends on more than any single number, accuracy included.

Monitoring forty portals doesn't prevent incorrect data from showing up. A platform can advertise complete EU coverage while quietly missing procurement activity that never reaches TED in the first place. It can look perfect on a dashboard and still hand a client a duplicate contract, a mistranslated field, or a deadline that's off by an hour. None of that shows up in a coverage number. It only shows up when someone compares a notice against the original publication and the two don't match.

Coverage
  • 40 countries
  • 65 portals
  • Millions of notices
Data quality
  • Complete
  • Accurate
  • Fresh
  • Unique
  • Reliable

That comparison, and the mismatch it sometimes turns up, is what this article is actually about.


Why Coverage Is the Easy Number to Sell

It's worth being fair to coverage first, because the instinct to lead with it isn't irrational.

Coverage is easy to count. It's easy to put in a sales deck, a homepage headline, or an investor update. It's easy to compare against a competitor: forty portals sounds better than fifteen, and that comparison takes about two seconds to make. None of that is true of data quality. You can't glance at a platform and know whether its deduplication logic works, or whether its timestamp parsing accounts for daylight saving time.

Most procurement intelligence buyers ask about portal count first, too, because it's the natural opening question. Coverage tells you whether a platform even has the data you need, before you get to whether that data is right. It rarely gets asked about second.

Treating it as the primary metric is where the trouble starts. Coverage answers "do you have this?" It says nothing about "can I trust what you have?" In procurement, where a single number often determines whether a bid gets submitted, a deadline gets met, or a supplier gets flagged, the second question is ultimately the more consequential one.

A coverage number can always go up. You add a portal, the number grows, the homepage gets updated. Trust doesn't move the same way. It builds slowly, across months of a client not having to double-check your data, and it can be undone by a single mismatched notice.


What a Coverage Number Doesn't Tell You

Here's the uncomfortable part: a platform can be technically monitoring a portal and still be wrong about what's on it. Procurement data quality fails in more than one direction. Data can be missing, duplicated, incorrect, or silently stale, and coverage only ever speaks to the first of those.

Coverage doesn't tell you whether duplicate notices were merged correctly when the same tender appears on two different portals. It doesn't tell you whether a timestamp published in one timezone was interpreted correctly for a buyer operating in another. It doesn't tell you whether a scraper silently stopped returning data three weeks ago while still reporting a clean run. And it doesn't tell you whether a portal's below-threshold notices, the ones that never make it onto the "complete" coverage map, are being missed entirely and invisibly.

Missing data and wrong data fail differently. You can't miss data you never knew existed. A contract that was never collected leaves no trace, nothing prompting anyone to go looking for it. Someone has to independently learn a notice should have existed before they can even notice it's absent.

An incorrect field works the opposite way. A client compares a notice they already know about against what your platform shows them, spots a difference, and the damage spreads beyond that one field. Once a platform gets caught being wrong once, people start silently re-checking everything else it's told them, and most of that re-checking never gets reported back. It just becomes the reason they trust the product a little less.

That's how a platform ends up fully "covered," and still wrong.


TED Isn't Europe: The Threshold Nobody Talks About

One of the most common misconceptions in procurement intelligence is that monitoring TED, the EU's central database for procurement notices, means monitoring European procurement, full stop. It doesn't.

TED primarily receives notices for contracts subject to EU-level publication requirements above the applicable threshold, and that threshold isn't one number. It varies by contract type (supplies, services, works) and by buyer: central government, sub-central authorities, and utilities can each fall under a different figure, and the whole schedule is revised every two years under the EU's WTO Government Procurement Agreement obligations. That variation is exactly the kind of detail a flat "TED equals European procurement" assumption erases. Below the applicable threshold, publication on TED isn't generally required.

That doesn't mean the contract doesn't exist. It means the contract may instead be published through a national or sub-national procurement system, in that system's own format, on its own schedule, depending on the rules that apply. BOAMP in France. TenderNed in the Netherlands. e-Vergabe in Germany. Each one runs independently, with its own field structure, its own quirks, and no obligation to mirror what's on TED.

A TED-only pipeline can genuinely, honestly believe its European coverage is complete. Nothing in its dataset contradicts that belief, because the contracts it's missing don't appear as errors. They simply never entered the system in the first place. No broken request, no failed parse, no red flag. The gap is invisible from the inside.

What a TED-only pipeline can and cannot see What a TED-Only Pipeline Can and Cannot See What TED coverage can see TED Above-threshold notices What a TED-only pipeline can't see National / sub-national systems Below-threshold / nationally published Coverage appears complete from inside the dataset, because the missing records were never collected in the first place.
Figure 1. What a TED-only pipeline can and can't see. The blind spot is invisible from inside the dataset itself.

Catching this gap takes looking for it on purpose: comparing what TED returns against what a single national portal publishes in the same week, for the same country, and checking whether the numbers line up. When they don't, TED usually isn't the one that's wrong. It's doing exactly what it was built to do. The dataset was never designed to be "all of European procurement" in the first place. That assumption gets added later, by the platforms consuming it.

Coverage, in other words, is only as complete as the sources a platform monitors, and the procurement activity those sources actually expose. The portals that don't show up in the count are, by definition, the ones nobody's looking for.


The Same Tender, Counted Twice

Duplication is a different failure mode from a missing gap, and it's arguably the hardest one to catch from the outside, because the platform doesn't look broken. It looks like it has more data than it actually does. The same multi-portal structure that creates the TED gap above creates this one too.

The same tender can, quite legitimately, appear in a procurement intelligence database twice: once from TED, once from the relevant national portal. Both records are accurate. Both describe a real contract. They just often don't look like the same contract. Field structures differ. Publication dates differ slightly depending on which system logged them first. Reported values can differ because of currency conversion, VAT treatment, or rounding. There isn't a universal identifier you can reliably depend on across TED and every national portal, and nothing in either dataset flags them as the same underlying tender.

When that duplication goes unmerged, contract counts become inflated in a way that's genuinely hard to notice internally. The dataset looks bigger, not wrong. Someone might search and find the same tender twice, or a total contract count might come out roughly double what an external, more careful source reports for the same period. That gap usually only surfaces once someone runs an outside crosscheck and the numbers don't reconcile.

Matching the same tender across two portals with no shared identifier Matching a Tender With No Universal ID TED record Buyer: Ministry of Transport Date: 2026-03-04, Value: €450,000 No shared ID field National portal record Buyer: Ministère des Transports Date: 04/03/2026, Value: 450 000 € No shared ID field Attribute matching: buyer + date + value + metadata → confidence score Same tender → merge, don't duplicate
Figure 2. Matching a tender across two portals with no shared identifier.

Without a shared identifier, the only reliable path is comparing attributes across records: buyer identity, publication timing, contract value, CPV classification, procurement procedure, and other metadata, with enough confidence to conclude two listings describe the same underlying contract even when the surface details don't match exactly. That matching logic isn't a minor implementation detail. It's the difference between a platform's contract count meaning something and being decorative.

Deduplication isn't really about making a database smaller. It's about taking the verification burden off every number a client sees, so they don't have to independently check it before using it.


The Timestamp That Was Technically Correct

Some of the most expensive bugs in procurement data aren't caused by anything crashing. They're caused by a value that's completely valid and just wrong in context.

Consider a hypothetical submission deadline for a public tender in Germany. The buyer intends the deadline in German local time, UTC+01:00. Now suppose the notice's XML instead specifies the timestamp with a UTC offset of +00:00. Every part of that timestamp is syntactically correct. It parses without error. It validates against the schema. Nothing about it looks broken.

It's also an hour off from what the buyer actually meant.

Syntactically valid ≠ semantically correct.

A UTC timestamp can be perfectly valid and still be wrong in context. It can describe a real, unambiguous instant while failing to represent the local deadline the buyer intended, unless the pipeline reading it validates what the offset was actually supposed to represent rather than treating the two as automatically equivalent.

This isn't the kind of bug that shows up in testing, because testing checks whether a timestamp parses, not whether it's been interpreted in the timezone the buyer had in mind. A field can pass every validation check a pipeline runs and still be functionally wrong. Nothing in the timestamp's own schema tells you whether the offset it carries reflects the buyer's actual intended local time, or simply whatever offset the notice's source system happened to write down.

An hour doesn't sound like much until you follow it to the consequence such a mismatch could cause. A bidder who submits at 4:59 PM, a minute before the real deadline, could be treated as late if the system is enforcing a cutoff of 3:59 PM instead. The timestamp itself never throws an error. Nothing in the pipeline flags it. Catching a discrepancy like this takes more than checking whether the timestamp parses. It requires validating the timestamp against the buyer's actual local conventions rather than trusting the offset printed in the source data.

Silent bugs, the ones that produce a plausible, validating, wrong answer, are more dangerous than the ones that throw an exception. Nothing prompts anyone to go looking for them until a real deadline has already been missed. That's the case for treating every timestamp field as a candidate for a second check, not just a successful parse.


When the Scraper Returns Zero and Nothing Complains

Some production failures are loud. A request fails, an exception fires, a log lights up red, someone gets paged. Those are, in a strange way, the easy ones.

The harder failures don't throw anything at all.

Many government procurement systems don't publish clear changelogs when their layout, schema, or access rules change. No notification, no versioning, often no visible signal that anything is different from yesterday. A scraper built against yesterday's structure can keep running against today's changed one. Technically succeeding. Structurally wrong.

Imagine a scraper that returned 1,200 notices yesterday. Today it returns zero. Every request still succeeds. The HTTP response is a clean 200. No errors are thrown, nothing crashes, and a monitoring setup built only to watch for failures sees nothing to flag. The dashboard stays green.

The path of a silent scraper failure, from clean run to client discovery The Silent Failure Path Scraper runs on schedule, against a portal whose layout changed overnight HTTP 200: request succeeds, no error thrown Parser completes without error, but the fields it's looking for have moved 0 notices returned: technically a valid, empty result Dashboard: all systems green, nothing in the pipeline knows anything is wrong
Figure 3. The path of a silent scraper failure, from a routine run to a client discovery weeks later.

Nobody gets paged for a dashboard that looks a little quieter than usual. That gap only closes when a client asks why a contract they know exists never showed up, which can be days or weeks after the failure started, long after the window to catch it quietly has closed.

Catching this requires monitoring built around a different question than "did this fail?" It has to ask "is this behaving like it normally does?" That means establishing an expected output range and flagging deviations from it, not just outright failures. A scraper doesn't have to return zero to be broken. It can go from 1,200 notices to 1,150 to 1,100, a decline gradual enough to look like normal variation, while quietly returning null buyer names on every record. That kind of drift is harder to catch than a hard failure, and in procurement integrations specifically, far more common.


Why Procurement Data Breaks the "Just Scrape It" Assumption

The difficulty with procurement data has less to do with collecting it than with the fact that it was never designed to be consistent in the first place.

TED alone shows how much variation exists inside a single, relatively well-structured source. Buyer information arrives as multilingual structured objects, not plain text. Notice and procedure types are internal codes, things like cn-standard or oth-single, meaningless without the right lookup table. When the EU rolled out its eForms standard, mandatory since October 2023, pipelines built around the legacy notice format had to be substantially rewritten just to keep working. And, as covered above, TED's own scope stops at the applicable threshold. A large share of European procurement activity is published only through national or sub-national systems, each maintained independently, in its own format, on its own schedule.

None of that inconsistency is a flaw exactly. Every portal, threshold, and notice format was built independently, by a different country, to solve a domestic problem, not to interoperate with anyone else's system. Consistency was never part of the brief.

The access layer varies just as much as the data itself. TED offers an open search API. No login, no API key, callable as-is. SAM.gov, the equivalent system in the United States, requires an API key up front even though the underlying data is public. Many national portals have thin or nonexistent API documentation, and some effectively require full browser automation to extract anything at all. A single client engagement can require all three approaches at once: a straightforward API call for one portal, authenticated API access for a second, full browser automation built from scratch for a third. Estimating the effort for portal one tells you almost nothing about what portal seven will actually require.

That unpredictability compounds over time, not just across portals. A fully functional scraper against one national portal held up for less than a week before the portal added an OCR-based verification step, which took real engineering effort to work around. Four days after that, the site's underlying structure changed completely, and most of the integration had to be rebuilt from scratch. Two unrelated changes, back to back, inside two weeks, on a single source. Whatever estimate held on day one was already wrong by day fourteen.

Generic scraping infrastructure doesn't inherently account for any of this. It treats a government tender portal the way it would treat any other webpage: pull the HTML, extract the fields, move on. The procurement-specific layer, validation, normalization, deduplication, monitoring, still has to be engineered on top, deliberately. Skip it and the result is clean-looking, structurally valid, incomplete data, with nothing about the output signaling that anything is missing until a client asks why a contract they know exists never showed up.


Where the Engineering Effort Should Actually Go

These failure modes matter most to a specific kind of team: organizations building procurement intelligence products, supplier discovery platforms, bid alert services, or internal procurement analytics. Anyone whose product only works if the data underneath it can be trusted every day, not just in a demo. They rarely surface during a prototype, when a handful of manually checked notices look fine. They surface once customers start relying on the data daily, at a volume where nobody's checking each notice by hand anymore.

Given that, it's worth being direct about where engineering time is best spent, because "add another portal" is usually the least valuable place to put it.

Adding a new portal is visible. It moves the coverage number, it's easy to announce, and it's simple to report upward. In practice, portal count is the easy metric to put in front of a stakeholder. Nobody needs a meeting to explain why sixty five is bigger than forty. Explaining why a scraper's null rate crept from two percent to eleven percent over a quarter is a much less convenient conversation, and a much more important one.

If the underlying validation, monitoring, and deduplication logic isn't solid, every new portal is just another surface for the same failure modes. One more source that can silently break, duplicate, or drift out of sync unnoticed. The work that actually protects data quality is less visible: monitoring that checks output volume and shape against historical norms, not just whether a request succeeded; validation that catches a field passing its schema check while still being contextually wrong; deduplication robust enough to match records with no shared identifier; schema and timezone handling treated as first-class concerns, backed by regression detection that flags a scraper's behavior changing even when nothing technically errors out.

None of that shows up on a homepage. All of it determines whether the platform holds up once a client actually cross-checks it, which is the real test, not the demo.


Weighing Coverage Against Accuracy

Knowing which sources you monitor is only the first part of the picture. The harder question is whether the records coming back from those sources are complete, correct, and dependable. Coverage and data quality aren't in conflict. A platform doesn't have to sacrifice one for the other, and coverage still matters. But they're not equally valuable, and treating them as though they are is where platforms get into trouble.

Growing from ten portals to fifty means very little if a client loses confidence in the fifty-first notice they check by hand and finds it doesn't match the source. A coverage number can always be grown. Add a portal, the map updates. Rebuilding trust after that isn't nearly as fast, and it usually happens only once the specific failure that caused the doubt has been visibly, verifiably fixed.

That's also why expanding coverage one reliably-integrated source at a time beats chasing the largest possible portal count as fast as possible. It's rarely the quicker way to grow a coverage number. It's the only way to grow one without quietly eroding what sits underneath it. A platform that's quietly right about fifteen portals is in a stronger position with its clients than one that's spectacularly wrong across forty.


What Should a Procurement Data Platform Measure Instead?

If portal count isn't enough to describe procurement data quality, what should replace it? Not another single number.

It's tempting to just swap "coverage" for "accuracy" and call it solved. This article's own title practically invites that swap. But accuracy on its own isn't sufficient either. A dataset can contain perfectly accurate records and still be incomplete, stale, duplicated, or unreliable over time. Data quality isn't a single score. It's the combination of conditions that determine whether a dataset can actually be used to make a decision.

In practice, that breaks down into six distinct, checkable dimensions: different ways of measuring whether procurement data is usable in practice, not six independent scores to average together.

Six dimensions of procurement data quality, framed as questions Procurement Data Quality: Six Questions, Not One Number Coverage Are we looking in the right places? Freshness How quickly does new or changed information arrive? Completeness Did we capture what was actually published? Accuracy Does the record match the source, right now? Deduplication Is the same procurement represented once, not twice? Reliability Does that stay true consistently, not just at launch?
Figure 4. Procurement data quality as six checkable questions, not one number.

None of these six replace coverage. Together, they explain what coverage alone can't. Two are worth separating out, since they're easy to conflate: accuracy is whether a given record is correct right now, reliability is whether the system keeps producing correct records consistently over time. A pipeline can be accurate today and unreliable tomorrow. That's exactly what a silent scraper failure is: a system that was accurate yesterday and has quietly stopped being so.

Coverage and completeness get conflated just as often, but they answer different questions too. Coverage tells you which sources you monitor. Completeness tells you whether you've actually captured everything those sources expose. With fifty portals in the pipeline, incomplete extraction from ten of them can still leave the dataset looking comprehensive from the outside.

Coverage is still on the list. It's a real dimension, not a discarded one. Portal count is one input into a data-quality system, and treating it as the whole definition of quality is what leaves duplicate contracts, mistimestamped deadlines, and silent gaps invisible until a client finds them first.

Being able to answer all six of these questions, not just the first one, is a fundamentally different claim than leading with a portal count. We'd rather see a platform explain how it validates freshness, completeness, and deduplication than simply advertise how many sources it touches. The next section shows two of these dimensions, completeness and accuracy, actually getting solved at the source instead of patched downstream.


Proof That Standardization Can Work: What CPV Gets Right

Not everything in procurement data is broken. If fragmentation is the problem behind the dimensions above, the natural next question is whether it can actually be reduced, not just measured and worked around. The EU's Common Procurement Vocabulary, CPV, and AusTender's move to a shared data standard are two concrete answers.

CPV is a classification system of more than 9,000 codes, codified in EU law by Regulation (EC) No 2195/2002 and mandatory for use on TED notices under the EU's procurement directives. Every contracting authority publishing on TED applies it, regardless of nationality or language. A road resurfacing contract in Poland and a road resurfacing contract in Portugal can carry the exact same CPV code. That's a small detail with a large implication: getting 27 countries, each running its own procurement system in its own language, to agree on and consistently apply a single shared classification is one of the few genuine wins for cross-border consistency in procurement data.

It's not flawless. Some codes are broad enough that two meaningfully different tenders can legitimately share one, and a buyer choosing between several plausible options will sometimes pick the closest available code rather than the most precise one. That's the same kind of judgment call that introduces noise anywhere metadata gets assigned by a human instead of derived automatically. But a single classification system covering the entire EU, in a data landscape where almost nothing else is standardized, shows that fragmentation isn't an inherent property of procurement data. It's the default outcome in the absence of a standard, and CPV is what happens when one actually gets built and enforced.

AusTender, Australia's federal procurement portal, tells a similar story from a different angle: not classification, but delivery. Before 2019, getting structured data out of AusTender meant scraping contract notice pages directly or relying on a custom export process, and neither approach followed a shared standard, so every integration ended up building its own field-mapping logic from scratch. A 2017 government review found that AusTender met only around a third of the data-field requirements of the Open Contracting Data Standard, OCDS, an open schema designed specifically to make procurement data comparable across systems. Australia's Department of Finance committed to closing that gap. In January 2019, AusTender began publishing a proper OCDS-compliant JSON API, replacing a landscape of separately reverse-engineered integrations with one consistent, documented format for everyone pulling from it.

CPV tackles classification, while AusTender's OCDS implementation tackled the way procurement data is delivered. They solve different problems, but both reduce the amount of interpretation a downstream pipeline has to perform, which is exactly where the failures earlier in this article come from. Neither happened by accident, and neither is typical. Most of the fixes described above exist precisely because most portals never get rebuilt around a shared standard. But both show that the fragmentation running through this article isn't a law of nature. When a data source commits to a real standard, a meaningful share of the accuracy problem gets solved before it ever reaches a pipeline at all.


From Coverage to Decision Confidence

Coverage tells you how much data a platform has. On its own, it doesn't tell you whether that data can be trusted. That's exactly why the six dimensions above exist: they're what determines whether the gap between "has the data" and "can trust the data" closes, or stays invisible until someone checks your numbers against the source and finds they don't match.

Procurement intelligence, in the end, isn't measured by how many portals a platform monitors, or even by any single number on that list. It's measured by how confidently someone can make a decision, bid on a contract, flag a supplier, report a market trend, based on what the platform shows them.

Procurement intelligence isn't built one portal at a time. It's built one trusted decision at a time.

If you're building this yourself and these problems are becoming operationally expensive, this is the layer ProcureDataLab works on: source integration, validation, deduplication, freshness monitoring, and ongoing scraper maintenance as portals change.

The goal isn't simply to collect more notices. It's to make the data dependable enough that your users don't have to check it twice.

Building procurement intelligence?

The difficult part isn't another portal.

It's keeping the data dependable as sources change. Source integration, validation, deduplication, freshness monitoring, and scraper maintenance — that's the layer we work on.

Talk to ProcureDataLab
Source integration
Validation
Deduplication
Freshness monitoring
Scraper maintenance

FAQ

Frequently asked questions

What is procurement intelligence?

Procurement intelligence platforms collect, structure, and analyze public procurement data — tenders, contract awards, supplier records — so businesses can track opportunities, competitors, and market trends across government buying activity.

Why isn't monitoring TED enough to cover EU procurement?

TED only receives notices for contracts above a set publication threshold, which varies by contract type and buyer. Procurement below that threshold is published only on national portals — BOAMP, TenderNed, e-Vergabe, and others — each in its own format. A TED-only pipeline can look complete while missing a substantial share of actual procurement activity.

What is CPV, and why does it matter?

CPV (Common Procurement Vocabulary) is a classification system of more than 9,000 codes that every buyer publishing on TED applies to describe what a contract is for, regardless of language or country. It's one of the few genuinely standardized layers in an otherwise fragmented data landscape.

How are duplicate tenders detected across portals?

There's no universal ID linking the same tender across, say, TED and a national portal. Duplicate detection instead relies on matching attributes — buyer identity, publication timing, contract value, CPV classification, and other metadata — with enough confidence to determine two records describe the same underlying contract.

Why do silent scraper failures matter more than obvious ones?

A scraper that returns zero notices with no error thrown looks identical, from a monitoring standpoint, to a portal that legitimately published nothing that day. Catching this requires monitoring that watches for unusual drops in output volume, not just failed requests.

What should I ask a procurement data vendor besides how many portals they cover?

Ask how they catch a scraper that's still running but quietly returning less than it used to. Ask how they merge duplicate contracts when two portals describe the same tender differently. Ask what happens when a source changes its layout without warning. Those questions tend to reveal a lot more about whether the data can be trusted than a portal count ever will.

Procurement Data Engineering
01

Coverage vs. Accuracy

You are here

More engineering notes in this series are coming soon. Back to all articles