Every analysis of public procurement begins with a deceptively simple statement: this company won this tender. Deceptively — because “we know who won” can mean four different things, and the difference between them decides how much the conclusion above is worth.
A winner is worth as much as its source
The split by origin looks like this:
- 118,900 records from the official open-format dump. This is the authoritative source — values in euro, contractor identifiers, distribution across lots. If the dump says it, it is so.
- 34,300 from the platform’s public API. Also a primary source, equally reliable, but with different coverage and a different refresh rhythm.
- 22,300 from reading the pages themselves. Where neither the dump nor the API carries the result, it is read off the published page. Reliable, but dependent on how the content was laid out.
- 6,500 from a language model over the text of the decision. The last resort: the decision exists only as prose, the winner is described in a sentence rather than a field. Here the record is inferred, not read.
The four are not equivalent and we do not pretend they are. That is why every record is tagged with where it came from.
Why this is a question at all
Public winner summaries often diverge from what the decision itself says. The reasons are mundane: a tender split into lots where the summary shows one contractor instead of five; a procedure cancelled and re-run; a contract signed by a consortium but recorded under the lead partner’s name.
Analyse a competitor from a summary like that and you get a plausible, false picture of their history. You will miss the lots they won as part of a consortium. You will credit them with contracts that were never concluded.
What this means for you
When you commission a competitor analysis and receive a list of tenders they have won, you should be able to ask “how do we know this” about any row — and get an answer other than “from the database”.
We work this way not because it is pleasant, but because the alternative is to blend four levels of certainty into a single number and call it a fact. The number would look cleaner. It would not be truer.
Anyone building a bid on competitor analysis is building on these records. They deserve to know what the records are made of.