How hiring-signal data is built: a methodology for recruitment agencies evaluating vendors
In short
A hiring signal is a live, dated job vacancy tied to a specific employer, not a generic industry classification. This page describes how two real vacancy datasets are built, covering Germany, Poland, the United Kingdom, France, the Netherlands, Switzerland, Lithuania and Estonia, with a combined 49,352 rows and a stated caveat: one dataset's "role" field is a search-query label, not a job taxonomy, so titles must be read directly rather than trusted as pre-classified. Neither dataset carries any Australian rows, and this page says so rather than implying coverage that does not exist.
On this page
What counts as a hiring signal
A hiring signal, in the sense this page uses it, is a specific job vacancy: a company name, a job title, a location and a date it was posted or last seen live. It is not an industry-level estimate of how many companies in a sector might be hiring. The distinction matters because a vendor can describe both as "hiring signals" while only one of them points a recruiter at a company with a real, current, open role.
This page describes the methodology behind two vacancy datasets built on that stricter definition, states exactly which countries they cover, and flags a specific labelling caveat that a recruitment agency evaluating any vendor's hiring-signal claim should ask about directly.
The distinction is not academic. An agency that buys into a vendor's hiring-signal claim on the strength of an industry-level estimate is effectively buying a general lead list with a more exciting name attached. Testing the claim against the definition used here, a specific, dated, named vacancy, is the fastest way to find out which kind of data a vendor is actually offering.
Two datasets behind the numbers
The first dataset holds 33,367 vacancy rows, with posting dates ranging from 2015-11-18 to 2026-08-05 and a scraping window from 2026-07-22 to 2026-08-05 for the most recently collected records. Each row carries a company name, job title, city, source and the date it was scraped, structured so a specific vacancy can be traced back to where it was found.
The second dataset, sourced from a job-board aggregator, holds 15,985 rows, with posting dates from 2019-04-17 to 2026-08-22 and a first-seen/last-seen collection window from 2026-07-17 to 2026-08-22. It tracks how many times a listing has been seen live, which is a separate freshness signal from the posting date alone.
Combined, the two datasets total 49,352 vacancy rows. The two are not simply additive for every country, since the first dataset spans nine countries and the second currently spans only two, so a recruitment agency comparing coverage across countries should read the per-country table below rather than the combined total, which can overstate depth in a country only one of the two datasets actually reaches.
Country coverage, row by row
| Country | Dataset 1 rows | Dataset 2 rows |
|---|---|---|
| Germany | 14,267 | 5,399 |
| Poland | 8,146 | 10,586 |
| United Kingdom | 4,154 | 0 |
| France | 3,877 | 0 |
| Netherlands | 1,634 | 0 |
| Switzerland | 1,247 | 0 |
| Lithuania | 34 | 0 |
| Estonia | 8 | 0 |
| Australia | 0 | 0 |
Germany and Poland carry the deepest coverage across both datasets. Lithuania and Estonia are present but thin, 34 and 8 rows respectively, which a recruitment agency evaluating Baltic hiring signals should treat as an early-stage sample rather than a mature dataset comparable to Germany's.
The United Kingdom, France, the Netherlands and Switzerland each appear only in the first dataset, with no corresponding rows in the second. An agency working across several of these countries should ask which specific dataset, or combination, a vendor is actually drawing from for each one, rather than assuming a vendor's "European coverage" pulls evenly from the same underlying source across every market.
Why a "role" field is not a job taxonomy
The first dataset includes a field called "role", and it is easy to assume that field is a classified occupation category, something like "driver" or "electrician" applied consistently across every row. It is not. The field is a free-text label recording which search query originally pulled that row, for example a query run for "kierowca" or "security", not a formal classification applied after the fact.
The practical effect is that two rows with the same "role" label are not guaranteed to describe the same kind of job, and two rows for genuinely similar jobs are not guaranteed to share the same "role" label if they were pulled by different queries. Any agency, this one included, reporting on this dataset needs to read the job title field directly and filter or group titles itself, rather than trusting "role" as a pre-built category.
This is a common failure mode in scraped job-market data generally, not something specific to this dataset. A query-driven collection process is efficient for gathering rows quickly across many searches, but it inherits the vocabulary of whoever wrote the search queries, not a standardised occupational classification. An agency asking a vendor whether its data uses a real taxonomy, such as a national or international standard occupational classification, versus a set of internal search labels, is asking a question that goes directly to how safely the data can be aggregated.
Freshness: how current the signal actually is
Both datasets carry a collection window that is more useful than the posting date alone for judging freshness. The first dataset's most recent scrape ran between 2026-07-22 and 2026-08-05, meaning any row collected in that window reflects what was live on a job board within two to six weeks of this page's publication date. The second dataset's first-seen and last-seen fields, spanning 2026-07-17 to 2026-08-22, go further, tracking whether a specific listing was still live the last time it was checked, not only when it first appeared.
A vendor that cannot state a comparable collection window, when the data was last pulled, not only when the underlying job was first posted, is not offering a freshness guarantee, only a historical archive with an unknown lag.
The gap between the oldest posting date and the most recent collection window is itself worth noticing. Both datasets described here carry posting dates going back several years, to 2015 in the first case and 2019 in the second, alongside a collection window of only a few weeks. That combination is normal for this kind of data, older rows persist in the archive even after the vacancy has closed, but it means a raw row count for a country includes historical vacancies that are no longer live, and a vendor should be able to say what share of its reported count is current versus archival.
What to ask a vendor claiming hiring-signal coverage
Ask for the specific country row counts, not a regional total. Ask what a "role" or category field in the vendor's own schema actually represents, a formal classification or a query label, since the difference changes how safely the field can be trusted for segmentation. Ask for the most recent scrape or collection date, not only the oldest and newest posting dates in the archive, since a wide date range can hide a dataset that has not been refreshed recently.
Ask, finally, where the underlying job postings come from, a single job board, several combined, or a paid aggregator feed, since the source affects both coverage and duplication. A vendor pulling from several boards without deduplication can inflate a row count with the same vacancy counted more than once, which is a different problem from thin coverage but produces a similarly misleading headline figure.
A vendor willing to answer all three specifically is showing a dataset built to be interrogated. One that answers only in general terms, "extensive coverage," "regularly updated," is asking to be trusted rather than checked.
Ask, in addition, whether the vendor can show a handful of live rows for the specific country and sector the agency cares about, rather than a coverage figure for the dataset as a whole. A dataset with a large total row count can still be thin in the one country and sector combination that actually matters to a given agency, and only a targeted sample reveals that.
Where the coverage stops
Australia carries zero rows in both datasets described on this page. Neither the general vacancy dataset nor the job-board aggregator dataset holds a single Australian hiring signal at the time of writing, despite a large Australian company database existing separately (4,770,875 active business entities on the Australian Business Register). Those are two different kinds of data, and a recruitment agency should not assume that a vendor with strong Australian company coverage also has Australian hiring-signal coverage.
The practical guidance for an agency targeting Australia specifically is to use company-register data, filtered by industry, size and state, as the targeting layer instead of a hiring signal that does not currently exist in this methodology, and to ask any vendor claiming otherwise to produce the underlying row count before believing the claim.
Stating this gap plainly is more useful to an agency than a vague promise of future coverage. A methodology page that hides its weakest market behind confident, general language about "global reach" is a page built to be read once and believed. A page that names the exact countries covered, and the one large market where the coverage is currently zero, is a page built to be checked and to still hold up afterward.
Frequently asked
What is a hiring signal, precisely?
Which countries does this hiring-signal methodology actually cover?
Why is the "role" field in a vacancy dataset unreliable as a job category?
Does this hiring-signal data cover Australia?
How should a recruitment agency check a vendor's hiring-signal freshness claim?
Want the accounts behind these numbers?
Book a short strategy call. We will show you which employers in your region and role family are hiring right now, and what we would write to them.
Book a strategy call