Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Counting live job adverts at scale revealed a useful but easily misunderstood picture: stale listings and repeated posts are common enough to measure, while a flawed text reader can make a pay-transparency statistic wrong. Praveen Kumar’s September 30, 2026 account describes TUNAI’s nightly crawler, the limits of its job-ad data, and three engineering failures that changed how its results should be interpreted.

What the 600,000 figure counts

TUNAI’s crawler collects live adverts directly from employers’ applicant-tracking systems, including Greenhouse, Workday, Lever, Ashby, Oracle HCM, and about fifteen others. The collection described by Kumar spans the UK, US, Canada, and Australia. His September 30, 2026 article puts the corpus at about 600,000 live adverts, recounted nightly. That is a dated description of TUNAI’s collection, not a census of every job listing, a count of hires, or a stable measure of the labour market.

The reported totals changed across snapshots: a September 29 dataset card described a corpus of 601,851 adverts, assembled from about 590,000 live adverts in its summary; a later TUNAI insights result, updated October 5, 2026, described 666,000. Those figures refer to different versions and should not be blended into one total. Each statistic also needs its own denominator: some findings use only adverts with a stated date, while others use a subset carrying a posting time.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The article and snapshot are TUNAI’s own account and data, not an independent audit of crawler coverage or every underlying advert. The dataset card identifies its snapshot as CC BY 4.0 and requests credit to TUNAI (jobs.tun-ai.com) with a link to the insights page.

What the collected adverts revealed

The following are observations reported by Praveen Kumar and TUNAI from their collected adverts. They describe that corpus and its parsing rules, not all hiring activity.

Some listings had very old dates

Kumar reported 34,000 live adverts dated more than a year earlier, equal to 6% of adverts that gave a date; the oldest displayed date was from 2010. He also reported 66,000 adverts older than six months. The September dataset snapshot gives the latter as 66,000, or 11% of adverts with a date, across 588,202 adverts. These are employer-supplied dates. An old date can indicate a stale listing, but Kumar notes that some employers keep evergreen adverts open to collect CVs, so age alone does not prove that a vacancy is fake or still actively being filled.

Rank #2
Sale
Storytelling with Data: A Data Visualization Guide for Business Professionals
  • Wiley
  • Language: english
  • Book - storytelling with data: a data visualization guide for business professionals

Some adverts were reposted repeatedly

TUNAI found 3,100 adverts reposted at least five times; one had been reposted 31 times. Reposting can reflect a hard-to-fill role or a listing kept alive. It is a reason to look more closely, not proof of fraud, a ghost vacancy, or a single underlying role: repost counts depend on how adverts are matched and grouped.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Posting patterns and unusual wording

In the September account, 20% of UK and US adverts went up on Friday and 4% appeared on weekends. The dataset card describes the day-of-week observation as based on employer-stated dates over the preceding 90 days. Separately, 31% of UK adverts that included a posting time appeared between 2 p.m. and 5 p.m.; the earlier snapshot excludes date-only midnight stamps from this time-of-day calculation.

The wording counts are similarly specific to TUNAI’s snapshot: 2,500 US job titles included an exclamation mark, compared with 79 UK titles. Three live titles asked for a “rockstar,” five for a “ninja,” and 140 used “champion.” They are counts of titles in the collected adverts, not measures of employer quality.

A later TUNAI dashboard result, updated October 5, 2026, reported weekend postings at 4% in the UK and 5% in the US. That later refresh should remain distinct from the September article’s 4% weekend figure; the totals and observed versions differ.

Three failures that changed how the data was handled

Exact duplicates from Oracle HCM

Kumar says Oracle HCM can expose the same requisition through multiple career sites. In 61,111 eligible Oracle adverts, he found 43,704 byte-identical records, grouped into 15,035 sets. Earlier, broader similarity matching had removed hundreds of real vacancies, because similar wording does not mean two adverts describe the same job.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The narrower rule in his pipeline links records only when source, title, place, and text are identical. It keeps the records in the corpus but publishes one public page for the group. This is a description of TUNAI’s handling of its own Oracle data, not a general benchmark for all Oracle HCM systems.

PostgreSQL ignored partial indexes

In another TUNAI pipeline issue, partial indexes were declared with WHERE live, while queries used live IS true. Although the predicates appear logically equivalent, PostgreSQL did not use those indexes in the described system. A match took 50 to 130 seconds. Kumar’s operational lesson is to inspect the idx_scan value for partial indexes rather than assuming a declared index is being used. The account does not establish that this mismatch will produce the same result across PostgreSQL versions, schemas, or query plans.

A truncated reader missed salary text

An earlier TUNAI insight said 94% of US adverts gave no pay. Kumar questioned the result and manually checked 200 US adverts that the system had classified as having no pay. In that sample, 43 adverts stated pay in text already held by the system; 42 of those put it after character 600, beyond the reader’s stopping point. He withdrew the 94% claim while the fix remained uncertified. The September dataset card likewise excludes facts based on advert pay text until the US pay-reading fix is certified.

The lesson is that a statistic can be precise and still be wrong if the extraction process overlooks relevant text. As Kumar put it, “when your number disagrees with everyone else’s, check your reader before you publish the surprise.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to interpret job-ad counts responsibly

Whether you are evaluating a dataset or citing one of its findings, a headline total is not enough. Check what was collected and how each value was derived.

  • Coverage: identify the countries, employers, and applicant-tracking systems represented. A collection from employer career sites is not automatically representative of all vacancies.
  • Snapshot and update cadence: record the date and whether the figure is a one-time snapshot or a refreshed total. “Live” counts change as adverts appear, expire, or are recrawled.
  • Definition of live: find out how the collector decides an advert is still live and whether records are retained after removal or reposting.
  • Duplicates and reposts: distinguish byte-identical duplicates, similar adverts, and repeated postings. Matching too broadly can collapse genuinely different vacancies.
  • Dates and denominators: note whether a statistic uses all records, only adverts with dates, or only those with explicit posting times. Date-only values should not silently become time-of-day evidence.
  • Text extraction: establish how much advert text is read and how pay or other fields are recognized. A failed parser can turn “not found” into a false claim that information was not disclosed.
  • Attribution and reuse: follow the dataset’s license and credit requirements. TUNAI’s September snapshot is identified as CC BY 4.0 and requests credit to TUNAI and a link to its insights page.

TUNAI also describes a CV-based service that ranks live adverts against a CV and a free ATS CV checker that does not require sign-in. These are different tasks: matching helps surface roles, while a checker examines how an ATS parses a CV. The available account does not establish comparative performance against other tools.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.