All systems operationalโ€ขIP pool status
Coronium Mobile Proxies
Web data collection guide

Indeed Job Scraping: Partner APIs and Hiring Lab Data in Python

An Indeed job-management integration and a labor-market research dataset answer different questions. Indeed documents partner APIs for employer workflows, while Indeed Hiring Lab publishes aggregate job-posting data. This guide explains the distinction and provides a tested Python example for the licensed public CSV, without presenting it as a feed of individual vacancies.

Coronium Technical TeamSources checked 7 min read

Before writing the collector

  • Use the partner product and employer scope required for a job-management integration.
  • Hiring Labโ€™s public Job Postings Index measures aggregate demand; it contains no individual job descriptions or applicant records.
  • Keep the observation date, series definition and source version with any calculated result.

Choose the source that matches the question

Select the data product before writing a collector.
RequirementRelevant routeWhat it does not establish
Manage a client employerโ€™s jobsAuthorized Indeed partner integrationAccess to unrelated employers
Measure hiring-demand changesHiring Lab aggregate CSVIndividual job descriptions or vacancy counts
Read employer-published openingsEmployer-approved ATS feedIndeed search position or publication status

Start with the record you need. If the task is to manage an employerโ€™s jobs, review Indeedโ€™s Job Sync API and Job Update API documentation and the integration requirements for that employer. A documented listing query does not establish permission to search every employerโ€™s inventory.

If the question is how hiring demand has changed, the public Hiring Lab repository offers country, sector and geographic series. It avoids collecting vacancy pages when an aggregate measure is the actual requirement.

For a list of openings maintained by an employer, an approved employer careers feed may be sufficient. The LinkedIn and employer-feed guide includes a Greenhouse example and explains why an ATS record is not proof of its publication or ranking on another platform.

Check the current employer API behavior

The Job Update overview, updated October 9, 2026, documents listing and updating jobs for employer and agency workflows. It describes different client setups and calls for confirming visibility with the agency support representative. The Job Sync overview covers creating, updating, expiring and inspecting jobs within an integration.

A September 15 release note matters when an existing integration appears to miss records: omitting jobFeedType returns only INTEGRATED_FROM_PARTNER. Set the appropriate filter when the authorized workflow requires crawled or hosted jobs. The same release explains that jobRequisitionId matching ignores case but does not trim whitespace, and that one requisition ID can match multiple jobs.

Record the employer, granted access and filters in the integrationโ€™s configuration. Test those details in the approved environment before interpreting an empty response as โ€œno jobs.โ€ This article does not provide a live partner credential or claim that an unregistered application can call those endpoints.

Understand the Hiring Lab index

The repositoryโ€™s methodology and column definitions describe daily observations refreshed weekly, with February 1, 2020 as the baseline. An index of 101 means the series is 1% above that baseline; it does not mean a 101% increase. The seasonally adjusted and unadjusted columns are separate series.

The US aggregate file currently labels its series total postings and new postings. Select the intended label explicitly. The documentation defines new postings as those on Indeed for seven days or fewer. Do not add the two series or treat them as independent counts.

Historical values can be revised. Retain the downloaded file, its retrieval time and a source commit or content hash for reproducibility. A figure calculated from todayโ€™s file can legitimately differ from an earlier saved snapshot; that difference needs investigation, not an automatic claim about newly created jobs.

Read the public CSV with Python

Download US/aggregate_job_postings_US.csv from the official repository and save it locally. Save the following script as indeed_index.py, then run python indeed_index.py aggregate_job_postings_US.csv. It uses the Python standard library and makes no network request.

import csv
import hashlib
import io
import json
import sys
from datetime import date
from decimal import Decimal, InvalidOperation
from pathlib import Path

SOURCE = 'https://github.com/hiring-lab/job_postings_tracker'


def latest_us_index(raw):
    reader = csv.DictReader(io.StringIO(raw.decode('utf-8-sig')))
    required = {'date', 'jobcountry', 'variable', 'indeed_job_postings_index_SA'}
    if not required.issubset(reader.fieldnames or []):
        raise ValueError('Unexpected Hiring Lab CSV columns')
    observations = {}
    for row in reader:
        if row['jobcountry'] != 'US' or row['variable'] != 'total postings':
            continue
        try:
            day = date.fromisoformat(row['date'])
            value = Decimal(row['indeed_job_postings_index_SA'])
        except (ValueError, InvalidOperation, TypeError):
            raise ValueError('Invalid date or index') from None
        if not value.is_finite() or value < 0 or day in observations:
            raise ValueError('Invalid index or duplicate observation date')
        observations[day] = value
    if not observations:
        raise ValueError('No US total-postings observations')
    day = max(observations)
    value = observations[day]
    return {
        'observation_date': day.isoformat(),
        'country': 'US',
        'series': 'total postings, seasonally adjusted',
        'index': str(value),
        'percent_change_since_2020_02_01': str(value - Decimal('100')),
        'source': SOURCE,
        'attribution': 'Indeed Hiring Lab, CC BY 4.0',
        'license': 'https://creativecommons.org/licenses/by/4.0/',
        'transformation': 'Latest US total-postings row; index minus 100',
        'input_sha256': hashlib.sha256(raw).hexdigest(),
    }


if __name__ == '__main__':
    print(json.dumps(latest_us_index(Path(sys.argv[1]).read_bytes()), indent=2))

The function selects the latest US total-postings observation, validates its date and numeric value, and reports the change from the baseline. It rejects duplicate dates instead of quietly choosing a conflicting row. A SHA-256 hash identifies the exact input used; the output also carries attribution, the license URL and a description of the transformation.

The example passed local fixtures for series selection, date ordering, decimal arithmetic, missing data, duplicate dates and invalid numeric values. It also ran against the public CSV downloaded on October 11, 2026: the latest included observation was October 2, with an index of 103.84 and a calculated change of 3.84%. Those are dated results from that file, not a live claim about todayโ€™s labor market.

The repository revision checked was e9811ce, committed October 6, 2026. Keep the output hash with the downloaded file when reproducing that result.

Carry attribution into charts and exports

Hiring Lab identifies the repositoryโ€™s data as available under Creative Commons Attribution 4.0. When sharing a chart or transformed table, retain the source attribution, link the license and identify modifications. The repository license also contains warranty limitations; the presence of a license is not a guarantee that an analytical conclusion is correct.

A useful chart caption names Indeed Hiring Lab, the country, the chosen series, the observation range, the baseline and the download date. Include the source link and state that calculations or filtering were applied. Keep those details in exported files as well as the dashboard.

That repository license applies to the material published there. Do not extend it to job descriptions, applicant records or other Indeed products merely because they share the same brand.

Give AI analysis a bounded dataset and reproducible calculation

An AI assistant can help explain a trend after the pipeline has selected and calculated it. Supply the series definition, observation dates, attribution and the computed result together. Ask it to distinguish an index movement from a count of jobs and to identify missing evidence before making a causal claim.

Calculate arithmetic in code. Keep the source file and transformation available for review rather than accepting a number generated from prose. If the model compares countries or sectors, verify that the series and normalization support the comparison.

This public aggregate workflow needs neither a browser session nor proxy rotation. For an authorized HTTP integration with a separate routing requirement, use the Requests configuration guide. Routing configuration does not change an APIโ€™s employer scope or turn an aggregate dataset into vacancy-level records.

Sources and review scope

Sources reviewed October 11, 2026. Primary documentation checked October 11, 2026. Examples passed local fixtures; the Indeed CSV example also ran against its public dataset. No private API credentials or live host calendar were used. Each guide states its access and coverage limits.

Frequently asked questions

Collect from another source

Choose an available data product and preserve its coverage, permissions and field meanings.

Related workflows

LinkedIn jobs and employer feeds

Read an employer-approved careers feed without confusing platform coverage.

Python Requests proxy configuration

Configure an authorized HTTP client explicitly.

Google data collection

Choose among search, site-performance and AI-grounding products.