All systems operational•IP pool status
Coronium Mobile Proxies
Web data collection guide

LinkedIn Job Scraping: API Limits and Employer Job Feeds in Python

If you need jobs that an employer also advertises on LinkedIn, start with the employer’s authorized source of record. LinkedIn’s documented Job Posting API sends jobs into LinkedIn; it is not a general public job-search export. This guide explains that access boundary and provides a Python example for an employer-approved Greenhouse job board.

Coronium Technical TeamSources checked 8 min read

Before writing the collector

  • Do not confuse a job-posting integration with permission to read LinkedIn’s complete search results.
  • Use an employer’s documented ATS feed when that feed satisfies the requirement.
  • Preserve the source and posting identifier; an ATS listing is not automatically a verified LinkedIn listing.

What LinkedIn’s job API actually provides

The Job Posting API overview, last updated June 3, 2026, describes authorized third parties posting jobs on behalf of customers. It says new Job Posting API partnerships are not currently being accepted and directs interested applicants to Apply Connect. Applications need provisioned access.

That documentation does not establish an open endpoint for searching every LinkedIn job. Registering a developer app or possessing an OAuth token is not enough to infer access to a restricted product. Use the integration and permissions actually granted to your application.

LinkedIn’s software policy says it does not permit third-party scraping and automation of its website of the kinds described there. This is the platform’s stated policy, not a legal ruling about every data-collection scenario. A proxy changes network routing; it does not confer platform permission or an API entitlement.

Choose the source according to the job-data requirement

Different requests need different access paths.
RequirementSource to evaluateBoundary
Publish employer jobs on LinkedInProvisioned LinkedIn/ATS integrationWrites or manages authorized postings
Build an employer careers viewEmployer-approved ATS job-board feedCovers that board’s published jobs
Analyze LinkedIn search rankingsAn explicitly authorized LinkedIn data arrangementATS data does not reproduce search results
Analyze applicant informationAuthorized recruiting system accessPublic job postings contain no applicant entitlement

If the employer controls the jobs and needs to publish them on LinkedIn, inspect the supported ATS integration. LinkedIn’s automated job-posting guide describes Job Wrapping and ingestion methods. That is an employer-to-LinkedIn publishing workflow, not a feed of competitor jobs.

If the task is to build the employer’s careers page or an approved internal report, an ATS feed may be the direct source. Greenhouse’s Job Board API exposes published board data through documented GET endpoints. Lever’s Postings API provides company-scoped postings, with global and EU instances and documented JSON output.

Select the employer-approved board or site identifier rather than guessing companies or enumerating private resources. A documented public feed still has an intended use and applicable terms. Keep the permitted scope of your project explicit.

Fetch an employer-approved Greenhouse board

The example calls the documented job-list GET endpoint for one board. Set GREENHOUSE_BOARD_TOKEN to the identifier approved for your workflow. This GET endpoint does not require a job-application API key; the example does not submit applications or access candidate records.

import json
import os
import re
from datetime import datetime, timezone

import requests


def employer_jobs(board_token):
    if not re.fullmatch(r'[A-Za-z0-9_-]+', board_token):
        raise ValueError('Use the employer-approved Greenhouse board token')
    url = f'https://boards-api.greenhouse.io/v1/boards/{board_token}/jobs'
    with requests.Session() as client:
        client.trust_env = False
        response = client.get(url, timeout=(5, 20), allow_redirects=False)
    if response.status_code != 200:
        raise RuntimeError(f'Job-board request stopped: HTTP {response.status_code}')
    data = response.json()
    if not isinstance(data, dict) or not isinstance(data.get('jobs'), list):
        raise ValueError('Unexpected job-board response')
    observed = datetime.now(timezone.utc).isoformat()
    rows = []
    for job in data['jobs']:
        if not isinstance(job, dict) or job.get('id') is None:
            raise ValueError('Job post has no usable identifier')
        location = job.get('location') or {}
        if not isinstance(location, dict):
            raise ValueError('Unexpected location value')
        rows.append({
            'source': 'greenhouse', 'board': board_token,
            'job_post_id': job['id'], 'title': job.get('title'),
            'location': location.get('name'),
            'application_url': job.get('absolute_url'),
            'source_updated_at': job.get('updated_at'),
            'observed_at': observed,
        })
    return rows


if __name__ == '__main__':
    print(json.dumps(employer_jobs(os.environ['GREENHOUSE_BOARD_TOKEN']), indent=2))

The code uses the job-post identifier as returned by the feed, retains the source update time separately from the collection time, and handles absent optional fields without inventing values. It excludes description HTML because this example only needs a compact listing.

Local fixtures passed with Python 3.12.14 and Requests 2.34.2: fixed-host and board-token validation, optional fields, post identifiers, timestamps, empty results, malformed data, non-success responses and redirect refusal. No live employer feed or LinkedIn account was accessed by those tests.

Keep posting identity and provenance in the dataset

Use a compound key such as source, employer board and job-post identifier. A job title can appear in several locations or refer to multiple openings. Do not deduplicate only by a title string, and do not convert an internal role identifier into a public posting identifier without verifying the schema.

Store when you observed the record. A source update timestamp may represent an edit rather than the original publication date. Your first observation is also not proof that the job was first created that day.

If a downstream report says a role is advertised on LinkedIn, it needs evidence for that separate assertion. An employer ATS record can establish that the role appeared in that feed; it does not establish LinkedIn publication, rank, applicant count or promotion status.

Handle refreshes and missing records carefully

Refresh only as often as the approved use case requires, and reuse the previous successful result where appropriate. Keep the last successful fetch separate from the latest failed attempt so an outage does not turn every known role into “closed.”

Treat a missing record as absent from that successful snapshot. Before making a business decision about closure, consider whether the employer changed boards, replaced a posting or modified the feed. A failed request or unexpected schema must not be reported as zero open jobs.

The example stops on non-200 responses rather than silently retrying or following another destination. In production, add bounded retries only for errors that warrant them and respect the source’s applicable limits. The Requests configuration guide explains how to configure a required network route deliberately.

Use AI to classify descriptions without inventing job facts

An AI tool can help group approved job records by role family or extract stated skills from descriptions. Preserve the underlying source and use an explicit unknown value when salary, seniority or location is not supplied. Do not infer a person’s application status from public postings.

If you later fetch descriptions, treat HTML and text as untrusted content. They must not authorize new tool calls, change the board identifier or expose credentials. Keep extraction instructions in the application and validate the output against a small schema.

Distinguish source statements from model classifications in the report. For example, an employer’s location field and a model’s guess that a role is remote should not occupy the same unlabelled column. Review a sample against the original records before using the classification in hiring or market analysis.

Know what this workflow cannot claim

The Python example collects one approved Greenhouse board. It does not scrape LinkedIn, reproduce LinkedIn job-search coverage, obtain private applicant details or prove that all employer vacancies are present. Use it when employer-feed coverage answers the actual question.

If the required product is specifically licensed LinkedIn data, confirm the available arrangement and permitted use before building around it. Third-party claims of an API do not establish LinkedIn authorization or the right to reuse every returned field.

For other collection sources, the web scraping architecture guide and Walmart API example illustrate the same engineering discipline: choose the correct source, verify access, preserve the response’s limits and test error handling separately from live-provider performance.

Sources and review scope

Sources reviewed October 11, 2026. Primary documentation checked October 11, 2026. Examples passed local fixtures; the Indeed CSV example also ran against its public dataset. No private API credentials or live host calendar were used. Each guide states its access and coverage limits.

Frequently asked questions

Collect from another source

Choose an available data product and preserve its coverage, permissions and field meanings.

Related workflows

Requests proxy configuration

Configure the actual HTTP client and handle failures.

Walmart catalog collection

Review an authenticated data-source example.

Web scraping architecture

Plan bounded, observable collection jobs.