Yelp Business Data Scraping: Places API in Python and AI Search
For a Yelp business-search integration, use the documented Places API and the plan your application is authorized to use. It returns selected business data with explicit coverage, request and retention limits. This guide provides a bounded Python request and explains how Places search differs from review excerpts and Yelpโs conversational AI API.
Before writing the collector
- Yelp business search is capped and excludes businesses without reviews; it is not a complete local-business census.
- A price level such as โ$$โ is a category, not a menu price or monetary amount.
- The documented cache limit matters: API access does not imply permission to build an unrestricted permanent archive.
Create the integration around the authorized product
Yelpโs Places introduction describes its current endpoint families. Business search, business details and review excerpts serve different purposes. Choose the endpoint that returns the data your application actually needs instead of treating every Yelp page as one interchangeable dataset.
The authentication guide requires an application and appropriate plan, with requests authenticated using a Bearer API key. It also points developers to the API terms and display requirements during setup. Verify the conditions accepted for your application before deploying a consumer-facing feature.
Use an authorized Places key for the example below. It does not register an application, enable paid access or collect a live dataset as part of the local tests.
Understand what a business search can return
| Field or behavior | Interpretation | Avoid inferring |
|---|---|---|
| Yelp business ID | Identifier within this source | A universal company identifier |
| Rating and review count | Values returned at collection time | A measure of all customers or all review sites |
| Price level | A relative category such as $$ | An exact menu or service price |
| Search result list | A bounded response to your query | A complete geographic census |
The business-search reference says an originating query can return up to 240 businesses and excludes businesses without reviews. It accepts a location or coordinates, with documented search filters. The reference also cautions that returned businesses may not lie strictly within the supplied location.
Define the intended geographic scope before using the response in a report. A search for a city name is not proof that every result is inside an administrative boundary. If that boundary matters, obtain appropriate location information and apply a documented geographic check.
Preserve the Yelp business identifier and source URL. Keep missing optional fields missing. Do not turn the returned price category into a numeric estimate, and do not describe the response as every business operating in the area.
Run one Places search with Python
Set YELP_API_KEY using your runtimeโs secret mechanism, install a supported Requests release and save this as yelp_businesses.py. The example makes one request for up to ten businesses. Review the transient output under the applicationโs retention rules; it is not an archival exporter.
import json
import os
from datetime import datetime, timezone
import requests
URL = 'https://api.yelp.com/v3/businesses/search'
def find_businesses(term, location, api_key):
if not all(value.strip() for value in (term, location, api_key)):
raise ValueError('Provide a term, location and authorized Yelp API key')
with requests.Session() as client:
client.trust_env = False
response = client.get(
URL, params={'term': term, 'location': location, 'limit': 10},
headers={'Authorization': f'Bearer {api_key}', 'Accept': 'application/json'},
timeout=(5, 20), allow_redirects=False,
)
if response.status_code != 200:
raise RuntimeError(f'Yelp request stopped: HTTP {response.status_code}')
data = response.json()
if not isinstance(data, dict) or not isinstance(data.get('businesses'), list):
raise ValueError('Unexpected Yelp response')
observed = datetime.now(timezone.utc).isoformat()
rows = []
for business in data['businesses']:
if not isinstance(business, dict) or not business.get('id'):
raise ValueError('Business has no usable Yelp ID')
rows.append({
'yelp_id': business['id'], 'name': business.get('name'),
'yelp_url': business.get('url'), 'rating': business.get('rating'),
'review_count': business.get('review_count'),
'price_level': business.get('price'), 'observed_at': observed,
})
return rows
if __name__ == '__main__':
rows = find_businesses('coffee', 'Boston, MA', os.environ['YELP_API_KEY'])
print(json.dumps(rows, ensure_ascii=False, indent=2))
The key is sent in the authorization header, not in the URL. Redirects and inherited proxy settings are disabled deliberately. If an authorized deployment needs a proxy, configure the HTTP client explicitly using the Requests guide.
Local fixtures passed with Python 3.12.14 and Requests 2.34.2: endpoint parameters, header-only authentication, optional fields, price-category preservation, empty responses, malformed results, non-success responses and redirect refusal. No live Yelp key was used, and the fixtures do not establish account permissions or data coverage.
Treat pagination as a limit, not a complete crawl strategy
Yelpโs Places FAQ states a maximum of 50 results per individual search request and 240 per originating query. Use the documented offset behavior only within that boundary. Increasing page count does not create an unrestricted business export.
The rate-limit guide describes per-second and daily controls, response headers that report the remaining allowance, and HTTP 429 when a limit is reached. Read the allowance for your own plan rather than copying a quota from an old tutorial.
The example stops on an unsuccessful response. A production worker should distinguish authorization failure, quota exhaustion and transient service failure before retrying. Put a maximum number of attempts and a total request budget around any retry logic.
Design storage and display around the API conditions
The Places FAQ permits caching Places content for up to 24 hours and says Yelp business IDs may be stored indefinitely. Follow the conditions applicable to your integration; the ability to print JSON does not turn the other fields into a permanent historical dataset.
Store identifiers and operational metadata separately from cached business content. Give cached records an expiry and ensure downstream jobs do not silently preserve copies beyond the permitted use. Review logs, analytics exports and backup behavior as part of that design.
Before displaying Yelp ratings or business content, use the current display requirements linked from the developer documentation and the terms accepted for the application. This article does not supply a compliant presentation component or authorize redistribution under a different productโs license.
Keep review excerpts separate from full review collection
The Places introduction describes its reviews endpoint as returning up to three excerpts. That is a limited product response, not access to every review or the full text of every review visible on a business page.
If your application needs a different review-data product, establish its availability and permitted use with the provider. Do not treat a business-search key as authorization for an unlimited review collector.
For analysis, label exactly which sample was returned and when. A small excerpt response cannot support claims about every reviewer, a complete sentiment history or the reason a rating changed.
Use Yelpโs AI product when the requirement is conversational
The Yelp AI API reference documents conversational search at POST /ai/chat/v2, with a query and optional conversation context. It describes multi-turn answers and indicates that keys need the required endpoint scope. This is a separate interface from the Places search request shown above.
Choose that interface when the product requires a conversational answer supported by Yelpโs data. Choose structured Places search when the application needs the documented business fields. Do not assume the same key, rate limit or output schema applies to both.
If an external AI tool processes Places content, first confirm that the proposed processing and retention fit the applicationโs agreement. Keep source links and uncertainty attached to generated claims, and do not convert missing fields into invented recommendations. A model-training dataset requires its own rights assessment; a working API call does not establish those rights.
Sources and review scope
Sources reviewed October 11, 2026. Primary documentation checked October 11, 2026. Examples passed local fixtures; the Indeed CSV example also ran against its public dataset. No private API credentials or live host calendar were used. Each guide states its access and coverage limits.
Frequently asked questions
Collect from another source
Choose an available data product and preserve its coverage, permissions and field meanings.
Related workflows
Apply routing and authentication to the correct client.
Compare another APIโs price and coverage limits.
Plan request budgets and observable failures.