← All articles

How to Scrape Online Reviews as JSON: Yelp, Tripadvisor, Booking and More

Reviews tell you things no product page will. Why guests leave a hotel unhappy. Which feature makes buyers switch software. What employees really think of a company. If you want that signal at scale, you need the reviews as data, not as a web page you scroll through.

Collecting them is harder than it looks. This guide covers why, which review plugins you can use to get the data as clean JSON, how to monitor new reviews on a schedule, and how to merge reviews from different sites into one dataset.

Why Reviews Are Hard to Scrape Yourself

Every review site fights you in a slightly different way:

  • Only a few reviews are in the HTML. A Yelp business page or an Airbnb room page shows a handful of reviews. The rest load page by page from the site’s own internal API, so your HTML parser sees almost nothing.
  • Pagination is the expensive part. Tripadvisor serves 10 reviews per page. To get 200 reviews you load 20 pages, and every page is another chance to get blocked.
  • Anti-bot systems guard the best sources. Several major review sites sit behind commercial bot protection such as DataDome or Cloudflare. A plain requests call gets a challenge page instead of reviews.
  • Login walls hide part of the data. Amazon shows its full review list only to signed-in customers. Glassdoor shows signed-out visitors three reviews per view.
  • Every site has its own schema. Booking.com scores out of 10 and splits text into “liked” and “disliked”. Glassdoor has pros and cons. Yelp has one text field. You end up writing a parser per site and fixing it every time the markup changes.

If you have built a review scraper before, you know the parser is the easy part. Getting the pages reliably is the work.

The Review Plugins at a Glance

ScrapeUnblocker has a set of review plugins that take a business, product or listing and return its reviews already parsed. You send one POST request and get JSON back. No account on the review site, no login, no CSS selectors.

PlatformEndpointReviews per callGood to know
Yelp/local/yelp-reviewsup to 500Every review page, one language per call
Tripadvisor/travel/tripadvisor-reviewsup to 200Hotels, restaurants and attractions, owner responses
Airbnb/travel/airbnb-reviewsup to 500Every review, with Airbnb’s translation
Booking.com/travel/booking-reviewsup to 270Newest 15 per language, 18 languages
Amazon/marketplace/amazon-reviewstop 5 to 13Rating, star histogram and the “Customers say” summary
Glassdoor/reviews/glassdoorup to 25A public sample, US site only
Capterra/reviews/capterraup to 500Pros, cons and reasons for switching
Trustpilot/reviews/trustpilot20 per pageTrustScore, star breakdown and company replies

Each plugin has its own page in the scraper catalogue with the full field list, a sample response and its public benchmark.

Your First Call

Here is the Yelp plugin in Python. The business parameter takes a Yelp URL, the business alias or its id:

import os
import requests

resp = requests.post(
    "https://api.scrapeunblocker.com/local/yelp-reviews",
    headers={"X-ScrapeUnblocker-Key": os.environ["SU_API_KEY"]},
    params={"business": "gary-danko-san-francisco", "max_reviews": "20", "sort": "newest"},
    timeout=180,
)
resp.raise_for_status()
data = resp.json()

biz = data["business"]
print(biz["name"], biz["rating"], biz["reviewCount"], biz["ratingHistogram"])

for review in data["reviews"]:
    print(review["date"], review["rating"], review["text"][:80])

The response has two parts. business is the summary: rating, review count, the 1-5 star histogram and review counts per language. reviews is the list itself, each with id, rating, full text, date, author, photos, reactions and the owner’s reply.

Paging works the same way across the plugins. You ask for a number of reviews with max_reviews or max_results, and the response tells you whether there is more (hasMore) and where to continue (nextStart on Yelp, nextPage on Tripadvisor and Capterra, nextPage or nextOffset on Airbnb). It also reports pagesFetched, which is what you are billed for.

Monitor New Reviews on a Schedule

Most review projects are not one-off downloads. You want to know when a new review lands, especially a bad one. The pattern that works:

  1. Ask for the newest reviews first (sort=newest on Yelp, sort=most_recent on Airbnb; Tripadvisor is newest first by default).
  2. Keep the batch small, because you only need what changed since the last run.
  3. Store every review id you have seen and only act on ids that are new.
  4. Where the plugin supports it, pass a since date to drop older reviews before they reach you.

Booking.com is built exactly for this. It only exposes the newest 15 reviews per language, so you poll it and pass since as a YYYY-MM-DD date.

import os
import sqlite3
from datetime import date, timedelta

import requests

API = "https://api.scrapeunblocker.com"
HEADERS = {"X-ScrapeUnblocker-Key": os.environ["SU_API_KEY"]}
since = (date.today() - timedelta(days=7)).isoformat()

TARGETS = [
    ("yelp", "/local/yelp-reviews",
     {"business": "gary-danko-san-francisco", "sort": "newest", "max_reviews": "20"}),
    ("tripadvisor", "/travel/tripadvisor-reviews",
     {"url": "https://www.tripadvisor.com/Hotel_Review-g60763-d93589-Reviews-The_Michelangelo_New_York-New_York_City_New_York.html",
      "max_reviews": "20"}),
    ("airbnb", "/travel/airbnb-reviews",
     {"room": "35283352", "sort": "most_recent", "max_reviews": "50"}),
    ("booking", "/travel/booking-reviews",
     {"hotel": "gb/the-savoy", "languages": "en,de", "since": since}),
]

db = sqlite3.connect("reviews.db")
db.execute("CREATE TABLE IF NOT EXISTS seen (platform TEXT, id TEXT, PRIMARY KEY (platform, id))")

for platform, path, params in TARGETS:
    r = requests.post(API + path, headers=HEADERS, params=params, timeout=180)
    r.raise_for_status()
    data = r.json()
    new = []
    for review in data.get("reviews", []):
        cur = db.execute("INSERT OR IGNORE INTO seen VALUES (?, ?)", (platform, str(review["id"])))
        if cur.rowcount:
            new.append(review)
    db.commit()
    print(f"{platform}: {len(new)} new reviews, {data.get('pagesFetched')} pages billed")

Run it from cron or any scheduler once a day. The first run fills the table; every run after that only reports what is new. From there, send the low ratings to Slack or email, or push everything into your warehouse.

Merge Reviews Into One Schema

Once you track several platforms, you want one table, not seven. The differences are small but real: Booking.com scores out of 10, Amazon calls its text body, Glassdoor and Capterra split it into pros and cons, and the reply field has a different name on every site. A small mapping table handles it:

FIELDS = {
    # platform: (text fields, date field, author path, reply field, rating scale)
    "yelp":        (["text"], "date", ("user", "name"), "ownerResponse", 5),
    "tripadvisor": (["title", "text"], "publishedDate", ("user", "name"), "ownerResponse", 5),
    "airbnb":      (["text"], "date", ("reviewer", "name"), "response", 5),
    "booking":     (["title", "positiveText", "negativeText"], "reviewDate", ("reviewer", "name"), None, 10),
    "amazon":      (["title", "body"], "date", ("author",), None, 5),
    "glassdoor":   (["title", "pros", "cons"], "date", ("jobTitle",), "employerResponse", 5),
    "capterra":    (["title", "comments", "pros", "cons"], "date", ("reviewer", "name"), "vendorResponse", 5),
}

def normalize(platform, r):
    text_fields, date_field, author_path, reply_field, scale = FIELDS[platform]
    author = r
    for key in author_path:
        author = author.get(key) if isinstance(author, dict) else None
    return {
        "platform": platform,
        "id": str(r["id"]),
        "rating": round(r["rating"] / scale * 5, 1) if r.get("rating") is not None else None,
        "text": "\n\n".join(r[f] for f in text_fields if r.get(f)),
        "date": (r.get(date_field) or "")[:10],
        "author": author,
        "replied": bool(reply_field and r.get(reply_field)),
    }

Every review now has a rating on a 1-5 scale, one text field, an ISO date and a replied flag. That is enough for dashboards, alerts and sentiment analysis across all your sources.

What You Can Build With Review Data

  • Reputation monitoring for hotels and restaurants. Track new reviews on Tripadvisor, Booking.com, Airbnb and Yelp in one place, alert on low ratings, and measure how often and how fast the owner replies compared to competitors nearby.
  • Product research on Amazon. The “Customers say” data lists the aspects buyers talk about (quality, fit, sound) with sentiment and mention counts. Compare it across competing products to see which weakness you can fix.
  • SaaS competitive intelligence. Capterra reviews include the reasons for choosing a product, the reasons for switching and the products a buyer switched from. Use mode=search with a keyword such as crm to find every product in a category, then see who loses customers to whom, and why.
  • Employer research. Glassdoor reviews come with six sub-ratings, CEO approval and business outlook. Track how they move for your company and the companies you hire against.
  • Multilingual sentiment. Booking.com returns reviews in up to 18 languages, and Airbnb returns its own translation next to the original text, so you can analyze everything in one language.
  • Datasets for AI. Normalized reviews are a clean input for an LLM that extracts topics, summarizes complaints or tags feature requests.

Know the Limits Before You Build

The plugins return what each site shows publicly, and that differs a lot:

  • Amazon returns the public slice of a product page: rating, histogram, “Customers say” and the top reviews. When Amazon shows a sign-in prompt instead of reviews, reviewsGated is true and the list is empty. It is not a full export.
  • Glassdoor shows signed-out visitors three reviews per view. The plugin merges several sort orders and star ratings to reach up to 25 per call. Treat it as a sample, not a history.
  • Booking.com gives the newest 15 reviews per language. A capped flag tells you when older ones exist that cannot be reached.
  • Yelp lists one review language per call. reviewsCountByLanguage tells you whether another pass is worth it.
  • Capterra only offers its Most Helpful order, with no date filter, and stops at 100 pages per product.

Billing follows the pages each site serves: one request per 10 reviews on Yelp, per review page on Tripadvisor (10 reviews, 15 for restaurants), per 50 on Airbnb, per 25 on Capterra, per language on Booking.com and per view on Glassdoor. Check pagesFetched in each response and size max_reviews to what you actually need.

One more thing: reviewer names and profiles can be personal data. Store only what your use case needs. Our guide on whether scraping websites is legal covers the basics.

FAQ

Can I get every review a business has? On Airbnb and Yelp, yes: the plugins page through all reviews, up to 500 per call. Tripadvisor gives up to 200 per call and you continue with the next page. Capterra reaches up to 100 pages of 25. Amazon, Glassdoor and Booking.com only expose a public slice, so you get a sample or the newest reviews.

How do I collect only new reviews? Sort newest first, keep each run small, store the review ids you have seen and pass a since date where the plugin supports it. The monitoring script above does exactly that.

Do I need an account on the review site? No. None of the plugins log in or need a partner API key. They read what the site shows to a signed-out visitor.

How much does it cost? Each billed page is one request, and one request is one credit: EUR 1.00 per 1,000 on pay as you go, down to about EUR 0.55 per 1,000 on the Ultimate plan. Blocked requests and timeouts are not billed.

Getting Started

Review data is only useful if it keeps flowing. The hard parts are the anti-bot walls, the pagination and the schema drift, and those are exactly what the plugins take off your hands, so you can spend your time on the analysis.

Browse the review plugins in the ScrapeUnblocker scraper catalogue, then create a free account and run the first example above with your own business or product. The free trial includes 500 requests, no card needed.

Try ScrapeUnblocker free

95%+ success rate · from 0.55€ per 1,000 calls · 500 free requests on signup.

Try it free → See pricing