Skip to content
Back
ScrapeAny Team

ScrapeAny Team

Scraping Foreclosure & Distressed Property Data

Scraping Foreclosure & Distressed Property Data

The Case for Chasing Distressed Inventory

Every investor knows the cliché: you make your money when you buy. What the cliché skips is where the good buys come from. Year after year, the widest margins sit in distressed inventory. Foreclosures, pre-foreclosures, short sales, REO, auction properties. Banks and trustees sell for speed, not price, and discounts of 10 to 30 percent below comparable retail listings are routine, not lucky.

The catch is that this inventory doesn't live anywhere in particular. It's spread across auction platforms, bank REO portals, county sheriff-sale pages, court dockets, and the foreclosure filters buried inside Zillow. There is no feed. The investors who consistently find deals are the ones watching all of these sources every day, which makes this a data collection problem wearing a real estate costume.

Timing makes it worse. A pre-foreclosure filing is a signal that decays by the day. Once a property shows up on a public auction listing, every investor in the county has seen it. The edge is seeing the filing within a day or two of it appearing, running your numbers, and reaching the owner or lender before the crowd does. You will not get there by checking a dozen county websites by hand. People try. It lasts about three weeks.

Where the Inventory Actually Lives

Five source categories, in rough order of how pleasant they are to work with.

Auction platforms

Auction.com is the big one, the largest online real estate auction marketplace in the US, listing bank-owned properties and foreclosure sales with auction dates, starting bids, occupancy status, and property details. It also runs serious anti-bot protection. A casual Python script starts eating 403s within the first few hundred requests; this is a site you scrape with real infrastructure or not at all.

Hubzu, owned by Altisource, carries REO and short-sale inventory and usefully shows bid history and reserve status right on the listing page, so you can gauge actual demand instead of just asking prices. Then there's the long tail: Xome, ServiceLink Auction, RealtyBid. Small individually, but lenders often route an entire portfolio to a single platform, so skipping the small ones means silently missing whole servicers.

Listing sites with foreclosure filters

Zillow, Redfin, and Realtor.com all surface foreclosure and pre-foreclosure inventory behind filters. Zillow's pre-foreclosure records come from public filings: properties that aren't for sale yet but have received a notice of default or lis pendens, with an estimated value and the filing stage attached. Mechanically this is ordinary listing scraping, and our Zillow guide covers it. One caveat we learned by getting burned: the foreclosure-specific fields change format noticeably more often than the standard listing fields do. Budget to fix that parser a few times a year.

County and sheriff-sale pages

This is where the freshest data lives, and where the pain starts. Foreclosure is a legal process run at the county level, and the US has over 3,000 counties. Judicial states like Florida, New York, and Ohio publish sale notices through court systems or sheriff's offices. Non-judicial states like California, Texas, and Georgia publish trustee sale notices, often through county recorders or legal newspapers.

In practice: some counties run modern searchable portals. Some publish a weekly PDF. Some post scanned images of paper notices. A few still publish only in the local legal newspaper, some of which keep their digital archives behind a paywall. Every county is its own scraper.

Aggregators

RealtyTrac, Foreclosure.com, PropertyShark in some markets. They do the county-level collection themselves and sell access, which is a legitimate shortcut. But they typically lag the source filings by days, coverage quality swings from county to county, and pricing scales with geography. The pattern we see most among serious investors: aggregator data for broad national coverage, direct county collection for the three or four markets they actually buy in. That split is sensible and we'd recommend it to anyone.

Bank and GSE REO portals

Fannie Mae's HomePath, Freddie Mac's HomeSteps, and individual bank REO pages like Wells Fargo's. Volume is nothing like the 2008 era, but the listings are clean, structured, and often priced to move. Low effort to add. Include them.

The Pre-Foreclosure Window

The distressed lifecycle runs through predictable stages, and competition stacks up at each one:

StagePublic SignalCompetition LevelTypical Discount
Missed paymentsNone (private)NoneN/A
Notice of Default / Lis PendensCounty filingLowHighest potential
Auction scheduledSheriff/trustee noticeMediumHigh
Auction saleAuction platform listingHighMedium
REO (bank-owned)Bank portal, MLSHighLower but predictable

The window between the initial default filing and the scheduled auction, typically 90 to 200 days depending on state law, is where you can negotiate directly with the owner, structure a short sale with the lender, or buy before the property ever reaches competitive bidding. A notice of default tells you three things: this owner has a problem, the clock is running, and almost nobody else knows yet.

Wholesalers build entire businesses on contacting these owners within days of the filing. That model has one dependency, and it's the pipeline. It does not work on aggregator data that's four days stale, because by then the owner's mailbox is already full.

Which Fields to Collect

Address and price get you nowhere in distressed. The fields that actually drive decisions: the filing type and date (NOD, lis pendens, notice of trustee sale, judgment), since the type tells you the stage and the date starts your clock. The auction date and location, including postponements, which are worth tracking as events in their own right because repeated postponements usually mean a workout is in progress. The opening bid or judgment amount, which is the lender's floor and one half of your discount math; an estimated market value from an AVM or your own comps is the other half. Outstanding loan balance and lender name where the filing discloses them, because the equity position decides whether a short sale, a subject-to purchase, or an auction bid is the right play. Owner name and occupancy status, since an occupied property carries an eviction timeline you have to price in. And the case number or trustee reference, which is how you follow one property across multiple documents.

What the county filing will not give you is beds, baths, and square footage. County clerks do not care how many bathrooms a defaulted property has. So the complete record is always a join: legal and financial fields from the filing, physical characteristics and value estimates from listing sites or assessor data, matched on address and parcel number. Address matching against legal filings is genuinely miserable. Filings use legal descriptions, misspell street names, and drop unit numbers, and unit numbers are exactly the thing that breaks address matching. We wrote up the whole normalization problem separately in our data quality guide.

3,000 Counties, 3,000 Formats

Scraping Auction.com is a solvable anti-bot problem. County data is a different kind of hard: not one difficult scraper, but hundreds of easy-to-medium scrapers that all break independently.

The good case is a searchable portal, often built on a common government vendor platform like Tyler Technologies, which means scrapers can be partially templated. Even then you deal with session handling, CAPTCHA walls on court systems, and per-search result caps. The medium case is a weekly PDF whose layout changes whenever the clerk's office changes software, or staff. The bad case is scanned images of paper notices, where OCR on dot-matrix-quality scans tops out well short of reliable and you budget for manual QA. And a handful of jurisdictions still publish only in print.

Then there's maintenance. County websites change without notice, and court portals go down for weekend maintenance more often than any other class of site we scrape. A pipeline covering 50 counties will see several scrapers break in any given month, and in our experience the breaks cluster on Friday evenings, for reasons nobody has ever explained to us. This is the honest reason most investors either accept aggregator lag or hand the problem to a managed service: the marginal county is cheap to add and expensive to keep alive.

If you're weighing build versus buy in general, our real estate scraping overview walks through the trade-offs. County fragmentation multiplies every cost on the DIY side of that ledger.

Compliance Notes

Foreclosure filings are public records. Collecting them is the same activity title companies and data vendors have performed for decades, and scraping public court and recorder data sits on solid ground.

The sensitivity is in the people, not the collection. Pre-foreclosure records identify homeowners in financial distress. If your use case involves direct outreach, TCPA and state solicitation rules apply, and several states restrict foreclosure-rescue solicitation specifically. That's a use-of-data question rather than a scraping question, but plan for it before you build the mailing list, not after.

On the platform side, auction sites and listing portals prohibit automated access in their terms. Scrape only what's visible without a login, keep request rates modest, and don't republish their proprietary content wholesale. Skip sealed or restricted court records entirely. And when a county gates bulk access behind a formal records-request process, consider just filing the request. It's often cheaper than the scraper would have been, which is a slightly deflating thing for a scraping company to admit, but there it is.

What We Deliver for Distressed Data

County fragmentation is exactly the problem a managed service exists to absorb. A typical foreclosure engagement with us covers the auction platforms, the listing-site foreclosure layers, and the county sources for your target geographies, merged into one dataset, with new NODs, lis pendens, and scheduled sales delivered within a day of publication. Speed is the whole point of this data, so we treat the daily cadence as non-negotiable.

Records arrive joined: each filing matched to property characteristics and a value estimate, so a row is decision-ready rather than a legal citation. The anti-bot infrastructure for Auction.com and the listing platforms is ours to maintain, and when a county changes its portal, that's our scraper to fix, usually before you'd have noticed it broke. Delivery is CSV, JSON, API, or direct to your database, on whatever cadence your acquisition workflow runs.

For an investment team the build-vs-buy math tends to settle quickly. One engineer babysitting 40 county scrapers costs more than the data does, and the scrapers still break on weekends.

See the Filing First

Distressed investing rewards whoever sees the filing first. If you're screening deals from a feed that's four days stale, you're bidding against everyone else who bought the same feed. A pipeline that pulls daily from auction platforms, listing sites, and county sources, normalized and joined, is a durable edge precisely because it's tedious to build.

Tell us which markets and which stages of the foreclosure lifecycle you care about, then talk to our team. We'll scope the sources and get a working data sample in front of you within days.

Ready to turn the internet into usable data?

Tell us about your project. We'll review it and get back to you within 24 hours.

Contact Us

Tell us about your scraping needs. Our experts will review your project and help you find the right solution. We typically respond within 24 hours.