Skip to content
Back
ScrapeAny Team

ScrapeAny Team

Commercial Real Estate Scraping: LoopNet, CoStar & Beyond

Commercial Real Estate Scraping: LoopNet, CoStar & Beyond

Everything Good Is Behind a Paywall

Commercial real estate runs on data most participants can't afford. A CoStar subscription commonly runs tens of thousands of dollars per year per market, with multi-year contracts and per-seat pricing. For a large brokerage that's a line item. For a small investment shop or an independent broker, it's often the single biggest data expense on the books, or simply out of reach.

Meanwhile a meaningful slice of CRE market data sits in plain sight on public listing platforms: asking lease rates, asking sale prices, available square footage, broker contacts, days on market. It's fragmented and it's asking-price data rather than closed-transaction data, but it's public and it's current.

So let's set expectations honestly, because this is where CRE scraping conversations usually go wrong. You will not replicate CoStar's closed-deal comps database by scraping. Nobody will, that data isn't on the public web. What you can build is a live view of the asking side of the market (availability, pricing trends, new listings, time on market) across the platforms where inventory is publicly marketed. For a lot of use cases, that's actually the thing you needed.

What's Publicly Scrapable

LoopNet

LoopNet is the largest public-facing CRE marketplace in the US, covering office, industrial, retail, multifamily, and land, for sale and for lease. Public listing pages expose asking price or rent, building size, available suites, property type, year built, and broker information.

Two things to know before you touch it. First, LoopNet is owned by CoStar, which matters enormously for the legal calculus below. Second, listing depth depends on the advertiser's tier: premium listings carry rich detail while basic ones often say "contact broker" where the price should be. In our experience a noticeable share of LoopNet records come back with null price fields, and there's nothing to do about it except design your analysis to tolerate the gaps.

Crexi

Crexi has grown into the strongest LoopNet alternative, with a large for-sale marketplace and expanding lease inventory. Its pages are structured and comparatively generous: asking price, cap rate where disclosed, NOI on some investment sale listings, tenancy details, and auction listings with visible bid activity. If you only get one field pair out of a CRE scrape, make it Crexi's cap rate and NOI. There's not much else like it on the public web.

Brokerage sites

CBRE, JLL, Cushman & Wakefield, Colliers, Marcus & Millichap, plus hundreds of regional and boutique firms, all publish their own availability listings. Coverage overlaps with the marketplaces but not completely; some firms list on their own site first, or only. Each broker site is a small scraping target, but a portfolio of 50 to 100 of them gives you real depth in a market. Tedious, not hard.

Niche platforms and public records

Office-space platforms in the 42Floors lineage (many since consolidated or absorbed), coworking marketplaces, and sector-specific sites for self-storage or hospitality each cover slices the generalists miss. And the government layer rounds it out: county assessor records, SEC filings for REIT portfolios, municipal permit databases. More on that layer below, because it's underrated.

The Fields That Drive Decisions

FieldWhere It's PublicCaveat
Asking sale priceLoopNet, Crexi, broker sitesOften withheld on larger deals
Asking lease rate ($/SF/yr)All platformsNNN vs. gross basis varies — normalize carefully
Cap rateCrexi, some broker OMsBased on asking price and stated NOI, not closed pricing
Available SF / suite breakdownLoopNet, CrexiSuite-level data is where vacancy analysis lives
Property type/subtypeAll platformsTaxonomies differ per platform
Days on marketDerivable from scrape historyPlatforms rarely display it; you compute it
Occupancy / tenancyInvestment sale listingsSelf-reported by brokers
Broker and firmAll platformsUseful for market-share analysis

Two derived metrics deserve emphasis because they only exist if you scrape continuously rather than buying a snapshot. Time on market: track when a listing appears and disappears and you have a lease-up and sale-velocity signal that even the subscription datasets handle imperfectly. And asking-rate trajectory: repeated observations of the same suite reveal rate cuts and concession pressure quarters before they show up in anyone's published market report. Same principle that powers residential price tracking, applied to CRE.

The Section to Read Twice: CoStar

CoStar Group is the most litigious data company in real estate, and arguably in all of B2B data. It has spent two decades suing competitors, resellers, former customers, and scrapers, and it has won or extracted settlements repeatedly. It has historically used watermarked photos and seeded records to detect misappropriation. This is not a company that sends one warning email and moves on.

The rules that follow from this:

  • Never scrape CoStar's subscription products. Not with a borrowed login, not "just for research." Credentialed access is governed by contract, and breach of that contract is the fact pattern CoStar wins on.
  • Treat LoopNet as CoStar property, because it is. Public LoopNet pages are viewable without login, and scraping publicly accessible data has meaningful support under the hiQ v. LinkedIn line of cases. But LoopNet's terms prohibit automated access, and CoStar has shown it will pursue perceived misuse. Public surfaces only, modest volume, no photo harvesting, no wholesale republication. Anything more aggressive goes through counsel first.
  • Never republish scraped listing photos. CRE photography is copyrighted, CoStar owns an enormous photo library, and photo claims are the cleanest, most mechanical claims a plaintiff can bring. Collect the facts; leave the images alone.
  • Facts vs. expression is the line that keeps you safe. Asking rents, square footages, and addresses are facts, and facts aren't copyrightable. Descriptions and photos are expression. Build your dataset from the facts.

For the broader landscape (CFAA, terms-of-service risk, what "publicly accessible" actually protects), see our guides on real estate scraping law and web scraping legality generally. None of this is legal advice, and for CRE specifically, given CoStar's posture, paying for actual legal advice before you scale is money well spent.

The Free Layer Everyone Skips

The listing platforms get the attention, but the highest-trust CRE data in existence is free and sitting in government databases, and it pairs naturally with scraped listings.

County assessor records give you ownership entities, assessed values, building and lot sizes, and parcel numbers. Join a scraped listing to its parcel record and you learn who actually owns the asset (usually an LLC worth tracing) and how the ask compares to assessed value. Recorded deeds and mortgages give you actual transaction prices in many states via transfer tax stamps, plus loan amounts and lender identities. That's the closest public substitute for closed-comp data, with a lag of weeks rather than never. Building permits tell you where renovation and development money is flowing. And public REITs disclose portfolio-level occupancy and acquisitions every quarter in SEC filings, free.

None of this replaces listing data. All of it makes listing data more decision-grade. The strongest CRE datasets we deliver are joins: platform listing plus parcel record plus last recorded sale.

Who Actually Buys This Data

The clearest example we see is lenders. If you're underwriting a loan on an office building, live asking rents and rising sublease availability in that submarket are directly relevant to your risk model, and a quarterly PDF market report can't give you an early warning. Acquisition teams use the same feeds for deal sourcing: new investment listings across Crexi, LoopNet, and broker sites, with cap rates where disclosed, become an asking-price comp set for target submarkets.

Smaller brokerages and appraisers use scraped availability data to compete with CoStar-equipped rivals on surveys and pitch decks. Proptech startups bootstrap on public listings before layering in licensed sources; knowing exactly which fields are public determines what the MVP can promise. And corporate real estate teams just want to answer "what's available and what's it asking" across a few metros without buying a research platform.

What the Pipeline Actually Involves

CRE scraping is lower-volume but higher-complexity per record than residential scraping. A single metro might have a few thousand active commercial listings across platforms versus hundreds of thousands of residential, so politeness is easy and there's no excuse for hammering anyone.

The marketplaces do run modern bot detection. Expect fingerprinting-sensitive defenses rather than simple rate limits; the techniques in our anti-detection overview apply directly. But honestly, the scraping is the solved 20%. The real work is normalization. Lease rates get quoted per square foot per year in one market and per month in another (California industrial does this and it catches everyone eventually). NNN versus modified gross versus full-service quotes can differ by $10/SF in effective terms while looking identical in a column. Platform taxonomies disagree about what counts as "flex." Suite-level and building-level records get mixed. Making three platforms' records comparable is the actual job.

Dedup is mandatory too, since the same availability appears on LoopNet, Crexi, and the listing broker's own site with slightly different numbers. Address plus suite matching, with tolerance for formatting chaos, comes before any aggregate statistic means anything.

And store every observation with a timestamp from day one. Six months in, you have time-on-market and rate-trajectory history that nobody can buy retroactively. That history is the moat.

How We Run CRE Projects

CRE is a regular request for us, and the engagement shape is consistent. We collect from public listing surfaces only: LoopNet public pages, Crexi, brokerage sites, and public records. No credentialed platforms, and no listing photos in deliveries, ever. Normalization is built in (lease-rate basis conversion, property-type mapping across taxonomies, address and suite-level dedup) so delivered rows are comparable across sources. Every delivery cycle includes new listings, removals, and price or rate changes since the last one, which is where most of the analytical value lives anyway. The proxy and fingerprinting infrastructure is ours to maintain. Delivery is weekly or daily CSV/JSON, an API, or writes straight into your warehouse.

If the alternative is a CoStar contract you don't fully need, or an in-house project your team maintains forever, a managed pipeline on public sources is usually the pragmatic middle path.

Get a Live View of Your Market

You don't need a five-figure subscription to know what's listed, what it's asking, and how long it's been sitting. The asking side of the CRE market is public. It just takes disciplined collection and cleanup to turn it into a dataset.

Tell us your property types and target markets and talk to our team. We'll scope the public sources, flag the legal boundaries, and get you a working sample within days.

Ready to turn the internet into usable data?

Tell us about your project. We'll review it and get back to you within 24 hours.

Contact Us

Tell us about your scraping needs. Our experts will review your project and help you find the right solution. We typically respond within 24 hours.