📺 Stream EntrepreneurTV for Free 📺

Once Only for Huge Companies, 'Web Scraping' Is Now an Online Arms Race No Internet Marketer Can Avoid Companies don't just vacuum up data on competitors' prices; some gain advantage by distorting the picture competitors see.

By Eran Halevy

entrepreneur daily

Opinions expressed by Entrepreneur contributors are their own.

Tang Yau Hoong | Getty Images

In January 2017, news broke that Amazon had successfully managed to block bots from Walmart, which would scrape Amazon's listings "several million times a day." In the Reuters report, the Chief Executive of Boxed, a New York-based online wholesaler, spoke of scraping competitor prices every 20 minutes and adjusting accordingly, saying, "If we're not decently priced, we'll see it almost immediately [in sales declines]."

Web scraping is something of a secret. The original growth hack is used by Fortune 500 companies to stay competitive on price, inform strategy and measure customer sentiment.

Related: Who Owns the Data Your Business Uses? Not Knowing Could Hurt the Sale of Your Company.

Knowledge is power.

What started as a one-way tool to extract web data and increase competition for the benefit of consumers turned into an arms race in which the target websites try to sabotage the data collection in order to achieve a competitive advantage. Third-party services have emerged to help target websites identify and block competitors scraping their data.

More cunning is serving falsified information -- serving bots a higher-than-actual price, for example -- to foil the scraper's plan, rather than the mechanism.

To avoid the problem of falsified information (also called spoofing or cloaking) or getting blocked, companies have employed proxy networks, which are data-center-based routers through which they route, or proxy, their requests, to hide their identities. However, these networks can be identified by savvy companies. The need for a solution came in the shape of peer-to-peer networks (P2P), also known as the residential IP network.

P2P networks consist of consumers who are willingly routing some commercial requests through their IP in return for benefits (e.g: free use of applications, ad-free browsing, using the P2P network themselves and more). Thus, companies collecting intelligence through such networks can see the web as consumers see it without being at risk of getting spoofed or blocked.

The potential of scraping goes far beyond price wars. The internet is awash with unstructured data just waiting to be tapped.

Related: The Biggest Revelations and Strangest Moments From Mark Zuckerberg's Congressional Testimony

How companies use data scraping.

Some companies generate high-quality sales leads rather than buying contact lists and get higher quality prospects in the process. Some scrape job boards to find companies that are growing, and they monitor social media for firms that have just won funding.

For example, Proven is a skincare company that scrapes customer reviews to create highly personalized products. They've built a continually updated database of 8 million reviews, 100,000 beauty products and 4,000 scientific articles about skincare and the ingredients used in products. Their machine learning algorithm discovers the links between these to develop cleansers, creams and toners highly customized to age, skin type, ethnicity and conditions like acne. Customers fill out a questionnaire to fit them into an AI-assisted skin profile and are recommended a skincare regime.

The arms race is also rampant in the online advertising industry. For example, large ad publishers need to make sure that hackers don't use their programmatic advertising platforms to spread viruses and malware to the end user. So they constantly scrape the incoming ad servers to make sure the content is safe and legitimate.

The problem is that when the hackers recognize a publisher is calling their servers, they send a real ad so it appears all is well. If the ad publisher can appear as a regular online user, it will be served the fraudulent ad, which they can then prevent from being published. The ability to scan ad servers as regular consumers is how they keep their audience safe from fraudulent and potentially dangerous ads.

Get creative, and you can disrupt any industry with scraping.

Related: 4 Insanely Easy but Overlooked Tactics to Advance Your Entrepreneurial Career

Is it worth the fight? The bottom line is that web scraping is surreptitiously powering more online commerce than you realize. Fortune 500 companies remain competitive by algorithmically adjusting their prices in reference to the market, an impossible task without scraping.

Having these data collection machines be misled by the target websites means pricing based on false information. This is a strong enough motivation for businesses to win this scraping battle.

Eran Halevy

Freelance data security consultant and user acquisition expert

Eran Halevy is a freelance data security consultant and user acquisition expert. His nine years of experience includes working for IBM and Google.

Want to be an Entrepreneur Leadership Network contributor? Apply now to join.

Editor's Pick

Side Hustle

He Started a Luxury Side Hustle at Age 13 — Now the Business Earns More Than $10 Million a Year: 'People Want to Help You When You're Young'

Michael Morgan, now the owner of Iconic Watch Company, always had a passion for "old things" — and he turned it into a lucrative venture.

Thought Leaders

It's the End of the Entrepreneurial Era As We Know It

With the rise of advanced technologies and AI, are we losing all sense of the independent business person and entrepreneur?

Business Ideas

63 Small Business Ideas to Start in 2024

We put together a list of the best, most profitable small business ideas for entrepreneurs to pursue in 2024.