← BACK TO FEED
web scrapingDMCAGoogleRedditSerpApi

SerpApi Tells Google and Reddit to Get Off the Internet's Lawn

Google and Reddit have been using the Digital Millennium Copyright Act (DMCA) to sue web scraping company SerpApi, arguing it illegally circumvents their anti-scraping technology to sell access to search results data. However, a court recently dismissed most of Google's lawsuit, finding Google lacked standing since it doesn't own the content in its search results, and Reddit's similar case faces the same legal obstacle. Legal experts warn that while the lawsuits aim to curb AI scraping, this broader push to restrict public web access could harm not just AI companies but also researchers, journalists, and archivists who depend on open data access.

Google lost a court battle last week and immediately announced it plans to keep fighting. The opponent is SerpApi, a web scraper that sells structured access to search data. The legal weapon of choice is the Digital Millennium Copyright Act, wielded in a way that legal experts are describing, charitably, as unusual.

Google sued SerpApi back in December, accusing it of bypassing anti-scraping protections and flogging an unauthorised API built on Google search results. The claim rested on the DMCA's anti-circumvention provisions, with Google arguing that its scraping protections exist to guard copyrighted content belonging to third-party rights holders, particularly those who license material to appear in Google's knowledge panels.

Reddit had tried the same move a couple of months earlier, suing both SerpApi and Perplexity for scraping Reddit content as it appears in Google search results. Google apparently took notes, citing Reddit's lawsuit when announcing its own.

The core problem with both cases is that neither company owns the content they're claiming to protect. A judge saw through this fairly quickly in Google's case, granting SerpApi's motion to dismiss at an unusually early stage. The ruling came down to a basic standing issue: Google doesn't own the content in its search results and hadn't demonstrated it was acting as an authorised representative of anyone who does.

Meredith Rose, senior policy counsel at Public Knowledge, told Ars Technica the outcome was notable. Courts don't often dismiss cases this early. Her reading: Google simply didn't plead enough about what copyrighted material it was actually protecting.

Reddit's hearing on a similar motion happened shortly after that ruling landed, which was unfortunate timing. Rose's assessment of Reddit's position is blunt. A rights holder, an exclusive licensee, or the party that deployed the technical protection measure can bring a DMCA circumvention claim. Reddit is none of those things when it comes to its content appearing in Google's search index.

Google isn't out of the game entirely. The judge left a narrow door open, giving it 21 days to file an amended complaint. The theory Google is reportedly pursuing is that some knowledge panel content is explicitly licensed from third parties who directly authorised Google to protect it. If that argument holds, Google might salvage a very limited version of the case.

But that argument carries its own risks. Rose pointed out that Google cannot easily claim knowledge panels are full of copyrighted material without raising questions about all the content it surfaces without explicit licenses. That opens a fair use fight Google would probably rather avoid. Threading that needle without implicating its own broader indexing practices will take some careful lawyering.

SerpApi, for its part, is framing this as a fight for the open web. Its client list includes Nvidia, Uber, and Adobe, all of which depend on its service for structured search data access. The company says business has kept growing during the litigation, but the legal uncertainty has been a drag on customers.

The company's broader argument is simple: public search results are public. Google is the world's largest web scraper. Asking a court to block others from accessing information that sits in the open while you index it yourself is a position that requires a certain confidence in judicial patience.

SerpApi was similarly pointed about Reddit, accusing it of trying to position itself as a toll collector over content its users created, content Reddit doesn't own. The filing characterised Reddit's real goal as consolidating enough control to eventually monetise that content, with users bearing the cost.

Rose broadly agrees with SerpApi's position on the merits, even if she's not endorsing the company personally. The wider concern she raises is about what this wave of litigation means for everyone else. Since mid-2023, a creeping re-enclosure of the web has been underway as platforms scramble to cut off automated access in response to AI scraping. That sweep catches far more than AI training pipelines. Researchers, journalists, archivists, and public health analysts all depend on the kind of large-scale automated access that is now getting caught in the crossfire.

Motivations vary across publishers. Some want licensing revenue. Some are protecting infrastructure. Some have genuine ethical objections to AI training. The result is the same regardless: the open web quietly closes.

As for how the Reddit case lands, Rose admitted even she isn't certain. Copyright law has a habit of producing unpredictable outcomes. In her words, a lot of judges just decide it on vibes.

Worth noting: Advance Publications, Reddit's largest shareholder, also owns Condé Nast, which owns Ars Technica, where much of this reporting originally appeared.

READ NEXT
Google Burned Through More Cash Than It Made Last Quarter. Thanks, AI.EU Forces Google to Open Android and Share Search Data. Google Is Not Pleased.EU tells Google to open up Android and share search data. Google is not pleased.