District court largely denies motion to dismiss Reddit’s DMCA anti-circumvention claims against data-scraping tool provider SerpApi and AI answer engine Perplexity, finding that CAPTCHA and bot-detection systems preventing automated scraping of Reddit content in Google search results qualify as technological access control measures under DMCA, even if that same content is available to human users.
Plaintiff Reddit, Inc., operates a “blue-chip” online discussion forum that generates traffic and data inputs from more than 100 million unique users per day. Its platform hosts billions of posts and comments across hundreds of thousands of communities known as “subreddits,” which house “nearly two decades of human conversational data” that is “invaluable to AI companies” for training large language models.
Reddit has licensing agreements with Google for its content to be displayed in Google search results and for use of that content in AI tools offered by Google and OpenAI. Central to these agreements are Reddit’s content protection policies, which include Google’s technological protective measure named SearchGuard, which uses JavaScript challenges and CAPTCHAs to block automated systems from accessing Google search engine results pages (SERPs).
Reddit brought Digital Millenium Copyright Act (DMCA) claims against SerpApi LLC and Perplexity AI, Inc., alleging that SerpApi provides a scraping tool that circumvents SearchGuard by using proxy servers and technologies that “mimic human behavior.” Perplexity allegedly uses SerpApi’s tool to access Google SERPs, extract snippets of Reddit content, and feed that data into a retrieval-augmented generation database to power its AI responses—essentially circumventing Google’s SearchGuard system. Reddit’s first amended complaint asserted claims for circumvention of access control measures under Section 1201(a)(1)(A), trafficking in circumvention technology under Sections 1201(a)(2) and 1201(b), and New York common law claims for unfair competition, unjust enrichment and civil conspiracy.
SerpApi and Perplexity moved to dismiss. The district court denied the motions as to all of Reddit’s DMCA claims, except its Section 1201(b) claim for trafficking in technology that enables copying of copyrighted works. The court also dismissed Reddit’s state law unjust enrichment and unfair competition claims as preempted by the Copyright Act, but allowed the state-law civil conspiracy claim to move forward.
As a threshold matter, the court found that Reddit had Article III standing. As to injunctive relief, the court held that Reddit plausibly alleged injury to the copyrights in its own authored content (tens of thousands of posts and comments), which has a “close historical analogue” to copyright infringement. The court also found that Reddit adequately alleged reputational harm based on defendants’ undermining of Reddit’s protections of users’ content, paying particular attention to the defendants’ end run around Reddit’s “deletion” feature that the court found enables candid participation in discussions on sensitive topics such as “mental health, fertility, addiction, trauma, personal relationships, and other sensitive subjects.”
As to monetary damages, the court found Reddit plausibly alleged lost profits, reasoning that SerpApi’s tool eliminated Perplexity’s incentive to enter into a licensing agreement with Reddit as other AI companies had done. The court next addressed whether Reddit falls within the zone of interests that the DMCA was enacted to protect. Citing Section 1203’s broad grant of a cause of action to “[a]ny person injured,” the court held that Reddit’s injuries are closely related to the DMCA’s purpose of making digital networks “safe places to disseminate and exploit copyrighted materials.” The court rejected SerpApi’s argument that only copyright owners may bring DMCA claims, noting that the statute’s broad language covers anyone “arguably within the zone of interests” protected by the statute, not merely legal or beneficial owners of copyright.
Turning to Reddit’s circumvention claim, the court found all elements plausibly pled. First, the court held it plausible that at least some snippets of Reddit content appearing in Google SERPs possess the minimal level of creativity required for copyright protection, rejecting SerpApi’s argument that snippets are per se uncopyrightable due to their length (here the court noted that the rule against copyrighting “words and short phrases” applies to one-to-four-word expressions like “names, titles, and slogans,” not multi-sentence excerpts of the kind at issue here).
The court also held that SearchGuard qualified as a “technological measure that effectively controls access” under Section 1201(a)(3)(B), rejecting defendants’ argument that a measure cannot “control access” if it leaves the same content accessible to human users. The court reasoned that SearchGuard does not leave content “freely readable” but instead differentiates between automated and human users, analogous to “a facial-recognition technology programmed to open the door of a home for residents but not for other visitors.”
Reddit also plausibly authorized Google to implement SearchGuard. Looking at the user agreement’s grant of rights to Reddit and Reddit’s partnership agreement with Google, the court held that Section 1201 does not require the copyright holder to have specifically authorized the particular technological protective measure that was circumvented but requires only that the holder broadly authorized the implementation of measures to control access. The court distinguished the Northern District of California’s Google LLC v. SerpApi decision, which dismissed Google’s own DMCA claims for failure to allege facts regarding licensing agreement terms, because Reddit’s complaint went beyond a bare allegation that Google “ha[d] licenses” by alleging specific restrictions on use and requirements to protect against unauthorized access.
Lastly, the court found circumvention was plausibly alleged as to both defendants. As to SerpApi, the complaint alleged its tool uses proxies, fake user-agent strings and “ludicrous speed” features to bypass SearchGuard. As to Perplexity, the court held that Reddit plausibly alleged that Perplexity was a direct circumventor because it sets the parameters for circumvention and because it “conduct[ed] the actual queries itself.” In this vein, the court clarified that Section 1201(a)(2) covers actors like SerpApi that manufacture circumvention tools, while Section 1201(a)(1)(A) “targets the use of a circumvention technology” and reaches actors like Perplexity. The court declined to resolve the circuit split on whether an “infringement nexus” is required but held that it was met in any case. The court also denied SerpApi’s motion to dismiss Reddit’s Section 1201(a)(2) trafficking claim, finding that the element of being designed or marketed for circumvention was not disputed.
The court granted SerpApi’s motion to dismiss the Section 1201(b) claim, however, holding that SearchGuard is an “access control” measure under Section 1201(a) and not a “rights protection measure” under Section 1201(b) (as SearchGuard blocks access to content but does not control what users do with content once it is obtained). The court noted that collapsing the distinction between the two subsections would be “contrary to the DMCA’s text and structure.”
As to Reddit’s state law claims, the court dismissed the unjust enrichment and unfair competition claims as preempted by the Copyright Act, finding that the gravamen of each claim was defendants’ unauthorized reproduction of copyrighted content and that circumvention did not add a qualitatively different “extra element.” The court permitted Reddit’s civil conspiracy claim to proceed, however, reasoning that the DMCA violation served as the underlying tort and that defendants’ agreement to engage in circumvention constituted an “extra element” qualitatively different from a claim of copyright infringement alone.
Summary prepared by Tal Dickstein and Ezra Isaacson
-
合伙人