top of page

INDIA’S FIRST RULING ON AI AND COPYRIGHT TUSSLE – What does it mean?

  • Writer: Aman Shankar
    Aman Shankar
  • 4 days ago
  • 9 min read

The Delhi High Court’s judgment in ANI Media Pvt. Ltd. v. Open AI OPCO LLC marks India’s first significant judicial engagement with the copyright implications of generative AI. At the interim stage, the Court refused to restrain OpenAI from using publicly available news content to train its large language models, holding prima facie that such internal use may fall within the “private or personal use, including research” exception under Section 52(1)(a) of the Copyright Act, 1957 [1].


While the Court left several technical and legal questions open for trial, the ruling is significant because it adopts a purposive and innovation-sensitive reading of copyright law, recognises the public interest in AI development, and provides early comfort to AI developers that training on freely accessible material may not, by itself, amount to infringement in India.


Although the ruling is only a prima facie view at the interim stage, it is likely to influence how Indian courts, policymakers and businesses approach the relationship between copyright protection and AI innovation. The judgment is also one of the first detailed judicial examinations of large language models under the Copyright Act, and this article distils its key holdings, unresolved questions and practical takeaways for businesses.


The Court’s public-interest analysis is central to the judgment. It noted that Section 52 of the Copyright Act reflects a balance between the exclusive rights of copyright owners and the broader public interest in creativity, learning and dissemination of knowledge. The Court also referred to India’s ambition to become a global AI hub, recognised that LLM development depends on access to large volumes of publicly available data, and cautioned against interpretations of copyright law that could unduly hinder domestic AI innovation. This policy orientation significantly informed the Court’s balancing exercise.

I. BACKGROUND - THE PARTIES AND THE CLAIMS


ANI Media sued Open AI alleging two distinct infringements -

(i) a “training claim” - that Open AI copied and stored ANI’s data to train its LLM; and

(ii) a “reproduction or output claim” - that ChatGPT’s responses reproduce ANI’s works in a manner constituting substantial copying.


Open AI denied both allegations and argued, inter alia, that


(i) training occurred entirely outside India;

(ii) ChatGPT does not store or reproduce copyrighted articles but merely learns statistical relationships between words;

(iii) any use of copyrighted works during training is protected as fair dealing under Section 52 of the Copyright Act. 


The Court framed four principal issues covering jurisdiction, copyright infringement in AI outputs, copyright implications of AI training, and applicability of fair dealing doctrine. 



II. HOW THE COURT ADDRESSED THE KEY ISSUES


While deciding that the Hon’ble Delhi High Court has territorial jurisdiction over the subject matter, the court decided and opined on the following key issues –


A.  On the Reproduction or Output Claim


ANI argued that LLMs do not merely “learn language”; they retain expressive elements of copyrighted works. ANI, inter alia, argued that - tokenisation and de-tokenisation reconstruct the original text; ChatGPT memorises training data; outputs often reproduce original expression rather than merely underlying facts; Open AI itself has acknowledged “regurgitation” in certain situations; Open AI has entered licensing arrangements with international publishers, implicitly recognising the need for copyright licences; and that the public availability of news articles does not extinguish copyright, and publication on the internet cannot be treated as an implied licence for AI training.


Open AI adopted a fundamentally different technical and legal narrative. It, inter alia, argued that ChatGPT does not permanently store articles in a retrievable form; the model learns statistical relationships rather than expressive content; copyright protects expression not facts; news enjoys only “thin” copyright protection because facts remain free for public use; that the examples cited by ANI related to articles published after the relevant model’s training cut-off date; that any occasional verbatim reproduction is an unintended and rare “regurgitation” event that Open AI actively seeks to eliminate; and that identical prompts can produce different outputs, demonstrating that ChatGPT does not simply retrieve stored copies.


On the reproduction/ output claim, the Court carefully distinguished between memorisation and copyright infringement. The Court noted that despite repeated and adversarial prompts including explicit instructions to reproduce content “exactly”, ANI could not (prima facie) demonstrate that ChatGPT outputs constituted “substantial reproduction” of its articles or that Open AI permanently stores ANI’s articles. The responses were generated through retrieval-augmented generation (“RAG”) [II] and these responses were not substantially similar to ANI’s copyrighted articles (even though essence was similar), since ChatGPT added its own commentary making it dissimilar. The Court also noted that news content carries a higher threshold for similarity because its purpose is to report factual events.


However, the court observed that issues relating to memorisation, model architecture and regurgitation involve disputed technical questions requiring evidence during trial.


B. Storage for training and fair dealing under Section 52(1)(a) of Copyright Act


This issue was the heart of the judgment. The Court undertook an extensive interpretation of Section 52(1)(a) of the Copyright Act and analysed the legislative history, purpose test, fairness test, public interest, economic impact and commercial use. The Court ultimately held, prima facie, that Open AI’s storage and use of ANI’s works for training LLMs falls within the realm of fair dealing.


The Court rejected the orthodox narrow reading of Section 52 as a mere exception, instead holding it to be an “integral part of the Act defining users’ rights” that must be given a “broad, liberal interpretation”. Applying a two-step framework, i.e., ‘Purpose Test’ followed by a ‘Fairness Test’, the Court held:


On purpose - Commercial use is not excluded from Section 52(1)(a)(i) of the Copyright Act because the legislature expressly excluded it elsewhere in the Copyright Act but not here.


The Court also noted that the use by Open AI is only for the purposes of training its LLM, which is completely an internal process and does not involve providing the copyrighted material of ANI to any third party. Even after the training is complete, the training data is never made available either in its natural language or tokenized form to any person. The term “private” is broader than “personal” and extends to a corporation’s internal training environment.


The Court observed that when Section 52(1)(a) of the Copyright Act was last amended in 2012, the legislature could not have imagined the advent of artificial intelligence (AI) and/or LLMs. However, when the Court interprets the aforesaid provisions in the light of modern-day technological developments, the Court has to give a liberal and purposive interpretation using the ‘doctrine of updating construction’.


The Court endorsed Open AI’s argument by stating that the expression “research” should be given an updating construction by considering the modern-day technical advancements. With the advent of these technologies, research/ learning is no longer confined to humans. It is now being done through Artificial Intelligence. However, ultimately the research is at the behest of humans and for the benefit of humans.


For example, Section 52(1)(i)[III] of the Copyright Act exempts the act of reproduction by a teacher “in the course of instruction” from the ambit of infringement. If in future, a human teacher is replaced by an AI bot or Robot to say that the said exception could only be used by a human would be a regressive view. Such an approach would limit societal progress. Therefore, the acts of further research cannot be confined to acts of human being alone and the same would extend to machine learning as well.


Thus, ultimately the Court held that, on a prima facie view, the process of training LLMs underlying ChatGPT undertaken by Open AI using stored literary work of ANI falls under “private or personal use, including research” as provided in Section 52(1)(a) of the Copyright Act and fulfils the purpose test.


On fairness - The Court formulated three India specific factors broadly aligned with the Berne Convention’s three-step test, examining whether use is limited to internal training of LLM (yes), whether it causes market substitution or economic harm to ANI (the functions of ChatGPT and ANI are quite distinct and ANI has not adduced to show loss in revenue or market share), and whether it serves the public interest (yes since LLMs advance education, research, and accessibility).


Hence, both the fairness and purpose tests were satisfied.


C. The Balancing factor


While undertaking the balance of convenience and irreparable harm test, apart from weighing the larger public interest in development of AI, the court noted that ANI’s own USD 7.5 million license offer to Open AI demonstrated that its claim was quantifiable in money, negating any irreparable harm. The Court also noted that ANI failed to deploy available technical opt-outs (robots.txt) and led no evidence of lost subscribers. An injunction, by contrast, would harm millions of Indian users of Open AI, impede India’s AI development ambitions, and impose economically unviable licensing burdens on LLM training.



III. SOME QUESTIONS THAT MAY ARISE


While the judgment is first of its kind in India, several aspects may invite a debate.


First, treating a multinational corporation’s industrial-scale ingestion of works as “private use, including research” is an expansive reading of Section 52 of the Copyright Act, Traditionally, research exceptions have been viewed in the context of human scholarship. Extending them to commercial AI model training arguably stretches the statutory language beyond what Parliament may have originally contemplated. One of the amicus briefs (Professor Arul George Scaria) observed that while the training process is private, the use is public. The Court did not fully resolve this issue.


Secondly, the market-harm analysis places a heavy evidentiary burden on content owners (ANI) while discounting lost licensing revenue as an evident injury to ANI. That Open AI licensed the Financial Times, Associated Press, and Condé Nast suggests a functioning licensing market whose value the Court arguably underweights. Similarly, treating ANI’s failure to deploy robots.txt as quasi-acquiescence shifts the burden of self-protection onto rights-holders rather than imposing an obligation on AI firms to seek permission.


Notably, this judgment being a prima facie view at an interim stage, these tensions are somewhat mitigated by the Court’s own emphasis that all findings are prima facie and that contentious technical questions and remain open for trial.


IV. KEY TAKEAWAYS FOR BUSINESS


The judgment carries important lessons for multiple stakeholders.


For AI developers,


  1. It provides significant interim comfort that model training using publicly accessible material may receive judicial protection under India’s fair dealing framework.

  2. The AI developers must document that training data was sourced from freely accessible, non-paywalled sources, and not shadow libraries or circumvented access controls.

  3. AI developers should expressly document the public-interest purpose of their models, including how their systems support research, education, accessibility or innovation

  4. The judgment reinforces the importance of implementing safeguards that minimise verbatim outputs, prevent regurgitation and maintain appropriate attribution mechanisms where retrieval systems are used.

  5. Training data should be used within a controlled internal environment to strengthen any reliance on the defence of private use, including research, by AI developers.

 

For publishers and copyright owners 


  1. The decision makes clear that ownership of copyright alone will not automatically justify injunctive relief against AI companies. Plaintiffs may need robust technical evidence demonstrating memorisation, substantial reproduction and actual market substitution and loss in revenue.

  2. Publishers and copyright owners should implement and maintain technical opt-outs, including robots.txt and crawler blocking, because ANI’s failure to deploy available opt-out mechanisms was treated as a factor against interim relief.

  3. The publishers can also consider paywalls or access controls, since freely available content weakened ANI’s position significantly.

  4. From an interim relief standpoint, content owners bear the evidentiary burden of demonstrating concrete market harm, bare assertions of traffic diversion or unjust enrichment will not suffice.

  5. The judgment distinguishes between the training stage and the reproduction or output stage for copyright infringement analysis. Litigants should therefore plead and evidence these claims separately, supported by technical material on memorisation, regurgitation, and substantial similarity.


In brief, the judgment gives India’s AI ecosystem an important, though interim, signal that copyright law will not be applied mechanically to obstruct technological development. By treating internal model training on publicly accessible material as potentially protected fair dealing, the Court has created meaningful breathing space for AI innovation. At the same time, the ruling is not a blanket license : AI developers must maintain strong data-governance safeguards, while publishers must adopt technical controls and build evidence-led claims of actual market harm.


The final trial will determine the contours of this balance, but the decision already marks a significant step in shaping India’s copyright framework for the AI age.



REFERENCES


[I] Section 52. “Certain acts not to be infringement of copyright - (1) The following acts shall not constitute an infringement of copyright, namely, -

(a) a fair dealing with any work, not being a computer programme, for the purpose of –

(i) private or personal use, including research;

(ii) criticism or review, whether of that work or of any other work;

(iii) the reporting of current events and current affairs, including the reporting of a lecture delivered in public;

Explanation - The storing of any work in any electronic medium for the purposes mentioned in this clause, including the incidental storage of any computer programme which is not itself in infringing copy for the said purposes, shall not constitute infringement of copyright……………”

[II] When an LLM refers to information on which it was not trained, it is using RAG technique.

[III] Section 52. “Certain acts not to be infringement of copyright - (1) The following acts shall not constitute an infringement of copyright, namely, -

(i) the reproduction of any work –

(i) by a teacher or a pupil in the course of instruction; or

(ii) as part of the questions to be answered in an examination; or

(iii) in answers to such questions;”

Please feel free to reach out to our Team to discuss any of the Technology Law, Competition Law, International Trade and Policy Issues.

Comments


Commenting on this post isn't available anymore. Contact the site owner for more info.
bottom of page