Advances in Consumer Research
Issue 8 : 178-186 doi: https://doi.org/10.5281/zenodo.22006117
Original Article
Training Data and the Reproduction Right: Is Ingestion by Large Language Models “Copying” under Section 14 of the Copyright Act, 1957?.
Loading Image...
 ,
Loading Image...
1
Research Scholar, Oriental University, Indore, Madhya Pradesh, India
2
Assistant Professor, Oriental University, Indore, Madhya Pradesh, India
Abstract

The training of large language models (LLMs) proceeds by ingesting vast corpora of text, a substantial portion of which is protected by copyright. This article asks a narrow but foundational question of Indian law: does that act of ingestion amount to "reproduction" within the meaning of Section 14 of the Copyright Act, 1957? Adopting a doctrinal and comparative method, the study disaggregates the training pipeline into six technically distinct operations and tests each against the statutory language of Section 14(a)(i), the definition of "infringing copy" in Section 2(m), and the storage-oriented amendments introduced by the Copyright (Amendment) Act, 2012. It argues that at least three of those operations — corpus acquisition, corpus curation and durable retention — constitute prima facie reproduction, that a fourth (tokenised batching) is defensible as transient and incidental, and that model weights are ordinarily not copies save in the narrow case of demonstrable memorisation. The article then evaluates the Delhi High Court's interim ruling in ANI Media Pvt. Ltd. v. Open AI OpCo LLC (2026), which held that such storage prima facie falls within the "private or personal use, including research" limb of Section 52(1)(a). It contends that this characterisation, while pragmatically attractive, strains the statutory text, conflates commercial secrecy with private use, and imports an open-ended proportionality analysis into a closed-list exception regime. The study concludes that the reproduction question in India cannot responsibly be resolved through interpretive elasticity and proposes a calibrated statutory text and data mining exception coupled with a collective licensing mechanism..

Keywords
Recommended Articles
Original Article
Women Empowerment Through Five-Year Plans in India (1951–2017): A Study in Policy Evolution and Social Development
Original Article
Geriatric Care Delivery under NPHCE in Urban Rajasthan: Evidence from Jaipur District
Original Article
AI Shopping Assistants and Electronic Purchase Intention: The Role of Trust and Privacy Concern
Original Article
Impact of AI and Gen-AI on Business Models
Loading Image...
Volume 3, Issue 8
Citations
312 Views
433 Downloads
Share this article
© Copyright Advances in Consumer Research