Under the proposal, if a model is trained on code protected by a CCAI license, the entire model is legally considered a derivative work of that code. This triggers a reciprocal obligation: the developer must also release the AI model under CCAI terms, including key transparency disclosures that proprietary companies typically keep secret .
As the authors explain in the PhilArchive version of their paper, the CCAI license rests on three pillars :
If an AI company trains a model on CCAI-licensed data, the license would mandate public disclosure of :
The aim is straightforward: prevent companies from taking community-maintained open-source code to build closed, commercial products while giving nothing back .
The paper was first released as a preprint in July 2025 and subsequently published in the Oxford Journal of International Law & Technology in early 2026 .
While theoretically compelling, the CCAI faces several significant legal obstacles identified by the authors themselves .
The most fundamental challenge is whether AI training qualifies as "fair use" under copyright law. If courts rule that training on copyrighted code is fair use — as suggested by some recent high-profile cases and settlements — CCAI's restrictions could crumble, because a model developer would not need permission to train in the first place . As an example, one source notes that a major AI company settled a copyright case for $1.5 billion, yet the judge still described AI training as "profoundly transformative" and fair use
.
CCAI's entire mechanism depends on a trained AI model being classified as a "derivative work" of its training data. It is far from settled that a neural network's learned weights, which encode statistical patterns rather than literal code, meet the legal definition of an adaptation or derivative work .
Copyright law differs dramatically across countries. A license enforceable under U.S. law may face an entirely different legal landscape in the EU, China, or other jurisdictions where AI training exceptions exist .
Even if CCAI clears the legal hurdles, enforcing it would be daunting. Modern AI models are trained on enormous, mixed datasets where tracing which lines of open-source code came from a specific repository is technically difficult .
The proposal represents a shift in strategy for the open-source community — from moral argument to legal mechanism . By attempting to make the license itself do the work of enforcing reciprocity, CCAI draws a line from training data through to the resulting model, creating a chain of obligation that current open-source licenses were never designed to handle
.
The debate over AI training and intellectual property is far from settled, and proposals like CCAI will influence both legal scholarship and the next generation of open-source licensing. Whether it survives courtroom scrutiny remains an open question — but the conversation it has started is already reshaping how developers think about releasing code in an AI-saturated world.