Was My Work in Your Training Data? Japan's AI Code Puts the Question in Writing

Introduction
From 26 October 2026, the Cabinet Office's Intellectual Property Strategy Promotion Office will accept filings from generative AI businesses that accept Japan's Principles-Code on intellectual property protection and transparency for generative AI. The government's Intellectual Property Strategy Headquarters published the code on 25 August 2026. It is not legally binding and works on a "comply or explain" basis. It applies to businesses outside Japan whose generative AI systems or services are provided to Japan, including, but not limited to, where Japanese nationals can use them.
Principle 1 asks a business to publish an outline of its models, training data and safeguards. Principles 2 and 3 let two kinds of requester ask a business about specific URLs or other information, limited to what the business can easily access and verify. A person who is in, or preparing for, legal proceedings over their work can ask whether information they identify is in the training data. A user of the service can ask the same about a URL where content identical or similar to their output appears. The answer bears on reliance, one of the two elements of a Japanese copyright claim against AI output.
Reliance is where the training data matters
Article 30-4 of the Copyright Act (著作権法, Chosakuken-hō) allows a work to be used for purposes that do not involve enjoying the thoughts or sentiments expressed in it, such as data analysis, to the extent necessary, unless the use would unreasonably prejudice the copyright owner's interests. The General Understanding on AI and Copyright, compiled by the Legal Subcommittee of the Copyright Subdivision of the Council for Cultural Affairs on 15 March 2024, reads this as allowing, in principle, the collection and reproduction of works as AI training data without permission.
The harder questions arise at the output stage. An infringement claim against AI output needs similarity and reliance (依拠, ikyo), meaning that the output was derived from the existing work. On reliance, the General Understanding sets out three cases:
- The AI user knew the existing work and used generative AI to produce something with its creative expression. Reliance is established.
- The AI user did not know the work, but the work was used in developing and training the generative AI. Access is then objectively established, so where a similar output is generated, reliance is normally presumed.
- The AI user did not know the work, and the work was not used in training. A similar output is a coincidence, reliance is not established and there is no infringement.
The General Understanding adds two qualifications. Even where the work was used in training, reliance may be denied if it is technically ensured that the creative expression of training works is not output, for example by filtering at the output stage, and the user argues the facts supporting that. And where reliance is presumed, the alleged infringer can argue that the work was not in the training data, as an indirect fact that negates reliance. The General Understanding also notes that the developer may be liable as the normative actor when a user is accused of infringement over a work the model was trained on.
So for an AI user who did not know the original work, whether that work was in the training data moves reliance one way or the other, subject to what the model's output safeguards can show. In practice, the business that built or provides the model is best placed to check it.
Principle 2: requests from rights holders
Under Principle 2, a person who is in, or preparing for, litigation, mediation, ADR or other legal proceedings to protect their rights in films, music, theatre, literature, photographs, manga, animation, computer games or other such works, or their lawyer or another agent authorised to act in court, can ask a covered business whether a URL or other information they identify is included in its training or validation data. The question is limited to information the business can easily access and verify. The code notes that it does not contemplate a creator asking, beyond that reference information, whether the work itself was among the crawl targets. A provider that cannot answer should give the name of the developer of the model it uses.
The request must give reasons showing that the requester falls within that group, state the purpose with a pledge not to use the answer for anything else, and identify the reference information and the grounds for asking this business about it.
This gives rights holders a way to learn, outside the court process and both before and during proceedings, whether a URL or other information connected with their work was in the training data. Under the General Understanding, use of the work in training is what supports a presumption of reliance. The code itself names other ways of collecting information, such as party inquiries under Article 163 of the Code of Civil Procedure (民事訴訟法, Minji Soshō-hō), which one party sends to the other while a suit is pending, and a motion for a document production order under Article 221.
Where the business would itself be the opposing party in the proceedings, the code asks the business to consider its response in light of its strategy for those proceedings. It allows measures against abusive requests, such as a certain fee or a cap on the number of requests within a set period, but warns against measures that would discourage, hinder or make people give up making requests.
Principle 3: requests from users
A person who generated content with the service can make a request by providing the content, the prompt used and the URL or other information where identical or similar content appears. They can ask whether that URL or information is included in the training data, again limited to information the business can easily access and verify. The request must also state the purpose and identify the grounds for asking this business about that information. The requester must pledge not to use the answer for any purpose other than the one stated, or for the purpose of filing a lawsuit or applying for mediation or ADR.
For a game studio or a publisher that uses generative AI in production, the worrying scenario is an output that turns out to resemble work published somewhere else. Under the General Understanding, the fact that the source was not in the training data is what such a user would want to be able to show. Principle 3 gives them a way to ask before an AI-assisted asset goes into a commercial release, and to keep the answer on file. Because the pledge ties the answer to the stated purpose, a user who expects to rely on it later should say so when asking.
Points for a foreign AI business
Treat answers as statements of fact
A Principle 2 answer is a factual statement about training data, made to someone who has said they are in or preparing for proceedings. Where a developer is also defending copyright claims elsewhere, the answer should be prepared with the same care as a litigation statement and kept consistent with the positions it takes in those cases.
Compare it with the EU training content summary
The European Commission's explanatory notice on Article 53(1)(d) of the EU AI Act describes a duty on providers of general-purpose AI models placed on the EU market, in so far as the models fall within the scope of the Act, to publish a sufficiently detailed summary of the content used for training, using a template from the AI Office. The duty applies from 2 August 2025, and models placed on the market before that date have until 2 August 2027. For content that the provider, or someone on its behalf, crawled or scraped from online sources, the template asks for the top 10% of internet domain names by size of content scraped. SMEs list the top 5% or 1,000 domains, whichever is lower. The notice also recommends that providers, voluntarily and in good faith, let rights holders find out on request whether content from specific domains not listed in the summary was scraped and used for training.
A provider preparing the EU summary will already have much of the material for Principle 1. The Commission has also pointed it towards answering domain-level requests. Japan builds more around that request:
- answering it is part of a comply-or-explain code, with a public list of businesses that have filed
- Principle 3 expressly extends it to users of the service, for content they have generated
- an explanation for not implementing Principle 2 must meet set conditions. A statement, with the necessary grounds, that disclosure is technically impossible counts. Saying only that a framework has not been built does not, unless the business also explains when it expects to complete it.
The AI news site ActuIA compared 24 published summaries from seven providers, including OpenAI, Mistral AI, ByteDance and DeepSeek, as of 25 September 2026. None of the 21 summaries that contain the domain section named a specific site; they described categories of sources instead, and the other three summaries omitted the section. ActuIA notes that its comparison is not a census. A rights holder reading such a summary cannot tell whether their own site was used, which is the question Principles 2 and 3 put to the business directly.
Objections from industry groups
In its comment of 26 January 2026 on the draft, the American Chamber of Commerce in Japan (ACCJ) urged the Intellectual Property Strategy Headquarters to "reconsider the entire framework". In the ACCJ's view, the comply-or-explain structure, combined with publishing the names of complying companies, risks enforcing substantive obligations without the legislative debate such measures normally require. On Principle 2, the comment said the standard of "preparing to take legal action" is imprecise and that the principle "bypasses established judicial procedures". The ACCJ also questioned applying the code to providers that merely make services available in Japan, pointing to the territorial nature of intellectual property rights and to the Berne Convention and TRIPS. Finally, it argued that asking whether a specific URL was used for training misreads how models are trained, because a trained model does not store its training data.
BSA | The Software Alliance, presenting to the government's study group on 21 April 2026, said the draft could create de facto binding obligations despite being framed as voluntary. It recommended deleting Principles 2 and 3 and leaving these questions to existing legal procedures and output-based analysis.
The final code kept Principles 2 and 3 and its reach to businesses outside Japan. It did add that it is not legally binding, narrowed its scope through new exclusions, and dropped the draft's statement that the government could set incentives in its programmes based on what businesses disclose. On the technical objection, Principles 2 and 3 ask whether a URL or other information is included in the data used for training and validation, not what the model's weights contain.
Absence from the list will be visible
Filing is voluntary, and the Cabinet Office will not review what businesses publish. It will, however, check that a filing business is actually operating and has published its acceptance at the stated address, and then add it to a public list with its name and website address. For Principle 1, the code asks stakeholders not to criticise a business immediately just because it has not disclosed some of the outline items. That passage is about partial disclosure, not about filing. Japanese enterprise customers and creative industry partners are still likely to ask about absence from the list, often in vendor due diligence.
A note from the US litigation side
This is outside my core practice, but having worked on US litigation alongside US counsel, I am interested in how this kind of "voluntary disclosure" in Japan will play into US proceedings. These are the questions I would put to US counsel:
- An answer to a requester in Japan is the business's own statement about its training data, made outside any court. Could it be used in a US class action, for example to help show which works were used?
- Would the pledge the requester gives in Japan limit how the answer is used in another country, or would that depend on the US court?
- Are the answer, and the work done to prepare it, protected by legal privilege, and does it matter who prepares it?
The choice in practice
- Accept all three principles. Once it files and the Office confirms the filing, the business appears on the list. It will need a request-handling process with legal review. Under Principle 1 it will also disclose how it handles its contact point for rights holders, including keeping records of responses, and the code encourages it to publish its policy for handling Principle 2 requests.
- Accept, but explain on Principle 2. This keeps the business on the list while it builds the process. The explanation needs grounds for technical impossibility, or a realistic completion date.
- Do not file. The code is not legally binding, so it attaches no sanction to this. But if a dispute reaches the courts in Japan, the same question can be put through procedures such as a party inquiry or a motion for a document production order, and the business will have less say over how it is framed.
Related reading
References
- Intellectual Property Strategy Headquarters, page on the formulation of and filing under the Principles-Code (Japanese).
- Principles-Code for the Protection of Intellectual Property and Transparency for the Appropriate Use of Generative AI (Japanese, official version).
- Principles-Code, provisional English translation.
- Filing form for acceptance of the Principles-Code (Japanese).
- Legal Subcommittee, Copyright Subdivision of the Council for Cultural Affairs, General Understanding on AI and Copyright (15 March 2024, Japanese).
- Agency for Cultural Affairs, General Understanding on AI and Copyright in Japan: Overview.
- Copyright Act, Article 30-4, and Code of Civil Procedure, Articles 163 and 221 (e-Gov, Japanese).
- Draft Principles-Code put out for public comment (26 December 2025, Japanese).
- American Chamber of Commerce in Japan, public comment on the draft Principles-Code (26 January 2026).
- BSA | The Software Alliance, recommendations on the draft Principles-Code, 11th meeting of the Study Group on Intellectual Property Rights in the AI Era (21 April 2026, Japanese).
- ActuIA, "EU AI Act summaries: 21 of 24 name no scraped sites in required field" (updated 7 October 2026).
- European Commission, Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models.
This article is for general information only and is not legal advice.
If you provide a generative AI model or service that reaches users in Japan, or use one in creative production, and want to think through how to respond to the code, you can reach us here.