📝 AI_Training_Data_and_Copyright.mdv4.5.1 · 2026-10-02

AI Training Data and Copyright: What Is Clear, What Is Not

Oct 3, 2026

Summary and scope

In Australia, copying copyright works into an AI training set without permission is likely infringement, because there is no US-style fair use and the government has ruled out a text-and-data-mining exception. Overseas courts disagree on whether training itself, or the trained model, counts as copying.

This is a desk review of public sources as at 3 October 2026. It is not legal advice and is not drawn from a case-law database (Westlaw, LexisNexis or similar). Anything relied on should be checked against the primary judgment or statute, and by a lawyer. I found no Australian court judgment on AI training.

Clearly inside copyright

Three acts are clearly within the copyright owner's control, at least in Australia.

Clearly outside copyright

These points come from long-standing copyright principles and my own knowledge, not from this review's searches, so check them before relying on them.

Contested

Courts have not agreed on whether training is copying, whether a trained model contains copies, and who bears the cost.

Question One answer Another answer
Is training on copyright works infringement? US trial courts in Bartz and Kadrey found generative training transformative and fair use The Third Circuit on 29 September 2026 rejected fair use for training a competing legal search tool, a non-generative tool built as a substitute for Westlaw
Does the trained model contain copies? Munich: yes, where works are memorised UK High Court in Getty v Stability: no, the weights represent learned patterns, not stored works
Does the licensing market count as harm? The Third Circuit acknowledged an emerging market for licensing training content Kadrey criticised the plaintiffs' market-harm arguments as half-hearted
Does training abroad escape local law? UK: Getty's training claim failed because training occurred outside the UK Munich: liability can attach where a service is offered in Germany

Still open in the US: the Second and Ninth Circuits have not ruled, and cross-motions for summary judgment in NYT v OpenAI were filed on 4 September 2026. The US Justice Department also filed a statement of interest arguing that LLM training is transformative. The UK Getty appeal has been allowed to proceed and may address training itself.

Australia in detail

Australia's defences are narrow, the policy line has been firm, and a leaked September 2026 proposal suggests it may be softening.

What the Act offers a developer

Policy timeline

Date Event
Late Oct 2025 Attorney-General Rowland rules out a text-and-data-mining exception after the Productivity Commission floated one, and refers next steps to the Copyright and AI Reference Group, including a possible small claims forum
15 Jul 2026 Prime Minister confirms mandatory national AI standards, an Office of AI and protection of copyright holders
Sep 2026 The ABC reports a leaked consultation paper, "AI on Australian Terms", with two options that would make training opt-out by default, paired with licensing or payment mechanisms

The government has stressed that no changes are finalised, and the Industry Minister has denied a reduction in copyright protection.

Likely test case. I found no Australian judgment on AI training. A test case would probably be brought by a rights holder or collecting society against a developer that copied works while training in Australia, for example on a local data centre. Australian authors and publishers' bodies, and local data-centre growth, make that a live prospect. Training done offshore raises the territoriality problem seen in the UK Getty case. That last step is my inference, not a reported plan.

The policy question

Whether training data should be proprietary is a policy argument, not a legal finding, and Australian law currently answers yes for unlicensed commercial copying. The two sides below are the strongest versions of each case, not my own view. I could not verify the quote attributed to Naomi Klein in the shared meme.

The case that training should be free (or opt-out)

The case that training should be licensed

The typewriter analogy and the "theft" framing both overreach. A typewriter keeps nothing of what its operator has read, whereas a trained model can memorise some works. But the law has no general rule that learning from a work needs consent: the question is whether the copying involved is permitted.