Maximising AI Utility Without Breaching Copyright
Date: 5 October 2026 Companion to: AI Training Data and Copyright: What Is Clear, What Is Not (3 Oct 2026)
0. Scope, status and caveats
This document answers a constructive question: how can AI deliver the most value while staying on the right side of copyright? The earlier review answered a different question (what is legal), so this one is a design brief, not a legal opinion.
- It builds on the 3 October 2026 desk review and on general principles of copyright law. It is not legal advice.
- Facts about the law and cases are taken from that review. Anything marked [from general knowledge] or [my inference] has not been verified in this document and should be checked against primary sources.
- Jurisdiction matters. The analysis leans on Australia (no fair use, no text-and-data-mining exception, narrow fair-dealing purposes), with notes on the US, UK and Germany.
- Technical effectiveness claims (for example, how well deduplication reduces memorisation) are described qualitatively. Treat any specific effectiveness figure you see elsewhere with suspicion unless it is sourced.
1. The core insight: copyright protects expression, not knowledge
Almost every practical strategy below follows from one principle.
| Protected |
Not protected |
| The specific way something is expressed: wording, structure, arrangement, a particular image, a recording |
Ideas, facts, methods, techniques, data, general style |
So the question for any AI system is not "did it learn from copyrighted material?" but:
- Input side: did building the system involve making unauthorised copies of protected expression?
- Output side: can the system reproduce protected expression, or substantially so?
- Territory: where did the copying and the use happen, and whose law applies?
An AI design that handles those three questions well retains most of its usefulness. The sections below work through each.
2. A map of where the legal risk actually sits
| Stage |
Activity |
Risk level (Australia) |
Why |
| Collection |
Scraping or copying works into a dataset |
High for protected works |
Reproduction is an exclusive right (s 31, Copyright Act 1968); no broad exception |
| Collection |
Using pirated copies |
Very high |
Courts treat this separately from training (see Bartz) |
| Storage |
Retaining a permanent library of copies |
High |
Same reproduction right; weakens any "transformative" argument overseas |
| Training |
The training run itself |
Unsettled |
Courts disagree on whether training is copying or whether a model contains copies |
| Model |
The weights |
Unsettled |
UK Getty: weights are learned patterns, not copies. Munich: memorised works can be "in" the model |
| Output |
Reproducing lyrics, passages, images on demand |
High |
GEMA v OpenAI turned heavily on how easily text could be prompted out |
| Output |
Original text in a general style |
Low |
Style is not protected |
| Retrieval |
Fetching and quoting a source at query time |
Moderate, manageable |
A quotation and linking problem, not a training problem |
Takeaway: the highest, clearest risks are at collection and output. The training question is contested but is only one link in the chain. Designing carefully around the other links shrinks the exposure considerably.
3. Strategy 1: Build on material you are free to use
3.1 Public-domain works
- Works whose copyright has expired can be copied freely [from general knowledge]. In Australia the term for most works is life of the author plus 70 years, with different rules for some categories; check the specific work.
- Large digitised collections from libraries and archives are a rich, legitimate source of historical text, images and recordings.
- Limitation: public-domain corpora skew old. They under-represent modern language, current science, contemporary culture and current code. A model trained only on them would sound dated and know nothing recent.
3.2 Openly licensed material
- Creative Commons, open-source software licences, and similar are usable on the licence's terms.
- Watch the conditions. Attribution (CC-BY), share-alike, and non-commercial terms impose real obligations. "Open" does not mean "unconditional". Non-commercial licences generally rule out commercial training.
- Permissively licensed code (MIT, BSD, Apache-2.0) is the best-understood category, though attribution notices still need to be preserved.
- Strong-copyleft code (for example, GPL-family) raises unresolved questions about whether outputs or models inherit obligations [from general knowledge; contested].
3.3 Government and institutional material
- Many governments and public institutions release data and publications for reuse. Check the specific terms. Crown copyright and equivalents differ by country.
3.4 Data you own or generate
| Source |
Strength |
Weakness |
| First-party data (your own documents, support tickets, product telemetry) |
Clean ownership, highly relevant |
Privacy law applies; may be narrow |
| User-contributed data with clear consent |
Clear permission, fresh |
Needs proper terms, opt-in design |
| Commissioned data (paying people to write, label, annotate) |
Full control, targeted |
Expensive, scales slowly |
| Synthetic data (generated by models, simulators, rule systems) |
Cheap, scalable, avoids source copyright |
Quality ceiling and "model collapse" risk; may inherit issues from the generator's own training |
Caution on synthetic data: if a synthetic dataset was produced by a model that itself was trained on unlicensed works, the legal position is unclear [my inference]. It reduces direct copying of identifiable works but does not erase upstream questions.
3.5 Facts, methods and ideas
- Raw facts and underlying ideas are unowned. The Telstra v Phone Directories line of Australian cases (from the earlier review, itself from memory) is a reminder that the compilation and authorship test matters, and that mere effort does not create copyright.
- A system that extracts the facts and methods from a body of text, rather than copying the text, has far lower exposure. This is the principle behind knowledge graphs, structured databases, and fact extraction pipelines.
- Practical note: extracting facts from a work still typically requires reading the work. If that involves copying it into a system, the reproduction question arises at that step even if the stored output is only facts. Whether transient processing is covered depends on the jurisdiction's exceptions (see Section 9).
4. Strategy 2: License what you cannot freely use
Licensing is the most direct route to the high-quality, recent, expressive material that makes AI most useful.
4.1 Models of licensing
| Model |
How it works |
Pros |
Cons |
| Direct deals |
AI developer negotiates with a publisher, archive or platform |
Clean provenance, tailored terms, often includes better data access |
Expensive, favours big players, leaves out long-tail creators |
| Collective licensing |
A collecting society licenses a repertoire on behalf of many rights holders, as in music |
Lowers transaction cost, reaches many creators, predictable for developers |
Needs governance, fair distribution, rights-holder opt-in or statutory backing |
| Statutory or extended licensing |
Law creates a licence scheme, with remuneration set by statute or tribunal |
Certainty and scale |
Needs legislation; design disputes over rates and who is covered |
| Per-use or royalty-based |
Payment tied to usage, citations or outputs |
Aligns payment with value |
Hard to measure; technical attribution is immature |
| Equity or revenue share |
Rights holders share in the developer's revenue |
Aligns incentives |
Complex, may favour the largest rights holders |
4.2 Why licensing is more viable than critics suggest
- The usual objection is that "licensing every source is impractical". That is true for individual negotiations, but collective schemes exist precisely to solve this, and music and broadcast licensing have done so for decades [from general knowledge].
- The Third Circuit's recognition of an emerging market for licensing training content (from the earlier review) matters: the more a licensing market exists, the harder it becomes to argue that unlicensed copying causes no market harm. Licensing both reduces risk and, in the US, strengthens the position of rights holders.
- Licensing often improves the product, because the developer gets clean, well-structured, well-labelled, current data rather than scraped noise.
4.3 What licensing does not solve
- Licences do not cover everything. Large parts of the open web have no clear, reachable owner.
- Terms can restrict uses (training allowed, but outputs limited; one model but not derivatives).
- Cost concentrates power: large developers can afford licences, small developers and open-source projects may not. This is a real policy problem, not just a business one.
5. Strategy 3: Separate learning from reproducing
This is where the technical design choices have the largest legal effect.
5.1 Why memorisation matters
- In GEMA v OpenAI, the Munich court's concern was not just that lyrics were used in training, but how easily they could be prompted back out (from the earlier review).
- A model that can output a work verbatim, or near-verbatim, looks like a copy. A model that has absorbed patterns but cannot reproduce specific works looks like a learner.
- The UK Getty v Stability reasoning (weights represent learned patterns, not stored works) is easier to sustain for a model that does not memorise.
5.2 Technical mitigations
| Measure |
What it does |
Notes |
| Deduplication of training data |
Reduces the number of times any one work is seen |
Repeated exposure is generally associated with higher memorisation [from general knowledge]; deduplication is a standard, low-cost mitigation |
| Memorisation testing |
Probes the model with prefixes of known works to see if it completes them |
Needs a reference set of works to test against |
| Output filtering |
Detects and blocks or alters outputs that closely match known protected text |
Imperfect; can be evaded; useful as one layer, not the only one |
| Training-time techniques (for example, differential privacy approaches, regularisation) |
Limit how much any single example influences the model |
Can cost accuracy; effectiveness varies |
| Refusal behaviour |
The model declines to reproduce lyrics, full articles, full books |
Easy to apply to the most obvious categories; a policy decision |
| Rate limiting and abuse detection |
Makes systematic extraction attacks harder |
Practical defence against deliberate extraction |
5.3 What this does and does not achieve
- Does: reduces output-side infringement risk, supports the argument that the model does not "contain" works, and aligns with how courts have reasoned so far.
- Does not: make unlicensed collection lawful. Reducing memorisation fixes the output problem; it does not retroactively authorise the copying at the collection stage in a jurisdiction like Australia.
- Honest limit: nobody can currently guarantee zero memorisation in a large model. The aim is to bring it down to rare, hard-to-trigger cases and to have credible processes for finding and fixing them.
6. Strategy 4: Retrieve instead of absorb
Retrieval-augmented generation (RAG) changes the legal and economic structure of the problem.
6.1 How it works
Instead of baking source content into the model's weights, the system:
- Keeps a searchable index of sources (licensed, public-domain, or otherwise permitted).
- At query time, retrieves relevant passages.
- Generates an answer grounded in them, ideally with citations and links.
6.2 Benefits
| Benefit |
Detail |
| Control |
Sources can be added, removed or updated; a rights holder can withdraw content without retraining a model |
| Attribution |
Answers can cite and link to the source, sending attention back to creators |
| Accuracy and currency |
Answers can reflect recent material that was not in the training set |
| Measurable use |
Because retrieval is logged, per-use payment becomes technically feasible |
| Smaller exposure |
The model itself can be trained on a narrower, cleaner corpus |
6.3 Risks and limits
- The index is still a copy. If a retrieval system stores copies of web pages or documents, the collection and storage problems return, even though the weights are clean. Search engines have long dealt with this; the position of an AI system that stores and then substantially reproduces content is less settled.
- Substantial reproduction at output. Long verbatim quotation of retrieved passages is still reproduction. Short, attributed quotation is safer, but the boundaries differ by jurisdiction and Australia's fair-dealing purposes are narrow.
- Substitution. If the answer is so complete that users never visit the source, the economic harm argument returns, regardless of how the system is built. This is the "market replacement" concern that rights holders raise.
- Paywalls and terms of use. Retrieving content that sits behind a paywall or restrictive terms raises contract and access questions in addition to copyright [from general knowledge].
6.4 Design guidance
- Prefer short quotations with attribution and a link over long extracts.
- Prefer summarising in the system's own words over reproducing structure and phrasing.
- Prefer licensed or open sources for the index, and treat paywalled content as off-limits unless licensed.
- Log retrieval so that usage-based compensation is possible.
7. Strategy 5: Respect opt-outs and keep provenance
7.1 Opt-out signals
- Honour robots.txt and emerging machine-readable rights reservations.
- Where a rights-holder registry or standard emerges, build to it.
- This matters more as policy moves. The leaked September 2026 Australian consultation paper described options that would make training opt-out by default (from the earlier review; not finalised). A developer already honouring opt-outs is well placed if that is adopted, and is better placed to argue good faith if it is not.
7.2 Provenance and record-keeping
| Practice |
Purpose |
| Dataset documentation (sources, licences, dates, collection method) |
Lets you answer a rights holder or a regulator quickly and credibly |
| Per-source licence metadata |
Allows removal or re-licensing of specific content |
| Takedown and complaint process |
A fast, visible way to deal with specific claims |
| Audit trail for training runs |
Shows what data went into which model version |
7.3 Why this matters beyond compliance
- Provenance makes licensing and compensation workable: you cannot pay creators if you do not know whose work you used.
- Clean provenance improves data quality and makes debugging model behaviour easier.
- It reduces the risk of inadvertently including pirated or stolen material, which the Bartz distinction shows is treated particularly seriously (from the earlier review).
8. Strategy 6: Share the value with creators
The strongest argument on the licensing side of the debate is market substitution: models can compete with the people whose work trained them. Policy and business models that address this directly reduce both legal risk and political backlash.
8.1 Options
| Mechanism |
Description |
Key question |
| Revenue share |
A proportion of AI revenue goes to a pool distributed to rights holders |
How is the pool split fairly? |
| Levy |
A charge on AI services or hardware, distributed through collecting societies (similar in concept to private-copying levies in some countries [from general knowledge]) |
Who pays, and who decides the rate? |
| Per-use royalty |
Payment when a work is retrieved, cited or demonstrably influences output |
Can influence be measured reliably? Current attribution science is immature |
| Licensing fees |
Payment for access to a corpus |
Does it reach individual creators or only large rights holders? |
| Small-claims or dispute forum |
A low-cost way for individual creators to pursue claims (floated in the Australian policy process, per the earlier review) |
Without it, rights exist on paper but are unenforceable for most individuals |
8.2 The distribution problem
Whatever the mechanism, three issues recur:
- Who gets paid? Large publishers and platforms can negotiate; individual authors, artists and musicians often cannot.
- How is it divided? Measuring a single work's contribution to a model's capability is technically unsolved.
- Who is missing? Creators outside the licensing system, or whose work cannot be traced, may receive nothing.
8.3 Why developers might want this
- It converts an open-ended legal risk into a predictable cost.
- It supports the argument that unlicensed copying causes no market harm because a functioning licensing market exists, which cuts both ways: a licensing market makes unlicensed taking harder to defend in the US fair-use analysis.
9. Jurisdiction and territory
9.1 Australia
| Feature |
Position (per the earlier review) |
| Fair use |
None. Australia uses fair dealing with five specific purposes |
| Text-and-data-mining exception |
Ruled out by the Attorney-General in late October 2025 |
| Research or study |
Unlikely to cover copying entire works for machine learning; commercial purpose makes it harder |
| Temporary copies (ss 43A, 43B) |
May cover some transient copies during learning, but not the copies that build the dataset; one 2026 article argues s 43B(2) excludes copies of infringing copies |
| Policy direction |
Mandatory standards and protection of copyright holders confirmed (July 2026); leaked September 2026 paper floats opt-out options with licensing or payment; no changes are finalised |
| Case law |
No Australian judgment on AI training found |
Practical consequence for an Australian-facing developer: assume unlicensed commercial copying of protected works into a training set is risky, and design around licensing, open material and retrieval.
9.2 United States
- US trial courts in Bartz and Kadrey found generative training transformative and fair use, with the pirated-library storage in Bartz treated separately.
- The Third Circuit on 29 September 2026 rejected fair use for training a competing, non-generative legal search tool, and acknowledged an emerging licensing market.
- Other circuits (Second, Ninth) have not ruled; NYT v OpenAI summary judgment motions were filed on 4 September 2026.
- Takeaway: the US is not a safe harbour. Fair use is fact-specific, and competing substitute uses fare worse.
9.3 United Kingdom
- Getty v Stability (High Court): no secondary infringement because the weights represent learned patterns; the training claim failed because training occurred outside the UK. An appeal has been allowed to proceed and may address training itself.
- Takeaway: territoriality was decisive. Where training happens matters.
9.4 Germany
- GEMA v OpenAI (Munich): memorised lyrics in a model and in outputs infringed; liability can attach where a service is offered in Germany.
- Takeaway: offering a service into a jurisdiction can create exposure even if training happened elsewhere.
9.5 The territoriality trap
Training abroad is not a clean escape:
- It may defeat a training-copy claim in the country where the claim is brought (the UK Getty result).
- It may not defeat an output or service-offered-here claim (the Munich approach).
- Australian policy discussion of local data-centre growth means onshore training is a live test-case scenario (from the earlier review; the likely test case is my inference, not a reported plan).
Design guidance: think about three locations separately: where data is collected, where training occurs, and where the service is delivered.
10. What AI can still do very well
A common assumption is that copyright-conscious design means a crippled AI. For most high-value uses it does not.
| Use case |
How it works within copyright |
| Reasoning, maths, logic, planning |
Depends on methods and ideas, not expression |
| Coding assistance |
Permissively licensed and first-party code; attribution notices; filters for verbatim reproduction |
| Explaining science, law, history, technology |
Facts and ideas are free; paraphrase rather than reproduce |
| Analysing a user's own documents |
The user supplies the material; the user's rights and permissions govern (check that they have the right to share it) |
| Translation, summarisation, rewriting of user-supplied text |
Operates on user-provided content |
| Enterprise knowledge tools |
First-party and licensed corpora |
| Scientific and technical tasks |
Large open-access literature, open data, public-domain material |
| Search and question-answering with citations |
Retrieval with attribution and short quotation |
| Writing in a general style |
Style is not protected; avoid reproducing specific works |
10.1 Where it costs something
| Use case |
Constraint |
| Reproducing or closely imitating specific protected works (lyrics, full chapters, paywalled articles) |
Legitimately restricted; should refuse or limit |
| Maximum breadth of cultural knowledge |
Harder without licences, because much recent expressive material is protected |
| "Write exactly like [named living author]" at close fidelity |
Style is free, but near-reproduction of specific works is not; also raises non-copyright issues such as passing off and moral rights [from general knowledge] |
| Frontier-scale general models |
Gain from breadth; licensing at that scale is expensive |
Honest assessment: the constraint bites hardest on breadth of expressive cultural content and on the economics of frontier-scale training. It bites far less on reasoning, coding, science, analysis and enterprise use. [My inference, not a measured result.]
11. The policy argument, stated fairly
Whether training data should be proprietary is a policy question, not a legal finding. Both sides below are the strongest versions of each case, not a verdict.
11.1 The case that training should be free or opt-out
- Reading and learning from public material is how people create, and a model is a form of learning.
- Facts, ideas and style are already unowned, and output is mostly new expression.
- Licensing every source is impractical, and strict rules may push training and investment offshore. The leaked Australian proposal is reportedly aimed at attracting AI investment.
- Broad access supports competition: a licence-only world favours incumbents who can afford it.
11.2 The case that training should be licensed
- Training makes commercial copies, and models can substitute for the creators whose work they use.
- An opt-out system forces creators to act to protect rights they already hold. Senator Pocock said the burden should not fall on Australians to defend them.
- An emerging licensing market, recognised by the Third Circuit, means unlicensed training may undercut a real market.
- Creators' consent and compensation are fairness questions, not only economic ones.
11.3 Where both sides have a point
- A typewriter keeps nothing of what its operator has read, whereas a trained model can memorise some works. The analogy overreaches, and so does the word "theft" (from the earlier review).
- But the law has no general rule that learning from a work needs consent. The question is whether the copying involved is permitted.
- The most defensible middle path combines free use of unprotected material, licensing for protected expression, strong anti-memorisation controls, and a fair route for creators to be paid.
12. A practical blueprint
12.1 For a developer
| Layer |
Recommendation |
| Data |
Prioritise public-domain, openly licensed, first-party, commissioned and licensed data. Document everything |
| Pirated and paywalled content |
Exclude. Treat as the highest-risk category |
| Training |
Deduplicate; test for memorisation; train where you have assessed the legal position, but do not assume offshore is a full escape |
| Model behaviour |
Refuse reproduction of lyrics, full texts and the like; add output filtering as one layer |
| Knowledge freshness |
Use retrieval over licensed or open sources rather than absorbing recent expressive content |
| Attribution |
Cite, link, quote briefly |
| Opt-outs |
Honour robots.txt and any rights-reservation standard; run a takedown process |
| Creators |
Join or build a licensing or revenue-share scheme; make it easy to be paid |
| Territory |
Map collection, training and delivery jurisdictions separately |
| Legal |
Get advice from a lawyer in each relevant jurisdiction; monitor the Australian consultation outcome and the US appeals |
12.2 For a user of AI tools
- Treat outputs as potentially derivative if you ask for a specific work or author's text.
- Do not feed in material you do not have the right to share.
- Where you publish AI-assisted work commercially, check the tool's terms and your own jurisdiction's rules on authorship. In Australia, content lacking human authorship may not attract copyright [from general knowledge; see earlier review's Telstra reference].
12.3 For a policymaker
- Clarify the position quickly: uncertainty chills both investment and creator confidence.
- Consider a statutory or collective licensing scheme with a workable distribution mechanism, so licensing is feasible for small developers as well as large ones.
- Pair any opt-out model with strong, simple, machine-readable opt-out standards, and with real remuneration, or the burden falls on the creators with the least resources.
- Provide a low-cost dispute route for individual creators.
- Address territoriality, so that offshore training for onshore service does not make local rights illusory.
13. What to watch next
| Item |
Why it matters |
| Outcome of the Australian "AI on Australian Terms" consultation |
Could move Australia to opt-out plus licensing |
| Australian Copyright and AI Reference Group work, including a small-claims forum |
Determines whether individual creators can enforce rights |
| First Australian court case on AI training |
Would replace inference with law |
| Second and Ninth Circuit rulings; NYT v OpenAI |
Will shape US fair use for generative AI |
| UK Getty appeal |
May address training itself, not just territory |
| Spread of licensing markets and collective schemes |
Strengthens the case that unlicensed taking causes market harm |
| Technical progress on attribution and memorisation control |
Underpins both per-use payment and output-side safety |
14. Summary
- Copyright protects expression, not ideas, facts or methods. Most of AI's reasoning, coding, scientific and analytical value does not depend on copying anyone's expression.
- The clearest risks are at collection and output, not in the abstract act of "learning". Pirated sources and verbatim reproduction are the danger zones.
- Free material goes a long way: public domain, open licences, first-party and commissioned data, and facts and methods.
- Licensing is the main route to high-quality expressive content, and collective schemes can make it workable beyond the largest players.
- Reduce memorisation through deduplication, testing, filtering and refusal behaviour. This supports the argument that a model learns rather than stores, though it does not authorise unlicensed collection in Australia.
- Retrieval with attribution keeps recent and expressive content outside the weights, keeps rights holders in control, and makes usage-based payment feasible.
- Keep provenance and honour opt-outs. This is both a compliance baseline and a prerequisite for fair compensation.
- Share value with creators. Market substitution is the strongest objection, and addressing it reduces both legal and political risk.
- Territory matters. Training offshore is not a clean escape from output or service-delivery liability.
- The cost is real but narrower than assumed: it falls mainly on breadth of cultural content and on frontier-scale economics, not on the core of what makes AI useful.
Appendix A: Items flagged for verification
| Item |
Status |
| Bartz settlement size |
Stated in the earlier review as unverified |
| Telstra v Phone Directories reasoning |
From memory in the earlier review |
| Deduplication and memorisation relationship |
General technical knowledge, not verified here |
| Public-domain term in Australia |
General knowledge; check the specific work and category |
| Private-copying levy comparison |
General knowledge |
| Likely Australian test-case scenario |
Inference |
| "Costs bite hardest on breadth, least on reasoning/coding" |
Inference, not measured |
| All case outcomes and policy dates |
Taken from the 3 October 2026 review; verify against primary judgments and official sources |
Appendix B: Source base
This document relies on the sources cited in AI Training Data and Copyright: What Is Clear, What Is Not (3 October 2026), which draws on law-firm commentary, news reports, and academic and practitioner articles. It was not compiled from a case-law database. Primary judgments and statutes should be consulted before any reliance, and a qualified lawyer should be engaged for any decision with legal consequences.