📝 Maximising_AI_Utility_Without_Breaching_Copyright.mdv4.5.1 · 2026-10-02

Maximising AI Utility Without Breaching Copyright

Date: 5 October 2026 Companion to: AI Training Data and Copyright: What Is Clear, What Is Not (3 Oct 2026)


0. Scope, status and caveats

This document answers a constructive question: how can AI deliver the most value while staying on the right side of copyright? The earlier review answered a different question (what is legal), so this one is a design brief, not a legal opinion.


1. The core insight: copyright protects expression, not knowledge

Almost every practical strategy below follows from one principle.

Protected Not protected
The specific way something is expressed: wording, structure, arrangement, a particular image, a recording Ideas, facts, methods, techniques, data, general style

So the question for any AI system is not "did it learn from copyrighted material?" but:

  1. Input side: did building the system involve making unauthorised copies of protected expression?
  2. Output side: can the system reproduce protected expression, or substantially so?
  3. Territory: where did the copying and the use happen, and whose law applies?

An AI design that handles those three questions well retains most of its usefulness. The sections below work through each.


2. A map of where the legal risk actually sits

Stage Activity Risk level (Australia) Why
Collection Scraping or copying works into a dataset High for protected works Reproduction is an exclusive right (s 31, Copyright Act 1968); no broad exception
Collection Using pirated copies Very high Courts treat this separately from training (see Bartz)
Storage Retaining a permanent library of copies High Same reproduction right; weakens any "transformative" argument overseas
Training The training run itself Unsettled Courts disagree on whether training is copying or whether a model contains copies
Model The weights Unsettled UK Getty: weights are learned patterns, not copies. Munich: memorised works can be "in" the model
Output Reproducing lyrics, passages, images on demand High GEMA v OpenAI turned heavily on how easily text could be prompted out
Output Original text in a general style Low Style is not protected
Retrieval Fetching and quoting a source at query time Moderate, manageable A quotation and linking problem, not a training problem

Takeaway: the highest, clearest risks are at collection and output. The training question is contested but is only one link in the chain. Designing carefully around the other links shrinks the exposure considerably.


3. Strategy 1: Build on material you are free to use

3.1 Public-domain works

3.2 Openly licensed material

3.3 Government and institutional material

3.4 Data you own or generate

Source Strength Weakness
First-party data (your own documents, support tickets, product telemetry) Clean ownership, highly relevant Privacy law applies; may be narrow
User-contributed data with clear consent Clear permission, fresh Needs proper terms, opt-in design
Commissioned data (paying people to write, label, annotate) Full control, targeted Expensive, scales slowly
Synthetic data (generated by models, simulators, rule systems) Cheap, scalable, avoids source copyright Quality ceiling and "model collapse" risk; may inherit issues from the generator's own training

Caution on synthetic data: if a synthetic dataset was produced by a model that itself was trained on unlicensed works, the legal position is unclear [my inference]. It reduces direct copying of identifiable works but does not erase upstream questions.

3.5 Facts, methods and ideas


4. Strategy 2: License what you cannot freely use

Licensing is the most direct route to the high-quality, recent, expressive material that makes AI most useful.

4.1 Models of licensing

Model How it works Pros Cons
Direct deals AI developer negotiates with a publisher, archive or platform Clean provenance, tailored terms, often includes better data access Expensive, favours big players, leaves out long-tail creators
Collective licensing A collecting society licenses a repertoire on behalf of many rights holders, as in music Lowers transaction cost, reaches many creators, predictable for developers Needs governance, fair distribution, rights-holder opt-in or statutory backing
Statutory or extended licensing Law creates a licence scheme, with remuneration set by statute or tribunal Certainty and scale Needs legislation; design disputes over rates and who is covered
Per-use or royalty-based Payment tied to usage, citations or outputs Aligns payment with value Hard to measure; technical attribution is immature
Equity or revenue share Rights holders share in the developer's revenue Aligns incentives Complex, may favour the largest rights holders

4.2 Why licensing is more viable than critics suggest

4.3 What licensing does not solve


5. Strategy 3: Separate learning from reproducing

This is where the technical design choices have the largest legal effect.

5.1 Why memorisation matters

5.2 Technical mitigations

Measure What it does Notes
Deduplication of training data Reduces the number of times any one work is seen Repeated exposure is generally associated with higher memorisation [from general knowledge]; deduplication is a standard, low-cost mitigation
Memorisation testing Probes the model with prefixes of known works to see if it completes them Needs a reference set of works to test against
Output filtering Detects and blocks or alters outputs that closely match known protected text Imperfect; can be evaded; useful as one layer, not the only one
Training-time techniques (for example, differential privacy approaches, regularisation) Limit how much any single example influences the model Can cost accuracy; effectiveness varies
Refusal behaviour The model declines to reproduce lyrics, full articles, full books Easy to apply to the most obvious categories; a policy decision
Rate limiting and abuse detection Makes systematic extraction attacks harder Practical defence against deliberate extraction

5.3 What this does and does not achieve


6. Strategy 4: Retrieve instead of absorb

Retrieval-augmented generation (RAG) changes the legal and economic structure of the problem.

6.1 How it works

Instead of baking source content into the model's weights, the system:

  1. Keeps a searchable index of sources (licensed, public-domain, or otherwise permitted).
  2. At query time, retrieves relevant passages.
  3. Generates an answer grounded in them, ideally with citations and links.

6.2 Benefits

Benefit Detail
Control Sources can be added, removed or updated; a rights holder can withdraw content without retraining a model
Attribution Answers can cite and link to the source, sending attention back to creators
Accuracy and currency Answers can reflect recent material that was not in the training set
Measurable use Because retrieval is logged, per-use payment becomes technically feasible
Smaller exposure The model itself can be trained on a narrower, cleaner corpus

6.3 Risks and limits

6.4 Design guidance


7. Strategy 5: Respect opt-outs and keep provenance

7.1 Opt-out signals

7.2 Provenance and record-keeping

Practice Purpose
Dataset documentation (sources, licences, dates, collection method) Lets you answer a rights holder or a regulator quickly and credibly
Per-source licence metadata Allows removal or re-licensing of specific content
Takedown and complaint process A fast, visible way to deal with specific claims
Audit trail for training runs Shows what data went into which model version

7.3 Why this matters beyond compliance


8. Strategy 6: Share the value with creators

The strongest argument on the licensing side of the debate is market substitution: models can compete with the people whose work trained them. Policy and business models that address this directly reduce both legal risk and political backlash.

8.1 Options

Mechanism Description Key question
Revenue share A proportion of AI revenue goes to a pool distributed to rights holders How is the pool split fairly?
Levy A charge on AI services or hardware, distributed through collecting societies (similar in concept to private-copying levies in some countries [from general knowledge]) Who pays, and who decides the rate?
Per-use royalty Payment when a work is retrieved, cited or demonstrably influences output Can influence be measured reliably? Current attribution science is immature
Licensing fees Payment for access to a corpus Does it reach individual creators or only large rights holders?
Small-claims or dispute forum A low-cost way for individual creators to pursue claims (floated in the Australian policy process, per the earlier review) Without it, rights exist on paper but are unenforceable for most individuals

8.2 The distribution problem

Whatever the mechanism, three issues recur:

  1. Who gets paid? Large publishers and platforms can negotiate; individual authors, artists and musicians often cannot.
  2. How is it divided? Measuring a single work's contribution to a model's capability is technically unsolved.
  3. Who is missing? Creators outside the licensing system, or whose work cannot be traced, may receive nothing.

8.3 Why developers might want this


9. Jurisdiction and territory

9.1 Australia

Feature Position (per the earlier review)
Fair use None. Australia uses fair dealing with five specific purposes
Text-and-data-mining exception Ruled out by the Attorney-General in late October 2025
Research or study Unlikely to cover copying entire works for machine learning; commercial purpose makes it harder
Temporary copies (ss 43A, 43B) May cover some transient copies during learning, but not the copies that build the dataset; one 2026 article argues s 43B(2) excludes copies of infringing copies
Policy direction Mandatory standards and protection of copyright holders confirmed (July 2026); leaked September 2026 paper floats opt-out options with licensing or payment; no changes are finalised
Case law No Australian judgment on AI training found

Practical consequence for an Australian-facing developer: assume unlicensed commercial copying of protected works into a training set is risky, and design around licensing, open material and retrieval.

9.2 United States

9.3 United Kingdom

9.4 Germany

9.5 The territoriality trap

Training abroad is not a clean escape:

Design guidance: think about three locations separately: where data is collected, where training occurs, and where the service is delivered.


10. What AI can still do very well

A common assumption is that copyright-conscious design means a crippled AI. For most high-value uses it does not.

Use case How it works within copyright
Reasoning, maths, logic, planning Depends on methods and ideas, not expression
Coding assistance Permissively licensed and first-party code; attribution notices; filters for verbatim reproduction
Explaining science, law, history, technology Facts and ideas are free; paraphrase rather than reproduce
Analysing a user's own documents The user supplies the material; the user's rights and permissions govern (check that they have the right to share it)
Translation, summarisation, rewriting of user-supplied text Operates on user-provided content
Enterprise knowledge tools First-party and licensed corpora
Scientific and technical tasks Large open-access literature, open data, public-domain material
Search and question-answering with citations Retrieval with attribution and short quotation
Writing in a general style Style is not protected; avoid reproducing specific works

10.1 Where it costs something

Use case Constraint
Reproducing or closely imitating specific protected works (lyrics, full chapters, paywalled articles) Legitimately restricted; should refuse or limit
Maximum breadth of cultural knowledge Harder without licences, because much recent expressive material is protected
"Write exactly like [named living author]" at close fidelity Style is free, but near-reproduction of specific works is not; also raises non-copyright issues such as passing off and moral rights [from general knowledge]
Frontier-scale general models Gain from breadth; licensing at that scale is expensive

Honest assessment: the constraint bites hardest on breadth of expressive cultural content and on the economics of frontier-scale training. It bites far less on reasoning, coding, science, analysis and enterprise use. [My inference, not a measured result.]


11. The policy argument, stated fairly

Whether training data should be proprietary is a policy question, not a legal finding. Both sides below are the strongest versions of each case, not a verdict.

11.1 The case that training should be free or opt-out

11.2 The case that training should be licensed

11.3 Where both sides have a point


12. A practical blueprint

12.1 For a developer

Layer Recommendation
Data Prioritise public-domain, openly licensed, first-party, commissioned and licensed data. Document everything
Pirated and paywalled content Exclude. Treat as the highest-risk category
Training Deduplicate; test for memorisation; train where you have assessed the legal position, but do not assume offshore is a full escape
Model behaviour Refuse reproduction of lyrics, full texts and the like; add output filtering as one layer
Knowledge freshness Use retrieval over licensed or open sources rather than absorbing recent expressive content
Attribution Cite, link, quote briefly
Opt-outs Honour robots.txt and any rights-reservation standard; run a takedown process
Creators Join or build a licensing or revenue-share scheme; make it easy to be paid
Territory Map collection, training and delivery jurisdictions separately
Legal Get advice from a lawyer in each relevant jurisdiction; monitor the Australian consultation outcome and the US appeals

12.2 For a user of AI tools

12.3 For a policymaker


13. What to watch next

Item Why it matters
Outcome of the Australian "AI on Australian Terms" consultation Could move Australia to opt-out plus licensing
Australian Copyright and AI Reference Group work, including a small-claims forum Determines whether individual creators can enforce rights
First Australian court case on AI training Would replace inference with law
Second and Ninth Circuit rulings; NYT v OpenAI Will shape US fair use for generative AI
UK Getty appeal May address training itself, not just territory
Spread of licensing markets and collective schemes Strengthens the case that unlicensed taking causes market harm
Technical progress on attribution and memorisation control Underpins both per-use payment and output-side safety

14. Summary

  1. Copyright protects expression, not ideas, facts or methods. Most of AI's reasoning, coding, scientific and analytical value does not depend on copying anyone's expression.
  2. The clearest risks are at collection and output, not in the abstract act of "learning". Pirated sources and verbatim reproduction are the danger zones.
  3. Free material goes a long way: public domain, open licences, first-party and commissioned data, and facts and methods.
  4. Licensing is the main route to high-quality expressive content, and collective schemes can make it workable beyond the largest players.
  5. Reduce memorisation through deduplication, testing, filtering and refusal behaviour. This supports the argument that a model learns rather than stores, though it does not authorise unlicensed collection in Australia.
  6. Retrieval with attribution keeps recent and expressive content outside the weights, keeps rights holders in control, and makes usage-based payment feasible.
  7. Keep provenance and honour opt-outs. This is both a compliance baseline and a prerequisite for fair compensation.
  8. Share value with creators. Market substitution is the strongest objection, and addressing it reduces both legal and political risk.
  9. Territory matters. Training offshore is not a clean escape from output or service-delivery liability.
  10. The cost is real but narrower than assumed: it falls mainly on breadth of cultural content and on frontier-scale economics, not on the core of what makes AI useful.

Appendix A: Items flagged for verification

Item Status
Bartz settlement size Stated in the earlier review as unverified
Telstra v Phone Directories reasoning From memory in the earlier review
Deduplication and memorisation relationship General technical knowledge, not verified here
Public-domain term in Australia General knowledge; check the specific work and category
Private-copying levy comparison General knowledge
Likely Australian test-case scenario Inference
"Costs bite hardest on breadth, least on reasoning/coding" Inference, not measured
All case outcomes and policy dates Taken from the 3 October 2026 review; verify against primary judgments and official sources

Appendix B: Source base

This document relies on the sources cited in AI Training Data and Copyright: What Is Clear, What Is Not (3 October 2026), which draws on law-firm commentary, news reports, and academic and practitioner articles. It was not compiled from a case-law database. Primary judgments and statutes should be consulted before any reliance, and a qualified lawyer should be engaged for any decision with legal consequences.