Consultation on Copyright in the Age of Generative Artificial Intelligence: Submissions O-T

The information on this Web page has been provided by external sources. The Government of Canada is not responsible for the accuracy, reliability or currency of the information supplied by external sources. Users wishing to rely upon this information should consult directly with the source of the information. Content provided by external sources is not subject to official languages and privacy requirements.

O

OCAD University

Technical Evidence

Background: 

At OCAD University, Canada’s oldest and largest art and design university, a key strategic goal is to increase access to emerging technologies, and enable students to use, create and innovate with these technologies skillfully and responsibly. As such, Artificial Intelligence (AI) and Generative Artificial Intelligence (GAI) are topics the institution is grappling with through a faculty advisory, a working group with external experts organized by the University’s Cultural Policy Hub, and input from students, faculty, and staff. 

AI and GAI represent an incredible opportunity for artists. Many OCAD U students, faculty, and staff members from across the university’s community use AI, GAI, and multimodal models (such as ChatGPT, Midjourney, Dall-E, Stable Diffusion, Bard, GrammarlyGo, among others) as part of their work. For example, the university’s administrators may use ChatGPT for general purpose copy drafting, while faculty and students experiment with image and music GAI tools and explore how they can be used to enhance their creativity. 

But these tools also have the potential to disrupt the standards and regulations that are in place to protect artists and the work they create. The university also has researchers who have formed a group exploring the ethical dimensions of the recent and rapid rise of GAI; they are working to establish guidelines and recommendations on how their peers can use GAI responsibly, both in pedagogical and creative contexts. 

Overall, the institution has recognized that GAI technologies have significant and immediate implications for conceptions of creativity, authorship, as well as approaches to pedagogy, and has responded with a focus on developing critical AI and GAI literacy by:

Defining the technology and its applications for students and faculty; 

Establishing the affordances and limitations of this technology; 

Situating the work around AI and GAI within the university’s broader anti-racist and decolonial framework and acknowledging how these technologies provide potential opportunities and drawbacks for marginalized students, faculty, and staff; and

Striving for equitable access to GAI tools, particularly for Indigenous, Black, people of color, neurodivergent, and marginalized faculty, staff and students.

Some of the university’s students and faculty are wary of GAI’s potential and how it could impact their creative practice. Members from a few departments have serious concerns about the potential of GAI tools to replace entirely the need for human intervention or participation in certain creative processes, and the accompanying threat to their livelihood. As a result of these different viewpoints, the university’s approach seeks to balance the following priorities: 

Promote innovation and experimentation in the development of new and powerful creative tools;

Train students and faculty to understand and implement these tools so they can incorporate them into their practice and remain proficient in using them; 

Anticipate industry and sector trends in adopting these tools to ensure that students are future-proofed and ready for their transition into the professional realm;

Establish guidelines and guardrails and provide recommendations to government that will protect students and faculty’s creative work from exploitation; and 

Advocate when necessary for artists and creatives retaining control over how their work is used now and in the future.

This background informs the responses to the questions posed by ISED in the consultation. 

The university’s approach to AI and GAI is iterative and in progress. As such, the recommendations posed should be considered in this context. This consultation raises more questions than answers, and our primary recommendation to the government is that more time is needed for discussion and public consultation on the issues raised, as well as those not raised.

OCAD U encourages the government to reopen the consultation process with an additional lens to the following issues, which this most recent survey does not address:

Bias in datasets used to train AI and the perceived threat of increased bias if AI developers are restricted to using materials from the public domain for TDM and machine-learning

The recognition of Indigenous sovereignty in the development, training, and application of AI and changes to copyright and intellectual property law

The existing and future impacts on human rights in AI development and implementation

The implications for education around ethical approaches to pedagogy in the age of GAI

Text and Data Mining

Recommendations: 

Develop regulations to ensure transparency around datasets. With these regulations should come oversight and compliance.

Put safeguards in place that track, record, and disclose how an artist's work is being used in text and data mining (TDM) activities.

Put the onus of developing those mechanisms and obtaining permissions or licensing for the use of content in TDM activities on AI tool and model developers, not content creators.

Datasets are largely comprised of customer data that we originally understood to be used for internal purposes but were then used for social media and targeted advertising and are now being used by AI to create products that are being sold back to us. The components of the datasets were technically acquired legally, but complied to laws that didn’t envision this use case. Any laws regarding AI need to be generated in such a way that they reflect this new use of information and data previously, currently, and prospectively–these laws should benefit and protect users rather than the creators of the tools.

The issue of dataset transparency was recently addressed by the European Parliament proposed regulatory framework on AI. Their recommendations specifically addressed the issue of dataset transparency by establishing the following requirements:

Disclosing that the content was generated by AI

Designing the model to prevent it from generating illegal content

Publishing summaries of copyrighted data used for training

The European Parliament also noted that “high-impact general purpose AI models” might pose systemic risk and should be subjected to thorough evaluations. They continued that citizens would have the right to report “serious incidents” involving this technology to the European Commission. More information is required on how the EU intends to ensure compliance with these requirements and be proactive in ensuring that citizens’ and creatives’ rights are not being infringed.

Many artists and creatives, especially emerging ones, are the direct copyright owners of their work, and do not have the support of agents, producers, and/or legal experts to ensure that their work is not being exploited or used in ways that circumvent or contravene the copyright act. Even if their work is found to be used inappropriately or illegally, the authors of those works in most cases do not have the capacity or resources to expend enforcing how their work is used and getting the compensation they are owed for its use.

Given these material realities, it is critical that the Canadian government, too, put safeguards in place that automatically protect an artist's work from being used in text and data mining activities. There need to be mechanisms for oversight and compliance in place that track, record, and disclose how content is being used in TDM activities. Finally, the Canadian government should put the onus of obtaining permissions or licensing for such work on those compiling content to develop and use training datasets, not on the content creators themselves.

OCAD U’s initial recommendations on best practices and potential ways forward are as follows:

Develop regulations to ensure transparency around datasets. Oversight and compliance must be considered as part of these regulations.

As a rule, artists should have the option to opt their work into being used openly in TDM activities. They should not be expected to opt their work out of being used in TDM activities.

Licensing and subsequent compensation will likely need to be achieved through collective models and could be serviced in part through existing collective management organizations (or Collective Societies) such as SOCAN, CAARC, SODRAC, ACTRA, and others. 

It would be reasonable for artists to opt into or register with a Collective Society to be eligible for remuneration from TDM activities.

If effective regulation and transparency around datasets is not in place, mechanisms must be developed to track how an artist’s work is being used in TDM activities. Search engines like “Have I Been Trained?” already exist to help artists determine if their work has been used in TDM activities. This type of program would need to be expanded with cooperation from those developing training datasets for TDM and machine-learning activities. Still, this remains a defensive response that places the onus on creators and does not represent a meaningful or sustainable solution to the larger issue of a lack of transparency in datasets.

As for the levels of remuneration for the use of a given work in TDM activities, the government could adopt a recommended fee structure to be developed by a consortium made up of: experts on artistic remuneration (e.g. CARFAC, RAAV, IMAA, etc.); representatives from the tech industry and others developing GAI tools and engaging in TDM activities; legal experts on copyright and remuneration; and representatives from arts service organizations across disciplines. 

There are a number of obstacles that arise when considering how the above recommendation could be implemented:

How to develop a or multiple mechanisms to track, record, and disclose content being used in TDM activities? (Examples already exist, such as Hugging Face’s BigCode project: https://huggingface.co/datasets/bigcode/governance-card#2-data-and-model-governance).

How to regulate this process of tracking and recording?

How to establish realistic licensing regulations and parameters to not stifle innovation?

How to educate artists and creatives about their rights as well as any changes to existing copyright legislation?

Ultimately, the goal should be to ensure fair remuneration for artists that will reflect how their work is being used in TDM activities. The government could also consider whether artists and creatives should be retroactively compensated for the use of their work in TDM and ML activities that have led to the creation of significant and highly popular GAI systems like ChatGPT, DALL-E, Midjourney, etc.

Authorship and Ownership of Works Generated by AI

Recommendations:

The definition of an author or creator may need to be redefined to account for how GAI is involved in the creative process.

The revision of these definitions should be led by authors and creators working alongside policy makers, and not by those developing GAI technologies.

Continue consultation on this topic, alongside the Department of Canadian Heritage and creatives.

The uncertainty around authorship of GAI- and AI-assisted creative works could have significant impacts for the creative community. Current definitions of authorship as being attributable to a “natural person” fail to account for instances where GAI is involved in the input or output of creative materials. The existing definitions in copyright law may not be sufficient for protecting the creative work of artists and designers and also enabling them to fully leverage new tools.

Our existing understandings of authorship may now be resting on premises and assumptions that are insufficient. There needs to be greater consideration of where the work of creativity is located in existing processes that are changing with the rise of new technologies, as well as in new modes of creation that are becoming possible as those technologies become more accessible. The experts on these questions of how authorship and creativity has changed are authors and creators themselves; if policymakers intend to revise how they define these terms, creatives need to be directly involved in developing these new definitions and understandings.

This questions around authorship and ownership requires more consultation and discussion, which ISED should lead alongside the Department of Canadian Heritage and must involve creatives.

Infringement and Liability regarding AI

Recommendation:

The government needs to be proactive in putting regulations in place rather than encouraging innovation and then leaving it up to the courts to make legal determinations of liability in instances of infringement (or suspected infringement).

OCAD U has begun to develop frameworks to help students and faculty develop an awareness of how GAI applications might currently store or use their ideas and/or intellectual property without their explicit consent. Its current guidelines remind students and faculty to always attribute when and how they are using GAI tools in their work, alongside best practices for doing so. The guidelines also put the onus on the person using a GAI tool to “ensure [they] are familiar with the privacy policy of each application before [they] use it [and] never enter someone else’s work into a GAI app without their permission and consent” and encourage faculty members to remind students of this in their course outlines and assignment instructions.

While these measures may help students and faculty protect themselves in the short term, there is much more clarity required from government on how the content used to train these AI tools should be tracked and regulated, how the content creators whose work is used should be compensated, how authorship is determined when these tools are involved, and where liability lies if copyright is infringed on by GAI tools.

Consultation with many experts suggests that businesses that build commercial AI applications can be expected to advocate for as little oversight as possible in tracking and recording what content they use to train their models. They are expected by and large to advocate for the deregulation of content and for free and open access to as much content as possible. This model would not only enable them to potentially build more powerful tools; it would allow them to do so at a significantly lower cost.

Unfortunately, the ensuing benefits to innovation would come at the cost of opportunities for artists and creative professionals to be compensated for the use of their work. As an institution that prepares students for careers in the creative sector, OCAD U’s recommendations need to situate artists and creatives—and their opportunities to be successful in their work and practices—first.

The main barrier that the university anticipates in determining whether an AI system has accessed or copied copyright-protected content when generating an infringing output is the lack of implementation of sector-wide tools and standards to track, record, and disclose what content is being used in TDM and machine-learning activities. Furthermore, research on code LLMs has suggested that the lack of transparency can also hinder innovation by allowing only a small number of well-funded labs to participate in shaping the technology (see the Hugging Face report linked above). Implementing the earlier recommendation on tracking and disclosing this information would provide a record of what types of data were included in a machine’s training set and what it could be borrowing or stealing from in generating new work that could be infringing on copyright-protected material.

There must be greater clarity on where liability lies when AI-generated works infringe on copyright, and there needs to be greater efforts to educate artists and users of GAI tools on how to ensure that their practice is responsible and in compliance with any legislations or regulations that are adopted on this issue. Infringement as it relates to AI-generated works is both an opportunity for protection and a liability for artists, businesses, and developers alike. 

If the government relies on the courts to eventually set precedents and make determinations on the potential harms and unlawful outcomes of this technology’s proliferation, the harm will be most acutely experienced by marginalized individuals and communities. As such, the government needs to be proactive in putting regulations in place, regulations which should address or clarify:

Definitions around authorship when GAI tools are involved.

Protections from liability for coders, programmers, and other workers involved in developing AI tools and models on behalf of private companies.

The establishment of updated corporate liability models for potential infringing activities.

Processes for registering AI tools and models, as well as standards and guidelines for disclosing human involvement in AI-generated works.

Comments and Suggestions

OCAD University is Canada’s oldest and largest art and design university with a key strategic goal to increase access to emerging technologies, and enable students to use, create and innovate with these technologies skillfully and responsibly. Artificial Intelligence (AI) and Generative Artificial Intelligence (GAI) are topics the institution is grappling with through an advisory council, a working group with external experts organized by the University’s Cultural Policy Hub, and input from students, faculty, and staff. 

Overall, OCAD University encourages the government to reopen the consultation process with an additional lens to the following issues, which this most recent survey does not address:

Bias in datasets used to train AI and the perceived threat of increased bias if AI developers are restricted to using materials from the public domain for TDM and machine-learning

The recognition of Indigenous sovereignty in the development, training, and application of AI and changes to copyright and intellectual property law

The existing and future impacts on human rights in AI development and implementation

The implications for education around ethical approaches to pedagogy in the age of GAI

In this document, OCAD University makes a set of recommendations. Key among them is the need for continued consultation and consideration of these sensitive and quickly evolving topics. Further, it is critical that creatives be part of these ongoing discussions, and OCAD U recommends that ISED work closely with Canadian Heritage, and include the voices of creatives and creators as part of this work.

The definition of an author or creator may need to be redefined to account for how GAI is involved in the creative process. The revision of these definitions should be led by authors and creators working alongside policy makers, and not by those developing GAI technologies. The government should continue consultation on this topic, alongside the Department of Canadian Heritage and creatives.

Develop regulations to ensure transparency around datasets. With these regulations should come oversight and compliance.

Put safeguards in place that track, record, and disclose how an artist's work is being used in text and data mining (TDM) activities.

Put the onus of developing those mechanisms and obtaining permissions or licensing for the use of content in TDM activities on AI tool and model developers, not content creators.

The definition of an author or creator may need to be redefined to account for how GAI is involved in the creative process.

The revision of these definitions should be led by authors and creators working alongside policy makers, and not by those developing GAI technologies.

Continue consultation on this topic, alongside the Department of Canadian Heritage and creatives.

The government needs to be proactive in putting regulations in place rather than encouraging innovation and then leaving it up to the courts to make legal determinations of liability in instances of infringement (or suspected infringement).

The Ontario Public Service – The Ministry of Public and Business Service Delivery - Privacy, Archives, Digital and Data Division

Technical Evidence

The Ontario government is developing a Trustworthy Artificial Intelligence Framework to establish the rules and groundwork for how we leverage the benefits of AI within government with a human-centered approach. The framework consists of six core principles that were developed in consultation with both the public and experts, including the Information and Privacy Commissioner of Ontario and the Ontario Human Rights Commission. These principles include transparent and explainable, safe, good and fair, responsible and accountable, sensible and appropriate, and human centric.

These principles also form the basis of the policies, guidance, and tools that Ontario continues to develop and implement to ensure responsible use of AI. In particular, the framework is designed to ensure a human-centered approach throughout the entire lifecycle of an AI system, this includes the discovery, building, deployment, and management of the system. Overall, the framework’s objective is to ensure that AI is used safely and responsibly to mitigate common risks such as biased and discriminatory outputs and to offer better and faster programs and services across Ontario.

As Ontario is developing a suite of policy tools to guide the development, deployment, or adoption of AI tools, it is taking a risk-based approach to ensure the thoughtful and appropriate balance of human involvement when AI supports decision-making.

Text and Data Mining

The Ontario Ministry of Public and Business Service Delivery (MPBSD) recommends that ISED provides greater clarity when it comes to text and data mining (TDM) and generative AI systems. MPBSD recognizes the challenge in balancing the need for governance and innovation and the importance of providing the source information used in AI training for greater transparency. Legal frameworks that target exceptions should remain within the scope of academic research, meanwhile other areas such as commercialization and industry-based research should remain out of scope.

This should also include clearly defining when secondary infringement occurs and putting legal guardrails in place to deal with cases when users distribute AI generated content that infringes copyright law. Amendments in this area would need to consider the lack of control developers may have over the training data used in some cases. Hence, it will be important to evaluate the different purposes of models that will need to be measured against varying levels of developer control and accountability.

Amendments to section 29 and 30 of the Copyright Act could impact right-holders from receiving fair compensation for their work being used in TDM activities and hinder industry innovation. Additionally, stricter requirements could unintentionally create barriers for innovation among Canadian researchers and companies, with start-ups and small businesses being disproportionately impacted. MPBSD recommends ISED to consider conducting targeted engagements with civil society, startups and small businesses, publishers, open-source technology communities to gain a better understanding of the potential impacts an amendment to the Copyright Act could cause.

Further, MPBSD recommends that ISED continue to monitor current global trends and standards to ensure Canada’s approach is aligned and to consider best practices. The EU’s AI Act provides a valuable reference point in this area. Article 16 of the EU’s AI Act provides provisions that exempt research activities and AI components offered under open-source licenses. Meanwhile, Article 28b(4)c of the EU’s AI Act requires developers of Large Language Models (LLM) to provide a list of copyrighted materials used in their training datasets. It is important to note that the United States currently does not have an explicit exception for TDM in its copyright legislation, therefore, this can be a leadership opportunity for Canada.

We also encourage ISED to explore options to have accountability and transparency measures in place for AI developers to document, evaluate and/or disclose what copyright-protected content or data is used in the training of AI systems. However, differentiating between protected content and unprotected content may present a challenge for lawmakers in determining a reasonable threshold for disclosure. Therefore, policies in this space would promote the ethical use of AI while taking inspiration from the EU’s AI Act in requiring AI developers, particularly big technology companies, to inspect and disclose copyrighted materials used to develop their systems starting from the discovery phase.

Authorship and Ownership of Works Generated by AI

MPBSD encourages ISED to provide greater clarity on how the current copyright framework applies to the use of copyright-protected works and other subject matter used in the training of AI systems. While MPBSD recognizes that regulating this space is difficult due to the evolving nature of AI systems, we believe that legislative and policy interventions will be more effective in providing clarity around the authorship of AI generated works.

The uncertainty surrounding authorship or ownership of AI-generated works and other subject matter could impact the development and adoption of AI technologies by contributing to a lack of consensus and legal ambiguity. Establishing standards for the attribution of authorship and ownership could address this by providing greater clarity for users who may be unsure if they have copyright protection or not.

MPBSD recommends that the federal government propose a clarification or modification of the copyright ownership and authorship regimes considering AI-assisted or AI-generated works. This can be done by evaluating fair dealing policies and updating the definition of fair dealing to bring it within the context of generative AI. 

Due to the lack of jurisprudence on this issue, MPBSD recommends policies and legislation that support a flexible approach to fair dealing that would allow for a fair balance between innovation and the rights of creators. An example would be an opt-out mechanism for users who do not want their data used for training purposes. This would enable AI training to be deemed as fair, while providing users with a degree of control over their work.

When comparing the approaches in other jurisdictions, it's important to note that Canada’s understanding of this concept is much narrower. The United States allows the restricted utilization of copyrighted material, provided that the user introduces a distinctive element to the copyrighted work, resulting in a transformation of the original work.

Similar to the UK, MPBSD recommends that a guidance document addressing copyright and authorship concerns is released to encourage best practices. This could serve as an effective intermediate step while the federal government continues to work towards developing generative AI copyright policy and Canadian jurisprudence becomes more defined. The UK’s code of practice on copyright is a valuable reference, as their code aims to make licenses for data mining more available. A Canadian code of practice, complemented with clear policy directions, could encourage authors and creators who use AI tools to keep a log of their contributions in creating a final version of their work. This would help promote important AI principles such as, fairness, transparency, accountability, and privacy.

Infringement and Liability regarding AI

MPBSD recognizes the lack of evidence and legal precedents related to copyright infringement and liability raised by AI.

Canada’s current copyright framework is unclear as to whether the use of copyrighted works in the training of AI systems would result in an infringement. It is important for the government to define a clear threshold for human involvement (such as in the development, sale, or licensing of the AI application) that determines when an individual is considered to have authorized copyright infringement and may be held legally liable. This clarification is essential to prevent confusion among developers.

There are also concerns about existing legal tests for demonstrating that an AI-generated work infringes copyright, since establishing these legal tests often rely on the concept of substantial similarity between the original work and the allegedly infringing AI-generated work. Applying these tests to AI-generated content may be complex due to the unpredictable nature of AI models.

Additionally, determining whether an AI system accessed or copied specific copyright-protected content when generating an output poses significant transparency and data protection concerns. It also raises questions over how to regulate and monitor the usage of different outputs and what regulatory compliance will look like for end-users, especially in cases where end-users lack the necessary training to understand the implications of different prompts and outputs.

MPBSD recognizes the barriers and challenges to determining when an AI system accessed or copied a specific copyright-protected content when generating an infringing output. It is important to note that many AI systems are trained on large and diverse datasets from the internet, with significant amount of user-generated content with various degrees of authorship, hence, making it difficult to pinpoint and establish a link as to the exact sources of information used to generate an output.

MPBSD recommends ISED to consider defining criteria that would help identify to what extent developers would be liable for primary or secondary infringement when training AI models and creating AI-generated content. This will be helpful in determining whether an amendment to the Copyright Act should be made to exclude AI-generated content from copyright protection. However, given the complexity of the development and governance of AI models, this would require shared accountability delegated at various levels between the developer, owner and end-user. Any exceptions to criteria in this area should also consider fair use or fair dealing principles to foster innovation and support education purposes.

As mentioned earlier, MPBSD recommends ISED to consider approaches in other jurisdictions such as EU’s AI Act and UK’s approach on the governance of generative AI that will help inform best practices for Canada.

Comments and Suggestions

MPBSD commends ISED for addressing previous feedback provided to ISED (from Canadian Guardrails for Generative AI – Code of Practice) by narrowing in on an important area of law that will be heavily impacted by the growing development of generative AI systems. We appreciate the opportunity to provide feedback as part of the consultation on the Copyright in the Age of Generative Artificial Intelligence.

Oxford Disinformation & Extremism Lab

Technical Evidence

At the Oxford Disinformation & Extremism Lab, we do not collect copyright-protected data but we remain concerned about the potential for violent extremists, malicious actors, and state influence operations to use copy-right protected materials (including news media, video games, music, and images) to create deepfakes, memes, and propaganda using popular media.

Text and Data Mining

Copy-right protected content must be protected when Generative AI is both trained and used. However we also need to consider more serious use cases involving criminal and terrorism offences with this content. Users exploiting and editing copyright-protected content in GenAI platforms have used the tools to produce powerful propaganda that picks up on existing memes, trends, and movements. Further regulation (similar to the the EU Digital Services Act) can help limit this potential by requiring better classifiers in consumer LLMs.

Authorship and Ownership of Works Generated by AI

Further regulation (similar to the the EU Digital Services Act) can help limit the potential for copy-right abuse by requiring better classifiers and hash-sharing to be included in consumer LLMs.

Infringement and Liability regarding AI

Hash-sharing and image classifiers are not perfect at identifying content and companies should invest in developing more robust tools to identify copy-right protected content.

Comments and Suggestions

N/A

R

Re:Sound

Technical Evidence

Re:Sound recognizes the value of AI and is reviewing options to utilize AI in the future to assist with the process of matching sound recordings, enhancing data quality and conducting data analysis, in order to create efficiencies and better serve rights holders.

Text and Data Mining

Re:Sound is the Canadian not-for-profit music licensing company dedicated to obtaining fair compensation for artists and record companies for their performance rights. We advocate for music creators, educate music users, license businesses and distribute royalties to creators — all to help build a thriving and sustainable music industry in Canada.

Re:Sound administers the right of equitable remuneration under section 19 of the Copyright Act on behalf of performers and makers for the communication by telecommunication and public performance of their sound recordings. As this right is not a right of copyright under section 3 of the Act and does not involve the reproduction right, it would not appear to be directly engaged by text and data mining (“TDM”). In support of creators, Re:Sound submits that the use of their sound recordings in TDM must be both authorized and compensated.

The use of sound recordings of musical works in TDM has obvious value, providing an essential input to the creation of AI. Creators should be fairly compensated for the use of their works and have the right to authorize such uses. This can be done without impeding innovation. Digital streaming services license entire catalogues of sound recordings, AI developers can do the same. Copyright owners as well as collective societies such as Re:Sound routinely process data for millions of sound recordings. The volume of data required by AI developers is not an impediment to a fair licensing system that tracks the ingested works and fairly compensates their creators.

No new exceptions should be created under the Copyright Act for TDM. The creation of a sound recording requires considerable time, effort, resources, and talent. AI developers relying on TDM exceptions could use those original sound recordings and generate output that competes with the very creative work that made it possible in the first place. That runs entirely counter to the incentives to create that the Copyright Act aims to achieve.

AI developers should be required to keep records of the copyright-protected content used to train their AI systems in order to ensure that the creators of that content are fairly compensated for the use of their content.

Authorship and Ownership of Works Generated by AI

While the right of equitable remuneration administered by Re:Sound is not a right of copyright and not directly engaged by this issue, in support of creators, Re:Sound submits that the current Copyright Act is sufficiently clear that only a human creator is entitled to copyright protection. No modification of the Act is required at this time, however this issue could be revisited in the future based on developments in jurisprudence.

Infringement and Liability regarding AI

The right of equitable remuneration administered by Re:Sound does not include the remedy of infringement. In support of creators, Re:Sound submits that the existing provisions of the Copyright Act are sufficient at this time to address infringing AI-generated works. However, additional legislative measures are needed to address the following:

  • A requirement that AI developers keep detailed records of the copyright-protected works they ingest to facilitate licensing. Records should be auditable and include how the works were used in the development and operation of the AI systems.
  • Protection for Canadian recording artists against harmful deepfakes and voice cloning undertaken without their consent.
  • Proper identification and labelling of sound recordings generated with AI, particularly those designed to mimic a musical performer’s name, image, voice or likeness.

Comments and Suggestions

Re:Sound thanks the Ministries of Heritage and ISED for this opportunity to provide comments on such an important issue. Re:Sound encourages policymakers to take into account the principles of the Human Artistry Campaign (humanartistrycampaign.com):

(i) technology has long empowered human expression, and AI will be no different;

(ii) human created works will continue to play an essential role in our lives;

(iii) use of copyrighted works and the use of voices and likenesses of professional performers requires authorization and free-market licensing from all rightsholders;

(iv) governments should not create new copyright or other IP exemptions that allow AI developers to exploit creations without permission or compensation;

(v) copyright should only protect the unique value of human intellectual creativity;

(vi) trustworthiness and transparency are essential to the success of AI and protection of creators; and

(vii) creators’ interests must be represented in policy making.

Le Regroupement des artistes en arts visuels du Québec (RAAV)

Preuve de nature technique

Le Regroupement des artistes en arts visuels du Québec (RAAV) a la même mission depuis 30 ans : la défense des droits sociaux, économiques et moraux des artistes en arts visuels du Québec. Cette mission s'articule autour de 3 grands champs d'action : Représenter, défendre et outiller les artistes en arts visuels québécois.

Cette soumission s'appuie sur plus de 220 réponses uniques à un sondage national préparé par CARFAC et le RAAV pour solliciter les commentaires des artistes sur leurs préoccupations concernant les produits d'IA générative, des centaines d'heures passées collectivement à communiquer avec des artistes a travers des comités et rencontres avec des intervenants du secteur culturel au Canada et à l'étranger.

Une partie des artistes en art visuel, particulièrement en recherche, utilisent l'IA générative comme outil de création. Nos membres estiment que l'IA générative est et doit demeurer un outil, en ce sens que la créativité humaine doit toujours primer.

Fouille de textes et de données

Comme l'a noté la Cour suprême du Canada dans l'affaire CCH c. Barreau du Haut-Canada, "une œuvre originale doit être le produit de l'exercice de la compétence et du jugement de l'auteur. L'exercice de l'habileté et du jugement requis pour produire l'œuvre ne doit pas être si insignifiant qu'il puisse être qualifié d'exercice purement mécanique". Ces mêmes critères devraient être appliqués lors de l'évaluation de l'octroi de droits d'auteur à des œuvres produites ou assistées par l'IA. La saisie d'une série d'invites textuelles dans un générateur d'images d'IA est assurément un "exercice purement mécanique", qui n'exige pas de l'utilisateur qu'il fasse preuve de "compétence et de discernement".

Il peut toutefois y avoir d'autres situations dans lesquelles les œuvres d'art générées ou assistées par l'IA répondent aux critères actuels d'octroi du droit d'auteur. Par exemple, si un artiste conçoit un modèle d'IA, entraîne ce modèle avec ses propres œuvres d'art afin que le modèle puisse comprendre et interagir avec les données d'entraînement d'une manière unique spécifiée par l'artiste, le contenu généré par l'IA résultant de ce processus peut être mieux placé pour répondre aux critères du droit d'auteur. À cet égard, les lois existantes sur le droit d'auteur sont suffisantes pour traiter la question de la paternité et de la propriété, et aucune modification de la loi sur le droit d'auteur n'est nécessaire.

Qu'impliquerait une plus grande clarté de la FTD relativement au droit d'auteur, à la fois pour l'industrie de l'IA et les industries créatives au Canada? Des activités de FTD sont-elles menées au Canada? Pourquoi est-ce le cas ou non?

Comprendre comment fonctionne la technologie est essentiel pour permettre l'application du droit d'auteur. Une plus grande clarté permettrait de mieux appréhender le fonctionnement de la FTD, incluant la façon dont les œuvres et autres objets de droit d'auteur sont utilisés. Une clarification permettrait également de déterminer dans quels contextes l'analyse informationnelle est autorisée, ou non, par le régime actuel de droit d'auteur canadien.

Des activités de FTD sont actuellement menées au Canada, afin d'entraîner des modèles algorithmiques. Les activités de développement et d'entraînement de systèmes d'IA sont susceptibles d'impliquer la reproduction de contenus protégés par droit d'auteur (œuvres et autres objets de droit d'auteur tels que des prestations), sans que les titulaires de droits y consentent et reçoivent une juste rétribution. Ceci est évidemment problématique et il importe d'y remédier.

Les titulaires de droits font-ils face à des défis en ce qui concerne l'octroi de licences de leurs œuvres pour les activités de FTD? Le cas échéant, quelles sont la nature et la portée de ces défis?

Oui, les artistes titulaires de droits ne se font pas contacter au sujet de licences pour l'utilisation de leurs œuvres. En outre, il est difficile pour les artistes ou leurs représentants de déterminer quel contenu est utilisé dans le contexte de FTD et quelle est l'ampleur de cette utilisation. Afin de pallier cette lacune, il pourrait être envisagé d'imposer une obligation de transparence ou de tenue de registres auprès des entités développant et entraînant des systèmes d'IA. En utilisant ces mécanismes, les ayants droit pourraient disposer d'informations essentielles à la gestion de leurs droits d'auteur.

Quels types de licences de droits d'auteur pour les activités de FTD sont disponibles et ces licences répondent-elles aux besoins des personnes qui mènent des activités de FTD?

Diverses licences sont disponibles pour les activités de FTD impliquant l'exercice d'un droit réservé aux titulaires de droit d'auteur. Ces licences peuvent être négociées de gré à gré avec les titulaires de droits d'auteur ou être obtenues par le biais d'une société de gestion collective. Ces licences ne semblent toutefois pas être obtenues par les personnes menant des activités de FTD. Ceci crée évidemment un manque à gagner pour les titulaires de droits d'auteur qui peinent à obtenir une juste compensation pour l'utilisation de leurs contenus.

Si le gouvernement devait clarifier la portée des activités permises de FTD, quelles devraient en être la portée et les mesures de sauvegarde? Quel serait l'impact d'une telle exception sur votre industrie et vos activités?

Nous ne sommes pas favorables à l'adoption d'une exception générale permettant la FTD, laquelle serait prématurée et contraire aux engagements du Canada en vertu de divers traités internationaux, tels que la Convention de Berne, l'ADPIC et l'ACEUM lesquels précisent que toute limitation ou exception à laquelle le Canada entend assujettir un droit d'auteur doit être restreinte à certains cas spéciaux où il n'est pas porté atteinte à l'exploitation normale de l'œuvre, ni causé de préjudice injustifié aux intérêts légitimes de l'auteur. Ainsi, si jamais le gouvernement décide d'adopter une exception de FTD, il devra veiller au respect de ses engagements internationaux, par exemple, en veillant à ce que l'exception soit :

Si jamais le gouvernement décide d'adopter une exception de FTD, cette exception ne devrait pas s'appliquer aux droits moraux, mais uniquement aux droits dits « économiques ». Nous réitérons cependant que l'introduction d'une nouvelle exception n'est pas souhaitable, car elle préjudicierait les intérêts des créateurs.

Les développeurs de systèmes d'IA devraient-ils être tenus de tenir des registres ou de divulguer les contenus protégés par le droit d'auteur qui ont été utilisés pour la formation des systèmes d'IA?

L'environnement actuel ne permet pas aux titulaires de droits de savoir si leurs œuvres ont été utilisées pour former des modèles d'IA générative. Ce modèle de fonctionnement opaque encourage l'utilisation non autorisée d'œuvres d'artistes canadiens par des développeurs d'IA et empêche la négociation de licences.

De plus les développeurs et chercheurs du secteur de l'IA générative documentent déjà leurs données d'entraînement, par exemple, par le biais de fiches de données ou « model cards ». Les « model cards » peuvent documenter des informations structurées tels que les noms de domaines où ont été collectés les données d'entraînement. Ces développeurs disposent donc déjà d'outils permettant de documenter les données d'entraînement. L'introduction d'une obligation de transparence ou de tenue de registres ne devrait donc pas entraîner de coûts additionnels pour l'industrie de l'IA.

Quel niveau de rémunération serait approprié pour l'utilisation d'une œuvre dans les activités FTD?

La rémunération des artistes qui acceptent de céder leurs œuvres à des développeurs d'IA dans le but de former des produits d'IA générative devrait être déterminée par les artistes et les développeurs d'IA impliqués dans ces négociations, et non par le gouvernement.

Le gouvernement peut favoriser une solution fondée sur le marché en veillant à ce que les entreprises d'IA opérant au Canada se conforment aux lois canadiennes actuelles sur le droit d'auteur, sans exception, et à ce que les dossiers des œuvres protégées par le droit d'auteur qui ont été utilisées pour former des produits d'IA soient rendus publics.

L'IA générative a, et continuera d'avoir, des répercussions négatives sur les possibilités d'emploi dans les industries artistiques et culturelles. Tout en reconnaissant que les nouvelles technologies peuvent avoir de tels impacts, le gouvernement fédéral peut contribuer à stabiliser ces retombées en veillant à ce que les entreprises d'IA générative opérant au Canada respectent les lois sur le droit d'auteur qui soutiendront le développement de modèles d'octroi de licences.

Titularité et propriété des œuvres produites par l'IA

Y a-t-il des préoccupations quant à l'application des critères juridiques existants pour démontrer qu'une œuvre générée par l'IA viole un droit d'auteur (p. ex. des contenus générés par l'IA qui incluent la reproduction complète ou une partie substantielle d'une œuvre utilisée aux fins d'activités de FTD menées sous licence ou d'une autre façon)?

Il peut être difficile pour un titulaire de droit d'auteur d'identifier le contenu contrefait ou plagié, ainsi que la ou les personnes responsables de la violation et d'établir que la partie qui a violé le droit d'auteur a eu accès à l'œuvre originale, que l'œuvre originale était la source de la copie et qu'une partie importante de l'œuvre a été reproduite.

Les artistes n'ont pas eu la possibilité de négocier l'octroi de licences pour leurs œuvres qui ont déjà été utilisées pour former des modèles d'IA générative. Bien que les modèles de licence ne soient pas utilisés par de nombreuses entreprises d'IA générative grand public, de tels cadres commerciaux existent dans l'industrie de l'IA. Getty Images, par exemple, a publié un générateur d'images d'IA formé exclusivement à partir de son propre contenu. Getty rémunère les créateurs pour l'utilisation de leur travail dans le modèle d'IA.

La pluralité des intervenants, l'incertitude juridique, le manque de transparence quant aux systèmes de gestion des données, ainsi que l'opacité des systèmes d'IA sont autant d'obstacles pour que les artiste puisse faire respecter leurs droit. Cette opacité empêche les parties de négocier des conditions de licence et étouffe le développement de marchés de licence émergents. Pourtant, nous comprenons que les développeurs et chercheurs du secteur de l'IA documentent leurs données d'entraînement : une plus grande transparence sur ces données auprès des ayants droit est donc techniquement faisable.

La résistance des entreprises d'IA générative à s'engager dans des négociations de licence avec le secteur artistique est un autre défi majeur dans l'établissement d'une approche basée sur le marché pour le consentement et la compensation des œuvres d'art utilisées dans le TDM.

Meta, par exemple, a fait valoir que l'imposition d'un régime de licence après coup provoquerait le chaos dans le secteur et n'apporterait que peu d'avantages à chaque artiste, compte tenu de l'insignifiance de leurs œuvres respectives dans le contexte d'ensembles de données plus vastes. Cependant, OpenAI a récemment conclu un accord de licence avec Axel Springer, la société mère de Business Insider et Politico;Il est donc réalisable de mettre en place un régime de licence

Les arguments selon lesquels une seule œuvre protégée par le droit d'auteur est monétairement insignifiante dans le cadre de vastes ensembles de données ne peuvent pas être prouvés car il n'existe pas de cadres obligatoires dans lesquels les négociations de licence peuvent se dérouler. Quoi qu'il en soit, même si la valeur financière d'une œuvre individuelle est jugée faible, cela n'exclut pas le droit de l'artiste de consentir à l'utilisation de cette œuvre et d'être rémunéré pour celle-ci.

Lorsque les entreprises d'IA générative ont utilisé, sans autorisation, les œuvres protégées par le droit d'auteur d'artistes de tout le Canada pour former leurs modèles et accroître la valeur commerciale de leurs produits, elles ont commis une violation du droit d'auteur et doivent par conséquent assumer la responsabilité de ces actes.

Le fait d'exiger des entreprises d'IA générative qu'elles conservent et publient des registres des œuvres protégées par le droit d'auteur utilisées dans la formation de leurs modèles permettra de remédier à la violation du droit d'auteur à grande échelle qui a déjà eu lieu et fournira aux parties concernées les informations nécessaires pour négocier les conditions d'utilisation de ces œuvres. Cela permettra le développement de marchés de licences et renforcera les économies créatives du Canada, tout en accélérant potentiellement la croissance et la concurrence au sein de l'industrie de l'IA elle-même.

La Loi sur le droit d'auteur dispose de mécanismes suffisants pour déterminer la responsabilité en cas de violation de droit d'auteur.

Toutefois, afin de permettre une meilleure lecture de la responsabilité, le Canada pourrait imposer une obligation de transparence ou de tenue de registres auprès des entités développant et entraînant des systèmes d'IA.

Nous recommandons que les entreprises qui violent la loi sur le droit d'auteur, ou toute autre loi canadienne, ne bénéficient pas d'une exemption au motif que ces actions ont déjà eu lieu.

Existe-t-il des approches dans d'autres pays qui pourraient éclairer l'examen de cette question au Canada?

Oui. Dans son projet de règlement « EU AI Act », le Parlement européen a introduit une obligation de transparence, de sorte que les entités qui développent des systèmes d'IA devront publier un résumé suffisamment détaillé de leur utilisation de « données d'entraînement protégées par la législation sur le droit d'auteur », ainsi qu'une information appropriée, claire et visible qui distingue le contenu généré de l'original.

Violation et responsabilité en matière d'IA

Y a-t-il des préoccupations quant à l'application des critères juridiques existants pour démontrer qu'une œuvre générée par l'IA viole un droit d'auteur (p. ex. des contenus générés par l'IA qui incluent la reproduction complète ou une partie substantielle d'une œuvre utilisée aux fins d'activités de FTD menées sous licence ou d'une autre façon)?

Il peut être difficile pour un titulaire de droit d'auteur

D'identifier le contenu contrefait ou plagié, ainsi que la ou les personnes responsables de la violation

D'établir que la partie qui a violé le droit d'auteur a eu accès à l'œuvre originale, que l'œuvre originale était la source de la copie et qu'une partie importante de l'œuvre a été reproduite.

Les artistes n'ont pas eu la possibilité de négocier l'octroi de licences pour leurs œuvres qui ont déjà été utilisées pour former des modèles d'IA générative. Bien que les modèles de licence ne soient pas utilisés par de nombreuses entreprises d'IA générative grand public, de tels cadres commerciaux existent dans l'industrie de l'IA. Getty Images, par exemple, a publié un générateur d'images d'IA formé exclusivement à partir de son propre contenu. Getty rémunère les créateurs pour l'utilisation de leur travail dans le modèle d'IA.

La pluralité des intervenants, l'incertitude juridique, le manque de transparence quant aux systèmes de gestion des données, ainsi que l'opacité des systèmes d'IA sont autant d'obstacles pour que les artiste puisse faire respecter leurs droit. Cette opacité empêche les parties de négocier des conditions de licence et étouffe le développement de marchés de licence émergents. Pourtant, nous comprenons que les développeurs et chercheurs du secteur de l'IA documentent leurs données d'entraînement : une plus grande transparence sur ces données auprès des ayants droit est donc techniquement faisable.

La résistance des entreprises d'IA générative à s'engager dans des négociations de licence avec le secteur artistique est un autre défi majeur dans l'établissement d'une approche basée sur le marché pour le consentement et la compensation des œuvres d'art utilisées dans le TDM.

Meta, par exemple, a fait valoir que l'imposition d'un régime de licence après coup provoquerait le chaos dans le secteur et n'apporterait que peu d'avantages à chaque artiste, compte tenu de l'insignifiance de leurs œuvres respectives dans le contexte d'ensembles de données plus vastes. Cependant, OpenAI a récemment conclu un accord de licence avec Axel Springer, la société mère de Business Insider et Politico;Il est donc réalisable de mettre en place un régime de licence

Les arguments selon lesquels une seule œuvre protégée par le droit d'auteur est monétairement insignifiante dans le cadre de vastes ensembles de données ne peuvent pas être prouvés car il n'existe pas de cadres obligatoires dans lesquels les négociations de licence peuvent se dérouler. Quoi qu'il en soit, même si la valeur financière d'une œuvre individuelle est jugée faible, cela n'exclut pas le droit de l'artiste de consentir à l'utilisation de cette œuvre et d'être rémunéré pour celle-ci.

Lorsque les entreprises d'IA générative ont utilisé, sans autorisation, les œuvres protégées par le droit d'auteur d'artistes de tout le Canada pour former leurs modèles et accroître la valeur commerciale de leurs produits, elles ont commis une violation du droit d'auteur et doivent par conséquent assumer la responsabilité de ces actes.

Le fait d'exiger des entreprises d'IA générative qu'elles conservent et publient des registres des œuvres protégées par le droit d'auteur utilisées dans la formation de leurs modèles permettra de remédier à la violation du droit d'auteur à grande échelle qui a déjà eu lieu et fournira aux parties concernées les informations nécessaires pour négocier les conditions d'utilisation de ces œuvres. Cela permettra le développement de marchés de licences et renforcera les économies créatives du Canada, tout en accélérant potentiellement la croissance et la concurrence au sein de l'industrie de l'IA elle-même.

Devrait-on clarifier davantage la responsabilité dans les cas où une œuvre générée par l'IA viole les droits d'une œuvre déjà protégée par le droit d'auteur?

La Loi sur le droit d'auteur dispose de mécanismes suffisants pour déterminer la responsabilité en cas de violation de droit d'auteur.

Toutefois, afin de permettre une meilleure lecture de la responsabilité, le Canada pourrait imposer une obligation de transparence ou de tenue de registres auprès des entités développant et entraînant des systèmes d'IA.

Nous recommandons que les entreprises qui violent la loi sur le droit d'auteur, ou toute autre loi canadienne, ne bénéficient pas d'une exemption au motif que ces actions ont déjà eu lieu.

Existe-t-il des approches dans d'autres pays qui pourraient éclairer l'examen de cette question au Canada?

Oui. Dans son projet de règlement « EU AI Act », le Parlement européen a introduit une obligation de transparence, de sorte que les entités qui développent des systèmes d'IA devront publier un résumé suffisamment détaillé de leur utilisation de « données d'entraînement protégées par la législation sur le droit d'auteur », ainsi qu'une information appropriée, claire et visible qui distingue le contenu généré de l'original.

Commentaires et suggestions

Le 12 octobre 2023, le gouvernement du Canada a annoncé cette consultation sur le droit d'auteur à l'ère de l'intelligence artificielle générative, avec des soumissions à remettre avant le 4 décembre 2023, et prolongées jusqu'au 15 janvier 2024. La consultation publique est accueillie favorablement par notre association, qui voit en cet exercice une volonté du gouvernement de clarifier les incidences de l'IA sur le droit d'auteur. Cependant, nous craignons que ce court délai pour préparer des recommandations sur cette question complexe ne risque d'aboutir à des résultats inégaux. Certaines parties prenantes du secteur des arts et de la culture pourraient être injustement désavantagées en ayant à équilibrer des ressources et des capacités limitées tout en préparant une analyse réfléchie et documentée, alors que les industries de l'IA générative peuvent certainement allouer des ressources beaucoup plus importantes au processus. Bien que nous apprécions la prolongation du délai de réponse à cette consultation, les futures consultations bénéficieront de délais plus longs.

Notre association ne souhaite pas freiner l'avancement de l'IA, mais désire préserver l'équilibre que la Loi sur le droit d'auteur sous-tend, en veillant à ce que les intérêts des artistes et des titulaires de droits d'auteur soient préservés. En effet, notre association voit le potentiel de l'IA : cette technologie, si elle est adéquatement encadrée, pourrait alimenter la créativité, favoriser la découvrabilité de certains contenus et outiller les créateurs dans le respect de leurs droits.

Il est néanmoins essentiel de prendre conscience des impacts négatifs que l'IA peut avoir sur l'ensemble des secteurs, les fondements de notre société, ainsi que sur les droits des auteurs. Afin de freiner ces risques, notre principale recommandation est de veiller au respect de la Loi sur le droit d'auteur en s'assurant que le consentement des créateurs soit obtenu et qu'une rémunération juste et équitable leur soit versée lorsque leur contenu est utilisé à des fins de fouille de textes et de données (« FTD »). Nous recommandons également l'imposition d'une obligation de transparence auprès des utilisateurs. Spécifiquement, ce cadre devrait obliger la divulgation de toute œuvre utilisée dans le contexte de l'IA. Un tel mécanisme est une action faisable, qui ne pose pas de difficultés techniques et qui jetterait les premières bases de l'édifice, afin d'assurer une rémunération juste et équitable aux artistes et titulaires de droits d'auteur. Dans tous les cas, nous recommandons que le principe des « 3 C » (consentement, crédit et compensation) guide les actions du gouvernement, dans le contexte de cette consultation publique et des possibles amendements à la Loi sur le droit d'auteur qui en découleront.

S

The Samuelson-Glushko Canadian Internet Policy & Public Interest Clinic (CIPPIC), Centre for Law, Technology and Society, University of Ottawa

Technical Evidence

Questions 1-3 are inapplicable to CIPPIC.

Q4) The development of AI is human reliant. While self-programming AI systems have been theoretically formulated by many scholars, to date there has been no successful system of this kind due to current computational constraints. There has been success developing self-modifying AI systems using code-generating language models, meaning that the system is able to manipulate its own hyperparameters to improve its operation. Still, such a model does not involve actual AI self-programming but requires a human to develop and implement the model itself. Similarly, programs such as Codex (Open AI) are able to generate code from natural language inputs. Like all current, publicly available generative AI systems, however, this system was built and made available by a human developer, and further requires prompting by a human user to produce the desired output. As a result, humans are still integral to the development of current AI systems.

Q5) CIPPIC is a public interest technology law clinic, meaning that our area of work primarily involves monitoring and intervening in policy issues and discussions arising at the intersection of law and advancing technologies. In the Canadian legal landscape, AI-assisted and AI-generated content can assist lawyers in reviewing contract formalities, generating memos and factums, as well as with research on specific legal topics as prompted by the user. Legal clients may also use generative AI-systems to ask for suggestions on how to approach a legal issue they are facing, though this should not be considered legal advice.

Text and Data Mining

Q1) Further clarity around copyright and TDM could shed light on: A) The nature of the copyrighted content being scraped for TDM purposes and whether the type of content has implications for TDM-based infringement; B) How broadening authors’ rights to capture TDM may shift copyright’s balance and in so doing create risk and uncertainty for innovative activity (shifting to protection of ideas rather than expressions; violations of tech neutrality; unforeseen effects on non-technological applications of learning techniques); and C) How authors who would like to prevent their copyrightable digital works from being scraped for TDM purposes may do so effectively. Such clarity would ensure that innovators are aware of any copyright-imposed limitations on TDM activities. The creative industry benefits from clarification on author rights and limitations when it comes to preventing their work from being subject to TDM and use in AI-training datasets.

Q2) TDM activities are conducted in Canada across a variety of sectors including research, health record analysis, business intelligence, and AI development. Generally, TDM activities play a significant role in assessing the Canadian population and consumer patterns, making it a key component to informed decision-making, even when separated from generative AI applications. Importantly, TDM describes a practice, not a technology. TDM can occur in analog form and is ultimately a human practice. TDM in the Gen-AI context is just one application of a wider practice that has been a staple of innovative research for some time. Accordingly, any policy proposal to address TDM must consider the implications of the change outside of the AI industry, and for practices involving or similar to TDM such as structuring and indexing information and innovating with technological processing systems (ex. search engines and plagiarism detection). It is worth considering whether we are in the midst of a panic with the emergence of a powerful new technology. Any move to subject the development of AI to the controls and risks inherent to the copyright regime must proceed on the basis of an accurate understanding of how and when AI models interact with works. Equally, addressing the challenges of AI will involve an appreciation of the balancing purposes of copyright law and the nature of author’s rights. Any proposed legislative amendment should proceed only with a clear-eyed appreciation of its consequences for copyright’s virtuous balance.

Q3) Canadian copyright holders often face challenges in licensing their works for TDM activities, in addition to facing challenges in preventing their copyright protected work from being used in TDM. The lack of clarity in Canadian law regarding how TDM may infringe copyrights, namely through unauthorized reproduction of the work, leaves rights holders with uncertainty as to their entitlement to licenses, the specific nature of licensing rights, as well as the proper content of licensing agreements. The further lack of direction on the applicability of fair dealing to TDM also introduces uncertainty as to when infringement occurs. More so, the diverse nature of TDM activities and works sought for TDM has led to a lack of standardization in licensing agreements, as well as undue complexity in licensing terms. This poses challenges for rights holders assessing fair compensation and defining the scope of licenses for use of works.

Q4) In Canada, there are various approaches to licenses for TDM activities. The most common form for publicly accessible data is terms of use agreements, which lay out the scope of permissible data use and any specific conditions for use of TDM results. For example, the terms of use may indicate that TDM is only permitted for non-commercial purposes. Permissible use of the data may be dependent on the payment of a subscription or one-time fee to access the copyrighted material. For many of these sources, the individual seeking to conduct TDM agrees to the terms of use by performing the TDM activity itself, rather than agreeing to a negotiated license with the copyright holder.

Since Canadian copyright law fails to address TDM, those seeking to perform TDM must investigate the terms of use for each specific data source to ensure they have the necessary permissions. Similarly, they must ensure that their proposed use of the data aligns with what is actually allowed under the agreement. As a result, the current scheme poses several challenges since there is no clear, consistent approach to licenses for use, which leaves T&D miners unaware of potential risks regarding copyright infringement and any legal obligations under licenses. Those conducting TDM on a smaller scale face a resource and knowledge gap when compared to big data corporations, leaving them at a greater risk of violating license terms or being priced-out from purchasing licenses. Additionally, those seeking to complete TDM activities often face unduly limitations on data access, which hinders research and innovation pertaining to AI development (ex. the license may limit the TDM results to specific word-limited extractions, rather than the whole of the results). The inconsistencies among licenses further limits accessibility, as the differences in terms of use may prevent comprehensive TDM-driven research where several sources are used.

Q5) TDM activities should be permissible and not cause infringement as long as the training data is not reproduced in any resulting generative output, meaning that the tech sector should be able to use copyrightable works in training datasets for AI in a manner that avoids infringement by reproduction. Canadian copyright protects original expressions of an idea as demonstrated through the exercise of an author’s skill and judgement, but not the idea itself (CINAR v Robinson, 2013 SCC 73 at para 24; Copyright Act, s. 5). Infringement by reproduction occurs where a substantial portion of the original work was copied, which is a question of quality and requires a holistic comparison of the works as a whole (CINAR at para 26). Applying this to TDM, copyright-protects works would likely not be infringed through use in AI training data. Rather than resulting in a durable reproduction, TDM provides the AI with data from which the model can extract knowledge from. The model is essentially deriving meta data to advance its capability to mimic human intelligence. Any technical, temporary reproductions arising from this process benefits from a number of user rights designed to facilitate innovation and its ensuing scientific, economic, and creative benefits. Fair dealing and the exception for temporary reproductions for technological processes both address these benefits. In other words, TDM activities alone do not give rise to an infringing reproduction, publication, or performance of the work. Extending owner’s rights to TDM would unduly extend protection to ideas, information, and data rather than the expression of the work.

Q6) Considering the normative purpose of the Copyright Act, maintaining a balance between the public interest in the encouragement and dissemination of works of art and intellect and obtaining a just reward for the creator, AI developers should disclose the use of copyright-protected content when used in the training of an AI system. Disclosure of the use of copyrighted content provides due acknowledgement to the original creator of the work even when the content itself is not being reproduced in any way, thereby upholding their moral rights to be associated with the work. At the same time, the public receives the benefit of AI systems with stronger and more accurate technological capacity, and AI developers are able to advance the technology through TDM activities without fear of unmerited legal claims. A way to address this normative position would be to specifically include TDM among those qualifying purposes of fair dealing that oblige the user to mention the sources and, if given, the author.

Q7) If Parliament concludes that TDM results in copyright infringement, owner and author remuneration should be addressed through thoughtful application of the current licensing framework. AI Firms are entering into licenses with copyright owners to gain access to their content for TDM activities. These licenses reflect both risk mitigation (avoiding expensive litigation) and the value of content towards enriching training datasets rather than acknowledging existing liability. Structured or edited data, for example, may have greater value than unstructured data for some purposes. If Canada were to adopt this approach, which minimizes the harm to innovation and ensures that markets remain competitive and open to new entrants, Canada should adopt a remuneration model, not an exclusive rights model, for compensating authors. This system could also address authorship – as opposed to ownership – entitlements to ensure that authors obtain discrete benefits from use of their works (rather than exclusively publishers).

Q8) The US has yet to adopt any legislative frameworks specific to TDM. Rather, US courts have applied the doctrine of fair use to permit for TDM, justified by the “transformative use” factor of the test (USC Title 17 – Copyrights, s. 107). Japan has updated its Copyright Act to permit for use of copyrighted works for machine learning, which includes TDM activities. Importantly, Japan only imposes compensation for use where there has been enjoyment of a copyrighted work. Since no one is enjoying the work during TDM, there is no infringement (Article 30-4, Japan Copyright Act). Europe has implemented Articles 3 and 4 to Directive 2019/790 to allow TDM under an opt-out system, such that authors can prevent having their works mined. Since 2014, the UK has provided a copyright exception for TDM under s. 29A of the Copyright, Designs and Patents Act.

Authorship and Ownership of Works Generated by AI

Q1: The current copyright framework does not explicitly address AI-generated works, leading to ambiguities in cases where works are created with significant mixed human and AI involvement. The Copyright Act should be updated to explicitly state that authorship is the exclusive domain of humans. The primary benefit of such an amendment is to head off needless litigation and administrative burdens imposed by actors attempting to assert copyright authorship for algorithms, which, among other things, lack legal personhood to hold such rights.

Copyright theory, legal doctrines, and its underlying rationales, require a human author. The Copyright Act already assumes authors are human; otherwise, s. 6 of the Act, tying the term of copyright to the lifespan of the author, would be meaningless when considering AI. For copyright to vest, a work requires the exercise of skill and judgment, and the work must not be so trivial that it could be characterized a purely mechanical exercise (CCH v Law Society of Upper Canada, 2004 SCC 13 at para 16 [CCH]).

As long as human input is required, AI should be viewed as a tool instead of an author. The originality and creativity elements will need to be judged on a case-by-case basis by examining the input (i.e., prompt) as well as the work itself. A generic input like “a picture of a cat” arguably lacks the originality criteria for copyright protection. The more specific and creative the input, the stronger the case for copyright protection of the resulting work. 

Outputs that lack original human input are unauthored and fall into the public domain.

Q2: The human providing specific inputs to an AI should hold the copyright in a generated work, especially when inputs are not generic, which aligns with the principles of existing copyright law. However, it would be beneficial for the Canadian Government to clarify these aspects to address the evolving landscape of AI-generated works. This approach maintains technology neutrality while ensuring that copyright continues to protect human creativity and expression.

Q3: There are 3 broad national approaches addressing authorship of AI-generated works in copyright law (WIPO Conversation on Intellectual Property and Artificial Intelligence). First, the United States, Australia, and most continental European countries requires human creativity in copyright law and does not extend copyright protection to AI-generated works (see Thaler v. Perlmutter, case No. 1:22-cv-01564, (D.D.C. 8/18/23), at p. 2).

Secondly, the United Kingdom, New Zealand, South Africa, and India award authorship through legislation to the human that arranged the work, and broadly permits fully autonomous or sentient AI to author works (see s. 9(3), Copyright, Designs, and Patents Act of 1988 (CDPA) – UK). The United Kingdom allows a copyright to subsist in AI-generated works by attributing authorship of the works to the human, corporate, or AI machine author that simply arranged the final copyrighted work. Thus, the United Kingdom relies on skill of labour or sweat of brow to determine who arranged the work (CPDA s. 9(3)).

Finally, China and Japan use the judicial system to incrementally expand upon existing legislation by attributing copyright authorship to human programmers and companies that create code dictating AI’s creative decisions and by declining to extend authorship to AI. For example, in Shenzhen Tencent v. Shanghai Yingxun (2019), the Chinese judiciary extended copyright protection to AI-generated works and attributed authorship in the final work to the human author or organization that created the AI. In Gao Yang et al. v. Golden Vision (2020), high-altitude photographs taken automatically by AI merited copyright protection because although humans did not click the shutter-release button to take the photograph (the AI made this decision), humans were solely responsible for making creative decisions that influenced the high-altitude photographs, such as the shooting angle, video recording mode, and video display format.

Our suggestion is that Canada follows the Chinese approach to authorship by attributing copyright authorship to the humans or corporations that create the code or prompt dictating the AI’s output. This approach allows copyright to extend to AI-generated works by focusing on the creativity and originality of the input, rather than judging the amount of creativity and originality in the creation of the output. It also avoids creating sui generis rights or legal fictions to accommodate AI-generated creations. This approach is more open to considering AI’s role as a tool in the creative process.

Infringement and Liability regarding AI

Q1: As previously described, whether a substantial portion of the work has been infringed is a question of quality rather than quantity and requires a holistic comparison of the works as a whole (CINAR at para 26). In other words, infringement can be found for both literal and non-literal copying where the substantial quality of the work is reproduced (CINAR at para 27). As the Supreme Court has ruled, an assessment of substantial copying focuses on whether copied features constitute a substantial part of the original work of the author; thus, the alteration of copied features or its integration into a notably different work may not preclude an infringement claim if a substantial quality of the work has been copied (CINAR at para 39).

As a result, generative AI systems may face infringement claims if their outputted works copy a substantial portion of the quality of the original work that the system was trained on (i.e., substantial reproduction of the copyright protected training data). This will be extremely difficult to monitor considering the expansiveness of most AI training data sets, as elements from a multitude of different works may be combined to produce a generative output. Infringement may be clearer when the prompter asks the AI system to generate an output addressing the specific expression of an author, artist, musician, or other copyright-protected creator. The law currently accommodates this issue by holistically assessing each alleged infringement on a case-by-case basis; copyright infringement is a matter of degree, nuance, and context, and whether a substantial part of a work has been copied is a flexible and fact-specific notion (CINAR at paras 26 & 40).

Due to the nuanced nature of this test, concerns are likely to focus on where an AI-generated work, trained by TDM activities, including copyright-protected expression, is being commercialized. In this situation, creators may raise issues related both to their moral and economic rights. For moral rights, authors may raise concerns related to how the integrity of their work has been manipulated by the AI system, as well as the loss of association with a substantially copied derivate of their original work. Considering that economic rights include the right to authorize reproductions, copyright owners may also raise concerns when another party exercises any of the exclusive rights associated with their work without consent (Copyright Act, s. 3; Théberge v Galerie d'Art du Petit Champlain Inc, 2002 SCC 34 at para 12).

Q2: The vast majority of modern-day AI uses deep learning. Deep learning is a subset of AI and uses artificial neural networks to mimic the human brain’s learning process. The one pitfall of deep learning is that generated outputs are developed using black-box algorithms, meaning that users are unable to see how the deep learning system actually makes its decisions. While we understand the training data inputted to the AI model and can observe the outputs created, we have no way of knowing how the AI actually came to that output. The only indicator we would likely have is the user-generated code used as the original scaffolding for the AI’s learning patterns.

In other words, considering both the black-box algorithm issue and the expansive breadth of datasets, there is no way for us to assess whether or not the AI accessed a specific copyright-protected work. Our knowledge is limited to the inclusion or exclusion of the copyright-protected work in the original dataset.

However, the core principle of how generative AI models function does not implicate copyright infringement: AI researchers do not design AI systems to reproduce training data; they design them to abstractly “learn” from training data. Training data influences the algorithm, but the algorithm does not reproduce the training data.

In rare cases, generative AI models reproduce copyright-protected expression. A precise understanding of how, why, and when this occurs should predicate any policy conclusions the Canadian government draws from this phenomenon. CIPPIC’s understanding of this phenomenon suggests that it arises from specific technological phenomena combined with user-specific prompts. AI researchers are better able to describe the scope and limits of this phenomenon.

Larger concerns arise from copyright protection for the reproduction of abstract concepts of expression, such as fictional characters (ex. the CINAR case). Caselaw will, over time, provide greater clarity over the limits of expression of abstract ideas that reproduce expression subject to copyright protection. Developers of AI systems will need to consider mechanisms for identifying where such occurrences are likely to emerge.

Q3: How the Canadian Copyright Act applies to AI-generated outputs is unclear, leaving businesses and other organizations unsure of liability for copyright infringement through AI-applications. As a result, Canadian enterprises are adopting a patchwork of risk-mitigation strategies. Examples include: restricting training data for AI software such that all data used is either licensed or public-domain; corporate indemnification clauses to protect end users from infringement claims where their use of AI was within the scope of the software’s terms and conditions; and author-applied tags to label their works as being non-TDM friendly.

Some enterprises have combined these approaches to create a robust framework protecting AI application users from infringement liability. For example, Adobe Firefly, which uses generative AI to alter an image based on user-inputted text prompts, has co-founded and implemented the Coalition for Content Provenance and Authenticity (C2PA) and the Content Authenticity Initiative (CAI). The C2PA and CAI work together to permit publishers, creators, and consumers to trace the origin of different types of media using an open technical standard. These tools allow users to add Content Credentials that allow the creator to indicate that generative AI was used in the production of the work. Information about the specific AI models used can be traced by the user, thereby helping to increase transparency and prevent the spread of misinformation regarding the use of generative AI in creative works.

Adobe has also made efforts to ensure that Firefly’s commercial character does not lead to infringement claims by training the AI model on licensed (Adobe Stock) and public domain content, while not training the AI using any subscribers’ personal content. Adobe Stock contributors whose works have been used to train Adobe Firefly are eligible for a “Firefly bonus compensation plan.” This plan represents a license agreement whereby contributors are paid for the use of their work as training data in the Firefly AI software. The purpose of Adobe’s approach was to eliminate the potential for copyright infringement claims, with Adobe going insofar as to offer intellectual property indemnification for any legal issues arising from its use. This indemnification clause provides that as long as the user has used the Firefly product in accordance with the terms and conditions, Adobe will compensate the individual for any IP-related legal claims that may arise. As a result, any generated work incorporating Firefly AI within the scope of its authorized terms will not be subject to personal copyright infringement liability. 

Q4: Yes. Further clarity is required in several areas related to AI systems generally, but specifically for generative AI systems and the works they produce. First, we need greater clarity on what constitutes a “substantial part” of a work is necessary for determining the liability of AI-generated works for infringement. There must be further clarification on the protection of artistic styles and whether that is considered as an unprotected idea or a protected expression of skill and judgement. Secondly, greater clarity is required regarding who would be liable for copyright infringement in these circumstances. Considering that copyright authorship requires an original expression through an exercise of skill and judgement, the legislature must provide clarity on whether or not an AIsystem can truly fulfil these criteria. At its core, AI is a mathematical and statistical computer science tool that analyzes training data for correlations and patterns, and then it uses those patterns to generate a predictive output. Essentially, the AI is regurgitating and recombining data to create a novel output. The issue of whether this constitutes an exercise of skill and judgement, such that the AI system is an author of an original, expressive work, must be resolved by the legislature. If it does, the AI itself would be liable, which would require a novel remedy of some kind. If it does not, is the human prompter who generated the specific output liable? Or does liability fall onto the corporation or individual who coded the AI system itself? 

Such clarity need not originate with legislative initiatives. In any event, special legislative amendments that depart from the usual copyright rules of liability for specific industries or technologies would violate the law’s neutrality. Experience has shown that such solutions are short-lived as the pace of the market and innovation inevitably leaves them behind.

Industry standards, consensus best practices documents, and litigation can all contribute to providing greater certainty around AI innovations. The government can play a role in facilitating multi-stakeholder initiatives that can lend themselves to this end.

Q5) Please see prior response to Authorship Section, Q3.

Comments and Suggestions

Two additional issues merit attention: A) vicarious liability; and B) liability for treatment of rights management information.

A) Vicarious Liability

As with any neutral technology that interacts with works, users can accidentally or deliberately produce AI outputs that infringe copyright. This raises the question of whether proprietors of AI systems used in this way may be vicariously liable for the infringements of its users.

In Canada, vicarious liability for copyright infringement arises primarily through (a) the authorization of an infringing act or (b) the provision of a service primarily for the purpose of enabling acts of copyright infringement (Copyright Act, s. 3(1) & s. 27(2.3). 

Whether a party “authorized” infringement is a question of fact, looking to whether the alleged authorize infringer sanctioned, approved, or countenanced the infringement (CCH at para 38). Authorization can also be inferred from the facts, meaning that both positive acts and sufficient indifference or passivity to known infringement can be grounds for a secondary infringement claim (CCH at para 38). Notably, liability for authorizing infringement does not occur where a person authorizes the mere use of technology that could be used to infringe copyright (CCH at para 38). Courts look to the knowledge the authorizer possessed of the infringing acts, and the degree of control the alleged authorizer exercised over the primary infringer.

Authorizing infringement is relevant to the discussion of copyright, TDM activities, and generative AI outputs because it opens the door to liability for programmers, providers, and users of AI systems who prompt the generation of an infringing work. Regardless of the identity of the author of the infringing AI-generated work, an expansive approach to authorization could extend liability throughout the chain of people responsible for building the AI, training it, and prompting the work’s creation. For example, if TDM is used on copyright-protected material to create training data for an AI, the sale or distribution of that training data to other parties could give rise to a secondary infringement claim. The programmers who train generative AI on such data sets and build the AI system such that it can reproduce a substantial, infringing portion of copyright-protected work could also face liability (see Liability Section, Q1).

The notion that infringement is not authorized where only equipment that could result in infringement is provided further complicates vicarious liability, as this rule indicates that only prompters should be liable for requesting the generation of an infringing work by the model. It must be clarified whether the issue is A) the potential for the AI model to reproduce a substantially similar output to its copyright protected training data or B) the ability for users to prompt such an output.

The manner in which courts have construed the authorization right should prove satisfactory in addressing such risks. AI services do no exercise control over any primary infringer producing outputs that infringe copyright. AI is neutral technology, and only in exceptional circumstances would the ordinary use case implicate a copyright infringement. Similarly, authorization does not require intervention to stop infringements on the part of an otherwise neutral by-stander, including the purveyor of technology used by another to infringe copyright. Copyright does not violate the liberty principle that underlies much of Canadian law: the law does not require Canadians to police the actions of our neighbours.

Expansive interpretation of the authorization provisions of the Act could result in the potential for broad vicarious infringement claims against all parties involved in the development of an AI system. For example, in Voltage Holdings, LLC v. Doe #1, 2023 FCA 194, the plaintiff has sought liability against internet subscribers on the basis of deemed knowledge of infringing acts, alleged control over the point of internet access, and an alleged failure to stop the infringements. Should courts abandon the traditional control test (that looks to the legal and personal relationship between allegedly infringing actors) in favour of assumed technological control, authorization liability could become a significant risk for AI actors. Similarly, if the courts were to overturn long-standing precedent and import into authorization a duty to police or intervene in copyright wrongs, this too would raise red flags.  It is worth commenting that radical policy shifts in the scope and reach of authorization would affect far more than just purveyors of AI services and should provoke a legislative reaction.

The Act’s prohibition on the provision of infringement enablement services should prove adequately focused on bad actors to avoid its inadvertent deployment against content-neutral, general application AI services.

B) Rights Management Information

The Copyright Act’s prohibition in s. 41.22(1) on altering or removing electronic rights management information (“RMI”) associated with an electronic copy of a work ought not to prove a concern for AI entrepreneurs. This is so for both TDM activities and with respect to outputs. RMI liability requires a that a claimant meet a “triple knowledge“ criteria: to be liable, a defendant must (1) “knowingly remove or alter” RMI, (2) without the consent of the copyright owner, and (3) in circumstances where the defendant “knows or should have known” that the removal or alteration “will facilitate or conceal any infringement of the owner’s copyright.” The particularity of these knowledge requirements, and their inapplicability to cases involved fair dealing or other exceptions to infringement – greatly limit the potential of RMI tampering claims to frustrate AI research and commercial applications.

SARTEC - Société des Auteur.e.trice.s de Radio, Télévision et Cinéma

Preuve de nature technique

Nous représentons des créateurs, et plus particulièrement des réalisateurs, artistes interprètes, auteurs de la radio, de la télévision et du cinéma, des acteurs et des musiciens. Spécifiquement, la SARTEC est l'association professionnelle des auteurs de langue française œuvrant à la radio, à la télévision, au cinéma et dans l’audiovisuel.

Nous soutenons la proposition des sociétés et associations suivantes sélectionnez les entités appropriées:

- Artisti,

- L’Union des artistes (UDA),

- L’Association des réalisateurs et réalisatrices du Québec (ARRQ),

- La Guilde des musiciennes et musiciens du Québec

Nous soutenons également les principes évoqués pas la Coalition pour la Diversité des Expressions Culturelles dans son mémoire, lesquels visent à préserver la créativité humaine. Ces mêmes principes guident nos réponses à ce questionnaire.

Notre association voit le potentiel de l’IA à titre d’outil  : plusieurs de nos membres s’en servent d’ailleurs comme source de recherche ou de point de départ de leurs processus de travail. Il est néanmoins essentiel d’encadrer l’utilisation de la technologie, particulièrement dans le contexte de la fouille de textes et de données (« FTD »), puisque le contenu des créateurs est actuellement utilisé à cette fin à leur insu, sans rétribution.

Fouille de textes et de données

Une plus grande clarté et transparence permettraient de mieux appréhender le fonctionnement de la FTD, incluant la façon dont les œuvres et autres objets de droit d’auteur sont utilisés, ainsi que les rôles et responsabilités des différentes parties prenantes. Ceci permettrait également de déterminer : (i) dans quel(s) contexte(s) l’analyse informationnelle est autorisée, ou non, par le régime actuel de droit d’auteur canadien et ainsi, (ii) quelles licences et rétributions doivent être versées aux titulaires d’œuvres et d’autres objets de droit d’auteur. 

Oui, des activités de FTD sont actuellement menées au Canada, afin d’entraîner des modèles algorithmiques. Les activités de développement et d’entraînement de systèmes d’IA peuvent impliquer la reproduction de contenus protégés par droit d’auteur (œuvres et autres objets de droit d’auteur tels que des prestations), sans que les titulaires de droits y consentent et reçoivent une juste rétribution. Ceci est évidemment problématique et il importe d’y remédier. En outre, il est essentiel que le consentement (de type « opt-in » et non « opt-out ») des titulaires de droits soit obtenu préalablement à toute reproduction de leurs contenus protégés, et qu’une rétribution juste et équitable leur soit versée en contrepartie de cette utilisation. L’obtention de ces consentements devra prendre en compte les particularités de chaque contenu reproduit. Par exemple, dans le cas des prestations fixées, un contentement distinct devra être obtenu auprès des artistes-interprètes si l’autorisation initialement consentie aux producteurs ne couvre pas la FTD. Afin de résoudre ces enjeux, le Canada devrait ratifier le Traité de Beijing, ce qui permettrait aux artistes-interprètes audiovisuels d’exercer un meilleur contrôle sur leurs prestations, notamment lorsque celles-ci sont incorporées dans des œuvre.

En outre, il est difficile pour les titulaires de droits d’auteur de déterminer quel contenu est utilisé dans le contexte de la FTD et quelle est l’ampleur de cette utilisation. Afin de pallier cette lacune, il pourrait être envisagé d’imposer une obligation de transparence ou de tenue de registres auprès des entités développant et entraînant des systèmes d’IA.

Diverses licences sont disponibles pour les activités de FTD impliquant l’exercice d’un droit réservé aux titulaires de droit d’auteur, à savoir la reproduction. Ces licences peuvent être négociées de gré à gré avec les titulaires de droits d’auteur ou être obtenues par le biais d’une société de gestion collective. Ces licences ne semblent toutefois pas être obtenues par les personnes menant des activités de FTD. Ceci crée évidemment un manque à gagner pour les titulaires de droits d’auteur qui peinent à obtenir une juste compensation pour l’utilisation de leurs contenus. Afin de remédier à cette problématique, plusieurs mécanismes peuvent être envisagés, tels que l’introduction d’un droit à une rétribution équitable pour la FTD ou un droit à rétribution via un mécanisme semblable à celui de la copie pour usage privé.

Nous ne sommes pas favorables à l’adoption d’une exception générale permettant la FTD, laquelle serait d’ailleurs contraire aux engagements du Canada en vertu de divers traités internationaux, tels que la Convention de Berne, l’ADPIC et l’ACEUM lesquels précisent que toute limitation ou exception à laquelle le Canada entend assujettir un droit d’auteur doit être restreinte à certains cas spéciaux où il n'est pas porté atteinte à l’exploitation normale de l’œuvre, ni causé de préjudice injustifié aux intérêts légitimes de l'auteur. Ainsi, si jamais le gouvernement décide d’adopter une exception de FTD (ce que nous ne recommandons pas), il devra veiller au respect de ses engagements internationaux, par exemple, en veillant à ce que l’exception soit : (i) limitée à des cas spécifiques (par exemple, à des fins de recherche) ; (ii) assujettie à des conditions d’application strictes (par exemple, l’accès à l’œuvre ou objet de droit d’auteur doit être licite) ; et (iii) assortie du versement d’une juste rétribution au bénéfice des titulaires de droits d’auteur, ainsi que d’un mécanisme de retrait (« opt-out ») pour les titulaires de droits d’auteur. Finalement, cette exception ne devrait pas s’appliquer aux droits moraux, mais uniquement aux droits dits « économiques ». 

Les développeurs de systèmes d'IA devraient être obligés de tenir des registres et de divulguer les contenues protégés par le droit d'auteur utilisés pour la FTD. Il s'agit, selon nous d'une obligation essentielle qui devrait être intégrée à la Loi sur le droit d'auteur.

Une rétribution devrait être versée pour toute utilisation d'une oeuvre dans une activité de FTD : Le niveau de cette rémunération doit être juste et équitable, basé sur les utilisations faites des contenus protégés. Dans tous les cas, la rémunération devrait être arrimée avec les autorisations obtenues et prendre en compte les particularités de chaque contenu reproduit. Par exemple, dans le cas des prestations

Comme exposé plus tôt, il n’est pas recommandé d’introduire une exception de FTD au Canada. Au contraire, il est essentiel de veiller au respect de la Loi sur le droit d’auteur en s’assurant que le consentement des créateurs soit obtenu et qu’une rétribution juste et équitable leur soit versée lorsque leur contenu est utilisé à des fins de FTD. Nous recommandons également qu’une obligation de transparence ou de tenue de registres soit imposées aux chercheurs et développeurs de systèmes d’IA générative, dans le contexte de la FTD. Si toutefois le Canada souhaite introduire une exception de FTD, il devra veiller à ce que cette exception respecte les balises internationales, soit d’application limitée et assortie d’un mécanisme de retrait (« opt-out ») pour les titulaires de droits d’auteur. À cette fin, le gouvernement canadien pourrait examiner la situation prévalant au sein de l’Union européenne, la Suisse et le Royaume-Uni.

Titularité et propriété des œuvres produites par l’IA

L’incertitude entourant la titularité ou la propriété d’œuvres et d’autres objets du droit d’auteur produits par l’IA ou à l’aide de l’IA est un ENJEU MAJEUR.

Cette incertitude a notamment des répercussions sur la rémunération des artistes tels que les musiciens, dont le contenu se retrouve « dilué » sur des plateformes telles que Spotify. L’absence de protection des créations « artificielles » a par ailleurs une incidence sur la protection des prestations des artistes-interprètes. En outre, il existe une incertitude entourant la titularité et la rémunération liées à une prestation « artificielle » incorporant la voix, l’image ou la ressemblance d’un artiste-interprète, alors que celui-ci n’a pas autorisé une telle incorporation. Également, selon la Loi sur le droit d’auteur, une « prestation » ne sera protégée que si elle est « rattachée » à une œuvre. Par conséquent, le droit des artistes-interprètes pourrait être mis en péril si ces derniers interprètent des créations « artificielles », non protégées par droit d’auteur. Nous recommandons donc que ces prestations soient protégées et ce, indépendamment du fait que les artistes interprètent ou exécutent des créations « artificielles », non protégées par droit d’auteur. À cette fin, plusieurs options sont envisageables, dont les suivantes : (i) revoir les définitions de « prestation » et d’« artiste-interprète » au sein de la Loi sur le droit d’auteur ; (ii) introduire les droits moraux pour les artistes-interprètes audiovisuels (par exemple, par le biais de la ratification du Traité de Beijing) ; et (iii) introduire des présomptions de violations des droits économiques et/ou moraux des artistes-interprètes lorsque leurs prestations (ou des composantes de celles-ci telles que la voix ou l’image) sont reproduites dans un contexte d’IA générative à leur insu.

Nous ne recommandons pas de protéger des créations « artificielles » dépourvues de créativité humaine. Quoique la jurisprudence canadienne ait défini les concepts d’originalité et d’autorat, ces concepts demeurent flous et malléables. Le gouvernement pourrait donc préciser qu’un « auteur », aux fins de la Loi sur le droit d’auteur, est OBLIGATOIREMENT un être humain. Ceci permettrait d’avoir plus de certitude quant à l’application de la Loi sur le droit d’auteur. Il est également recommandé de modifier la définition d’« artiste-interprète » et de « prestation » au sein de la Loi sur le droit d’auteur, afin que cette dernière ne soit plus uniquement rattachée à des « œuvres ».

Le Royaume-Uni, l’Irlande et la Nouvelle-Zélande sont souvent cités en exemple sur cette question. Les législations de droit d’auteur de ces pays attribuent en effet la titularité d'œuvres générées par ordinateur à la personne qui a pris les dispositions nécessaires à la création de l’œuvre créée. Nous ne recommandons toutefois pas d’emprunter cette voie, car ces dispositions ont été introduites dans un contexte étranger à l’IA générative. Or, cette technologie soulève des questions bien plus complexes. Au surplus, avant même d’adresser la question de la titularité des « œuvres générées par ordinateur », il convient de statuer sur leur protection. Afin de clarifier cette question, nous recommandons de clarifier certains concepts phares de la Loi sur le droit d’auteur, tel que l’autorat en précisant qu’un « auteur », aux fins de la Loi sur le droit d’auteur, est obligatoirement un être humain.

Violation et responsabilité en matière d’IA

Il peut être difficile pour un titulaire de droits d’auteur : (i) d’identifier le contenu contrefait, ainsi que la ou les personnes responsables de la violation ; et (ii) d’établir que la partie qui a violé le droit d'auteur a eu accès à l'œuvre originale, que l'œuvre originale était la source de la copie et qu’une partie importante de l'œuvre a été reproduite. Les mêmes difficultés s’appliquent dans le contexte des prestations, en ce sens qu’il peut être difficile pour un artiste interprète : (a) d'identifier la ou les personnes responsables de la violation ; et (b) d’établir que la partie qui a utilisé sa voix, son image ou sa ressemblance a eu accès à une prestation préexistante (plutôt que simplement la voix, l’image ou la ressemblance), que la prestation (et non simplement la voix, l’image ou la ressemblance) était la source de la copie et qu'une partie importante de la prestation a été reproduite.

La pluralité des intervenants, l’opacité des systèmes d’IA sont les principaux obstacles qui empêchent de déterminer si un système d'IA a accédé ou copié un contenu spécifique protégé par le droit d'auteur lors de la génération d'un extrant.

Les entreprises pourraient obtenir des licences, mais nous n’avons pas connaissance d’un tel phénomène. Alternativement, il est possible d’utiliser du contenu libre de droits, incluant des œuvres tombées dans le domaine public. Dans ce dernier cas, l’utilisation d’œuvres du domaine public est susceptible d’avoir une incidence négative sur la qualité des systèmes d’IA, puisque les données d’entraînement pourraient être désuètes.

La Loi sur le droit d’auteur dispose globalement de mécanismes suffisants pour déterminer la responsabilité en cas de violation de droit d’auteur. Toutefois, le Canada pourrait imposer une obligation de transparence ou de tenue de registres auprès des entités développant et entraînant des systèmes d’IA.

Certaines approches internationales pourraient éclairer l'examen de cette question : Dans son projet de règlement « AI Act », le Parlement européen a introduit une obligation de transparence, de sorte que les entités qui développent des systèmes d’IA devront publier un résumé suffisamment détaillé de leur utilisation de « données d’entraînement protégées par la législation sur le droit d’auteur », ainsi qu’une information appropriée, claire et visible qui distingue le contenu généré de l’original. Cette approche nous paraît louable, mais le Canada devrait aller encore plus loin. En outre, l’obligation de transparence canadienne devrait également s’appliquer aux prestations et à leurs composantes (voix, image et ressemblance de l’artiste-interprète), ainsi qu’aux résultats générés par ou avec IA.

Commentaires et suggestions

La consultation publique est accueillie favorablement par nos associations, lesquelles voient en cet exercice une volonté du gouvernement de clarifier les incidences de l’IA sur le droit d’auteur. Nos associations ne souhaitent pas freiner l’avancement de l’IA, mais désirent préserver l’équilibre que la Loi sur le droit d’auteur sous-tend, en veillant à préserver la culture canadienne, la créativité humaine, ainsi que les intérêts des titulaires de droits d’auteur. Pour ce faire, nous recommandons que les principes regroupés sous l’acronyme « A.R.T. » (Autorisation, Rétribution et Transparence) guident les actions du gouvernement, dans le contexte de cette consultation publique et des possibles amendements à la Loi sur le droit d’auteur qui en découleront. Par ailleurs, il est important que la consultation publique ne se limite pas aux intérêts des auteurs et autres titulaires de droits d’auteur sur des œuvres, mais qu’elle couvre également les intérêts des titulaires des droits dits « voisins », tels que les artistes-interprètes. L’IA générative bouleverse en effet grandement ces créateurs, notamment dans le contexte de l’hypertrucage (ou « deepfake » en anglais). À ce chapitre, les artistes-interprètes audiovisuels ne disposent pas de droits suffisants pour protéger leurs prestations, incluant dans le contexte de l’IA générative et de l’hypertrucage. Afin de pallier cette situation, il est recommandé d’étendre les droits exclusifs et les droits moraux de ces artistes, par exemple, en ratifiant le Traité de Beijing.

SaskBooks

Technical Evidence

1) Our organisation promotes and sells copyright-protected content. We do not use this content in training datasets; we may use its metadata, but not the content itself.

2) Our organisation does not use training datasets to develop AI systems.

3) As far as I know, there are very few measures taken to mitigate liability risks regarding AI-generated content infringing on existing copyright-protected work, as evidenced by the number of content creators and publishers now searching for their copyright-protected/licensed works in existing datasets.

4) Our organisation does not develop AI systems.

5) In our area of work, businesses may use AI-generated content for marketing and promotional purposes, or for graphic design. This practice is discouraged, as there is no way to currently license the use of copyright-protected material present in datasets.

Text and Data Mining

1) More clarity around copyright and TDM in Canada would mean the AI industry would be very clear on how and when licensing fees need to be paid to content creators and/or licensors.

2) TDM activities are being conducted globally; it's difficult with VPN technology to pinpoint exactly where data is being scrubbed and/ore stored.

3) Rightsholders are absolutely facing challenges in licensing their works for TDM activities. To date, we are not aware of any TDM operators paying copyright licensing fees. Therefore, the data being used to train/inform AI includes unlicensed copyrighted material.

4) We are not aware of any TDM copyright licenses in existence or being offered.

5) The impact of an amendment to the Act to clarify the scope of TDM activities would ensure content creators and licensors of copyrighted material would be able to earn revenue from the use of their copyrighted material. It would mean material already included in datasets would be revenue-generating rather than outright theft.

6) AI developers must keep records of and disclose what copyright-protected content is used in training of AI systems. Not to do so is plagiarism.

7) A per-use licensing fee for access to datasets, similar to what is done with educational/post-secondary/library licensing would be efficient and effective. Existing copyright collection agencies like Access Copyright would be the perfect agency for these collections and disbursements.

8) The work being done in the United Kingdom and the European Union on this topic is interesting and leading-edge.

Authorship and Ownership of Works Generated by AI

1) AI-generated material cannot be copyrighted. It is not an unique creation of intellectual property created by a person. It is content generated by algorithm using extant data (therefore, by definition, not unique. Even though those datasets include millions of records; material generated from their use cannot, by definition, be 'unique'). It is content generated by algorithm, not by a person. The person is a user, not a creator. It's doubtful these facts will have any effect on the development and adoption of AI technologies, but Canada should follow the UK's example and rule that AI-created material cannot be copyrighted as it is not the unique creation of a person.

2) The Government *must* clarify copyright and authorship in light of AI-generated works. Not to do so will be to undermine the entire purpose of copyright and the ownership of intellectual property. It will undermine content creators' ability to earn a living from their IP.

3) Again, the UK and EU is doing good work on this topic.

Infringement and Liability regarding AI

1) The extant tests for determining whether AI-generated work infringes on copyright are expensive, unreliable, and behind the development of the algorithms that generative AI uses. Considering the extant datasets have already been demonstrated to be using copyright-protected material without permission or license indicates the current tests do not work or are not being applied by the developers of AI.

2) Barriers to determining whether an AI system accessed or copied specific copyright-protected content include: cost of the programs; reliability of the programs; enormous datasets; reliability of the data included in the dataset (ie. are the datasets being provided for examination the same ones as were used in the AI system?); the capacity of content creators and licensors in being able to access, understand, and examine those datasets.

3) It is unclear whether most businesses truly understand the liability of using AI-generated material that has already been demonstrated to have been using unlicensed copyright-protected material. Businesses that want to commercialise these applications must ensure the datasets they're using are only using licensed works or works in the public domain. Regular audit of their datasets should be part of the business model/workflow.

4) There must be greater clarity on where liability lies with AI-generated works infringing on existing copyright-protected works. People using these works may not know they are using plagiarised material.

5) UK and EU

Comments and Suggestions

Canada must first repair the broken Copyright Act to ensure more clarity is provided around what constitutes "fair use". As it stands, businesses, individuals, and, frankly, pirates who want to use unlimited datasets can do so with Canadian material because these definitions are so widely interpretable as to make Canadian copyright almost useless.

The Schwartz Reisman Institute for Technology & Society (SRI) at the University of Toronto (U of T)

Technical Evidence

N/A

Text and Data Mining

Introduction

The Schwartz Reisman Institute for Technology & Society (SRI) is a research institute at the University of Toronto (U of T). Its mission is to be the world’s leading institute for innovative research and practical solutions to help ensure that artificial intelligence (AI) and other advanced technologies benefit all of humanity. Central to this mission is SRI’s belief that mitigating potential harms and unlocking the many benefits of AI will require new approaches in law and regulation. What follows is SRI’s response to the Government of Canada’s Consultation on Copyright in the Age of Generative Artificial Intelligence, issued by Innovation, Science, and Economic Development Canada (ISED).

Text and Data Mining (TDM)

If the Government were to amend the Act to clarify the scope of permissible TDM activities, what should be its scope and safeguards?

If the government were to amend the Copyright Act to clarify the scope of permissible TDM activities, it should do so in a way that makes it clear that copyrighted data can be used to train AI systems. Copyrighted works in TDM should not require additional authorization from rights-holders. The purpose of copyright law is to motivate the production of new creative works by offering some monopoly protection for a limited period of time to people who invest in producing new ideas. It does not protect ideas themselves—in fact, it encourages the free flow and exchange of them. Barring AI systems from training on these existing ideas runs counter to the principles of copyright law.

The boundaries around what copyright does and does not protect can be explored through the example of a painting. If I see your painting—at a gallery, museum, or posted online—there is nothing to prevent me from reproducing some aspect of it. I can decide to adopt your distinctive style of painting, take your idea to paint a cat exploring outer space, or even reproduce the same image you’ve created in my own style. The only thing copyright protects against is exact copying. The bounds of these protections are purposefully narrow, so as to encourage creativity and innovation. We want creators to draw inspiration from other creative works, to expand upon them, and to bring new creations to life. The line for copyright has always existed to protect creators just enough that they are motivated to continue creating, without stifling the ability of other creators to draw inspiration and make similar but iterative works.

It may be tempting to look at an AI system being trained on copyrighted material and suggest that the use of such materials should not be permitted, or that copyright holders should be compensated for this use. Yet copyright is not based on labour theory—we are not protecting creative works simply because someone put their labour into making those works. We are protecting them for economic reasons. Simply put, creators need some level of protection in order to receive the economic benefits of producing new creative works. Without these benefits, there would be little motivation to produce new works, and we would lose the innovation and continued creation required to stimulate the economy.

Existing copyright does not prevent new artists from studying the works and styles of artists who came before them. Rather, we want to motivate this free flow of ideas between creators as a way of prompting continued creativity and innovation. Barring AI systems from being trained on copyrighted material would therefore distort copyright law and its purposes.

    Technological Protection Measures

Having established that barring TDM goes against the core purpose of copyright law, let us quickly address technological protection measures, or TPMs. Access to copyrighted works can be prevented by using TPMs such as digital locks or digital rights management. As such, another question that may be worth exploring is whether, and to what extent, TPMs interfere with the data scraping activities discussed above.

Although the portion of copyrighted materials currently protected by TPMs is relatively small, this may increase to an overwhelming majority if rights holders remain unhappy with TDM, and turn to TPM as a way of pushing back against their materials being scraped from the web. As such, although allowing TDM is the right approach in keeping with copyright law, it is worth noting the existence of TPMs and other existing mechanisms that may allow rights holders to ‘rebel’ against this choice. The compensation scheme suggested in the section below would help in addressing these types of concerns.

    Individual licensing

The Consultation Paper notes that various stakeholders have been resistant to the idea of licensing, arguing that this will stifle innovation. This is a valid argument—individual licensing is costly, both in terms of actual monetary cost and the amount of time required to track down and bargain with every individual license holder. Further, most generative models are trained on a scrape of the web. Any kind of individual licensing system would make this type of training impossible, introducing serious barriers into the process of creating new generative models. This would run counter to the general purpose of copyright law, by demotivating the production of new works and ideas.

Authorship and Ownership of Works Generated by AI

Ownership of AI-Generated Works

The consultation paper presents the following three approaches for analysis:

1. Clarify that copyright protection apply only to works created by humans

2. Attribute authorship on AI-generated works to the person who arranged for the work to be created

3. Create a new and unique set of rights for AI-generated works

We believe that something akin to approach 3, creating a new and unique set of rights for AI-generated works, is the best approach. However, we will first quickly address options 1 and 2.

The approach that copyright protection only applies to works created by humans likely does not have very much staying power. As AI capabilities increase, determining what has been created by humans will become more difficult to distinguish. The consultation paper suggests that an AI-generated output could be copyrighted when a human author uses AI to generate the output, but in the process uses skill and judgement. Deciding, however, the appropriate level of skill and judgment for AI becomes murkier as the technology advances and new capabilities emerge. Strictly applying copyright to human-made works may cause future issues and calls for a more nuanced approach. Moreover, as noted above, the goal of copyright law is not to protect human labour per se but rather to encourage creativity and innovation. If machines become capable of creativity and innovation, it would seem appropriate to encourage that by providing the economic protection that copyright offers human creators.

The next proposed approach is to attribute authorship of AI-generated works to the person who “arranged for the work to be created.” This would presumably operate similarly to a business model, in which works created by the employees of a company in their capacity as an employee are owned by the business (or organization, academic institution, etc) that employs them. This is a viable approach, and having the existing example of works created by employees and owned by companies would help serve as a model for resolving problems that may arise.

However, the most viable approach seems to be introducing a new type of compensation scheme to motivate creators. Copyright provides us with a useful mechanism for incentivizing and compensating creators. However, in the age of generative AI, this mechanism begins to break down. Traditional copyright was not created to account for the speed at which AI systems can learn and process vast amounts of new information, nor the abilities of these systems to generate new works in a matter of seconds. It is not actually consistent with copyright to suggest that these systems cannot have access to ideas and images that exist—albeit with certain protections—in the public sphere. Yet, this suggestion is often posed in the context of generative AI, seemingly motivated by the feeling that it doesn’t seem fair that a company can scrape the whole internet for free, and then take control of vast amounts of wealth and market share.

We should not twist copyright into something it’s not on the basis of feelings about what is and is not fair. However, we should pay attention to the fact that copyright is an economic tool for motivating creativity and the continued production of new ideas and works; if we think this motivation will be hindered by allowing generative AI companies to capture too much of the market share, then we will need some kind of solution.

What we are faced with in the age of generative AI is a fairly momentous transformation of the nature of our production relationship. As a result, what we likely need is a method for sharing the surplus: the wealth that is generated from these systems. Canada could get in front of these coming complications by creating a new type of compensation scheme for creators. Such a compensation scheme could avoid the concerns that generative AI companies operating on a free scrape of existing ideas would capture too much of the surplus, dominating markets and demotivating the creation of new ideas and works by smaller creators.

This compensation scheme could take various forms, but should share the core economic motivations of copyright law. The government could decide to impose a tax on generative systems and divert the proceeds towards funding various projects that incentivize new ideas and works, such as education or stipends for creators. Or, it could be decided that what is needed is an intermediary organization that receives a share of the profits generated by an AI system that’s producing new works, and distributes those profits to the creators in its network (similar to the American Society of Composers, Authors and Publishers). Whatever form the compensation scheme takes, it should tackle our economic concerns (stifling the production of new ideas and works due to a lack of motivation for individual creators), rather than our moral ones (it doesn’t seem fair that this generative system can create images at a fraction of the speed and cost that humans can).

Infringement and Liability regarding AI

Infringement and Liability

The question of liability in the context of generative AI systems largely boils down to this: who should be liable for a generated work that violates existing copyright laws? The AI developer, who trained the model on copyrighted works, or the user who prompted the model to produce a work that duplicates or too closely resembles an original piece?

Generally speaking, liability should lie with the user who prompts the model. If a user wants to make a system reproduce Van Gogh’s Starry Night, and gives it directions to do that, then fault will lie with that user. This is likely to be the case in the majority of copyright infringements. However, should there be instances where a user is unfamiliar with Starry Night, and the generative system just decides to output an exact replica of this work because it remembers Starry Night from its training data, then liability will lie with the system developer. Determining who caused the copyright violation should be relatively straightforward in the majority of cases, simply requiring a retracing of the steps that caused the infringing output to be generated.

Comments and Suggestions

N/A

The Screen Composers Guild of Canada, and, The Songwriters Association of Canada

Technical Evidence

N/A

Text and Data Mining

2.1.1 Are there concerns about existing legal tests for demonstrating that an AI-generated work infringes copyright (e.g. AI-generated works including complete reproductions or a substantial part of the works that were used in TDM, licensed or otherwise)?

Existing legal tests are already sufficient to determine ingesting copyrighted materials into AI programs, without permission of the owner, is an infringing activity.

o  At the input level, ingesting copyright works into an AI program typically results in a permanent copy of those works being made and stored by the AI program, including for commercial purposes (e.g. supporting keyword based user prompts). As such, TDM activity directly engages the right of reproduction, and would not trigger the exception for ‘temporary reproductions for Technological Processes’ nor the fair dealing exception.

o  At the output level, generative AI models are designed to create works that compete in the marketplace with the very copyrighted works they are trained on. TDM activity that results in the generation of content that derives from, emulates or copies the characteristics of specific copyrighted works could potentially infringe copyright.

Generative AI content requires the making and storing of a copy of copyrighted works and the resulting work seeks to compete directly with the copyrighted source works in the marketplace.

As such, the Screen Composers Guild of Canada (SCGC) and the Songwriters Association of Canada (SAC) submit that no new legal tests are required to determine whether TDM ingestion or the creation of outputs infringe copyright. On the contrary, existing legal tests clearly indicate that exceptions linked to ‘temporary reproductions’ and ‘fair use’ do not apply to TDM activity used to program generative AI programs.

Nor are any new exceptions from these tests required. SCGC and SAC are deeply opposed to any additional exceptions in the Copyright Act to exempt AI companies from their legal obligations.

Any additional exception could be contrary to Canada's commitments under various international treaties, such as the Berne Convention, TRIPS and CUSMA (which specify that any limitation or exception to which Canada intends to subject a copyright must be restricted to certain special cases where it is not detrimental to the normal exploitation of the work, nor causes unjustified prejudice to legitimate interests of the author).

2.1.2 What are the barriers to determining whether an AI system accessed or copied a specific copyright-protected content when generating an infringing output?

A lack of substantive data and accurate recordkeeping are current barriers for determining the scope of infringing activity by AI programmers.

At the November 14, 2023 stakeholders’ roundtable hosted by ISED and PCH, a representative of a Generative AI company confirmed that their platform is capable of ingesting “the entire internet.”  That representative also confirmed that they have the capacity and capability to track and record the specific copyrighted works ingested in that process.

As such, the Screen Composers Guild of Canada (SCGC) and the Songwriters Association of Canada (SAC) submit that relevant records on TDM sources and activity must be kept by AI programmers to identify all stakeholders -- including authors (i.e. composers and lyricists), performers and any other owners of intellectual property in musical works or sound recordings ingested in the course of an AI system’s machine “learning”.

-----------------------------------

2.1.3  Are rights holders facing challenges with TDM activity and in licensing their works for TDM activity? If so, what is the nature and extent of those challenges?

Music is a licensing business.  The key challenge facing rights holders when it comes to licensing their works for AI programming purposes is that AI companies typically do not ask rights holders for a licence before ingesting their work. By definition, we cannot licence works that are taken without our knowledge or permission.

As TDM activity arguably engages numerous copyrights, various licenses are available for TDM activities. These licenses can be negotiated directly with copyright holders or obtained through collective societies. Generally speaking, AI companies do not seek these licenses before ingesting copyrighted works. This makes it impossible for copyright holders to obtain fair compensation.

The Screen Composers Guild of Canada (SCGC) and the Songwriters Association of Canada (SAC) note that there are ethical AI companies who actively obtain licenses for copyrighted material before using it for TDM purposes.  This eliminates any argument from less ethical AI companies that licensing models are impractical or impossible for TDM activity.

Numerous and highly respected Canadian collectives exist precisely to work with copyright users to ensure their whatever rights their specific uses engage, they are aligned with the Copyright Act. 

SCGC and SAC strongly submit that authors of works ingested into AI programs and platforms must be able participate fairly or equitably in all royalty revenue streams to which their work contributes.  A new right of remuneration for authors and creators whose work is used by AI developers, with or without consent, may be required.

However, for such a remedy to be practical, AI companies need to be proactive and transparent about their use of copyrighted material for TDM programming, and the brief history of this technology has already demonstrated that most AI companies will voluntarily join that discussion.

At a minimum, AI companies operating in Canada should be required whether by legislation or regulation, to (i) maintain accurate and transparent records of the copyrighted works ingested into their AI programs, (ii) to obtain consent in the form of a licence before ingesting it, and (iii) to pay compensation in exchange of such use.

Authorship and Ownership of Works Generated by AI

2.2.2. Should the Government propose any clarification or modification of the copyright ownership and authorship regimes in light of AI-assisted or AI-generated works? If so, how?

The Screen Composers Guild of Canada (SCGC) and the Songwriters Association of Canada (SAC) note the three broad options for managing AI-assisted and AI-generated works the consultation paper, and support Option 1:while the Copyright Act is based on a philosophy of providing incentives for human creativity —including, for example, the tying of the general term of copyright ownership to the life of a human author—SCGC and SAC respectfully submit that there would be practical benefit in clarifying, within the Act, that the author of a copyrightable work must be human.  Such a clarification would bring critical guidance and predictability to regulators, courts, creators and consumers.

SCGC and SAC unequivocally reject Option 2: the objective of the Act has always been to protect human creation, not to facilitate artificial creation. The Copyright Act is meant to evolve and adapt as technology and business models change. Option 2 would turn the Copyright Act into an instrument that disenfranchises rather than protects human creators.

Accordingly, SCGC and SAC support Option 1, and as noted, unequivocally reject Option 2.

We respectfully encourage the Government of Canada to be guided in its determinations on this question by the Principle 5 of the Human Artistry Campaign Copyright should only protect the unique value of human intellectual creativity.

Copyright protection exists to help incentivize and reward human creativity, skill, labor, and judgment - not output solely created and generated by machines.

Human creators, whether they use traditional tools or express their creativity using computers, are the foundation of the creative industries and we must ensure that human creators are paid for their work.

Infringement and Liability regarding AI

2.1.4. Should there be greater clarity on where liability lies when AI-generated works infringe existing copyright-protected works?

As generative AI programs can be manipulated and adapted for uses not originally intended, the role of those who deploy the system must be considered, not just that of the initial AI programmer. 

The Copyright Act is already clear on where liability lies for those who engage in unlicensed use or distribution of copyrighted material.

Plagiarism is plagiarism, whether the plagiarizer uses a pen, a typewriter, or a sophisticated computer program to reproduce copyrighted material without credit, consent and compensation.

Comments and Suggestions

2.1.5. Are there approaches in other jurisdictions that could inform a Canadian consideration of this issue?

The Screen Composers Guild of Canada (SCGC) and the Songwriters Association of Canada (SAC) respectfully encourage the Government of Canada to be guided by transparency obligations in the European Union’s AI Act, which require that AI-based systems must be transparent in their functioning so that users can understand how decisions are taken and the logic behind them. This includes providing an explanation of how an AI system arrived at its decisions, as well as information on the data used to train the system and the accuracy of the system.

Similarly, transparency obligations in the US Department of State and Commerce AI Guiding Principles should serve as a model for Canada. In particular Principle 11, which requires AI developers to implement appropriate safeguards before and throughout training, on the use of personal data, and material protected by intellectual property rights, including copyright-protected content.

Principle 11 further states that there should be appropriate levels of transparency on the use of datasets ingested in the AI programming process, and these organizations should comply with applicable legal frameworks.

It remains the position of SCGC and SAC that a similar principal should have been included in the ISED’s Voluntary Code of Conduct on the Responsible Development and Management of Advanced Generative AI Systems. 

We respectfully maintain that the Code should have required signatories to include an appropriate range of identifying information in the metadata of any work derived from “vast data sets” of copyrighted works ingested for TDM purposes. Absence of such a principle in the Code only exacerbates the ongoing negative impacts of generative AI on Canadian copyright holders.

At a minimum, AI companies operating in Canada should be required, whether by legislation or regulation, to (i) maintain accurate and transparent records of the copyrighted works ingested into their AI programs, (ii) to obtain a consent in the form of a licence before ingesting it, and (iii) to pay compensation in exchange of such use. These measures would address many of the TDM challenges facing rights holders, while keeping Canada in alignment with both the EU and US on these critical questions.

SECUR3D

Technical Evidence

How does your organization access and collect copyright-protected content, and encode it in training datasets?

A: We currently only access and train off of datasets that are wholly owned or licensed, open-source, and/or rights free data. Our organization is focused on visual IP and copyright protection so we try to employ ethical standards when it comes to training data to avoid hypocritical criticisms.

·         How does your organization use training datasets to develop AI systems?

A: Our organization deploys both proprietary software and advanced AI systems to facilitate digital content analysis and protection to our customers. Currently we use ethically trained datasets to analyze and compare 2D textures of 3D models, identify brand infringement, and detect explicit content.

Eventually we will be training new datasets to analyze and compare 3D meshes.

·         In your area of knowledge or organization, what is the involvement of humans in the development of AI systems?

A: Research and Development: Humans are integral to the conceptual and practical development of AI technologies. This includes conducting scientific research, formulating algorithms, designing machine learning models, and developing the underlying computational architecture.

Data Preparation: One of the most significant contributions of humans in AI development is in the realm of data preparation. This involves collecting, cleaning, and organizing data, which AI systems use for learning. The quality and diversity of this data are critical for the effectiveness and impartiality of AI models.

Training and Tuning: AI systems, especially those based on machine learning, require training. Humans are involved in this process by selecting appropriate training data, adjusting parameters (tuning), and refining algorithms based on the performance of the AI system during its training phase.

Ethical and Legal Considerations: Humans are essential in addressing the ethical and legal aspects of AI. This includes ensuring that AI systems are developed and used in a manner that is ethical, respects privacy, avoids bias, and complies with legal standards and regulations.

User Experience Design and Interface Development: The design of user interfaces and the overall user experience for AI systems are crafted by humans. This aspect ensures that AI systems are accessible, user-friendly, and effective in meeting the needs of the users.

Testing and Quality Assurance: Human oversight is crucial in testing AI systems for errors, biases, and performance issues. Quality assurance processes often involve human evaluators who can identify and rectify issues that automated testing might miss.

Integration and Application: Humans are involved in integrating AI systems into existing technological frameworks and in applying AI solutions to practical problems in various industries, such as healthcare, finance, transportation, and more.

Ongoing Monitoring and Maintenance: After deployment, AI systems require continuous monitoring and maintenance. Humans are responsible for overseeing these systems, updating them as needed, and ensuring they continue to function effectively and safely.

Policy Making and Governance: Policymakers and regulators, who are humans, are responsible for creating and enforcing policies and regulations that govern the development and application of AI technologies. This ensures that AI is used responsibly and beneficially in society.

·         How do businesses and consumers use AI systems and AI-assisted and AI-generated content in your area of knowledge, work, or organization?

A: Everyone is using AI in some form or another as a tool to expedite workflows and processes. From very basic rudimentary using ChatGPT for email responses, to using Midjourney to quickly concept out different visual styles and themes, to using Copilot to write and check engineering code. AI systems and tools are enhancing and fundamentally improving work and social lives, but the larger issue is these systems have and continually to be built unethically.

Text and Data Mining

·         What would more clarity around copyright and TDM in Canada mean for the AI industry and the creative industry?

A: Clarity would mean better ability for companies and individuals to understand where exactly the legal line(s) fall and hopefully operate more ethically than the current standard. Clearly defined and substantiated legal ground with equal outlines for repercussions would hopefully curtail bad actor behavior around copyright and TDM. As with any other policy or legislation, if clarity is absent, there is no way to determine what action falls within being lawful, and will skew behaviour as anything is permitted.

·         Are TDM activities being conducted in Canada? Why or why not?

A: TDM activities are certainly being conducted in Canada. While perhaps not as prevalent nor publicized as TDM practices being done by large US-based tech and AI firms, those ethical and unethical practices are definitely being done in Canada. The fact that the Canadian government is actively engaged in a consultation process to understand and address the implications of TDM activities within its copyright framework somewhat answers that question.

·         Are rights holders facing challenges in licensing their works for TDM activities? If so, what is the nature and extent of those challenges?

A: Without regulation around licensing work for AI/TDM purposes, there is no framework to do so, and content is being used without regard for licensing. There isn’t much regard for licensing given that the vast majority of data is readily available online and can be accessed to train AI models without repercussion. The current standard is fundamentally to take whatever you can get your hands on to train models, and face punishment after, if any. This needs to change. 

·         What kind of copyright licenses for TDM activities are available, and do these licenses meet the needs of those conducting TDM activities?

A: Non-Negotiated Licenses: These are commonly found in mass market products, such as software, and often take the form of click-wrap or browse-wrap licenses. In these cases, the user must expressly agree to the terms by clicking a button or checking a box. These licenses are typically non-negotiable and unilateral, meaning they are set by the licensor without room for modification by the user. It is important for users to review sections on "Authorized Uses," "Non-Permitted Uses," and "Intellectual Property" in these agreements, as they often contain clauses relevant to TDM activities​​.

Open and Public Licenses: These are more general licenses under which copyright holders release their works for public use without requiring special permission. A well-known example of this type of license is the Creative Commons licenses, which allow content to be used under specified conditions, making them suitable for certain TDM activities​​.

Legal Considerations and Fair Use: In some jurisdictions, like the United States, fair use provisions may apply to TDM activities, particularly for non-fully open access content where the user is not the copyright holder. However, contractual and licensing agreements can override standard copyright laws and may limit the use of content even in ways that might otherwise be considered fair use​​.

Intellectual Property Laws: Different countries have various intellectual property laws that can affect the use of content for TDM. While some laws allow TDM without explicit permission from rights holders, others require explicit permission. In cases where legal exceptions allow for TDM, aspects of a license or contract may still limit how content can be used for TDM​​.

·         Should there be any obligations on AI developers to keep records of or disclose what copyright-protected content was used in the training of AI systems?

A: Yes, all records should be required to be maintained and to be made available upon challenge to the system’s integrity.

·         Are there TDM approaches in other jurisdictions that could inform a Canadian consideration of this issue?

 A: Yes, a few examples below.

European Union - General Data Protection Regulation (GDPR): The GDPR is one of the most influential data protection regulations globally. It emphasizes data privacy and gives individuals significant control over their personal data. Canada could look at how GDPR balances data utility and privacy.

United States - Health Insurance Portability and Accountability Act (HIPAA): HIPAA sets the standard for protecting sensitive patient data in the U.S. While it's specific to healthcare, the principles of data protection and privacy can be informative for broader TDM strategies.

Australia - Consumer Data Right (CDR): Australia's CDR provides consumers with greater control over their data. It mandates that businesses give consumers access to their personal data and the ability to authorize secure access to this data by third parties.

Singapore - Personal Data Protection Act (PDPA): Singapore's PDPA governs the collection, use, and disclosure of personal data. It's an example of a framework that balances data protection with organizational needs and innovation.

Japan - Act on the Protection of Personal Information (APPI): Japan's APPI is another robust personal data protection law. It is known for its unique approach to data management, emphasizing both the protection of personal information and the utilization of personal data.

United Kingdom - Data Protection Act 2018: This act is the UK’s implementation of the GDPR. It contains provisions and exemptions tailored to domestic needs, which could be particularly instructive for a country like Canada, which has its own specific legal and cultural context.

Authorship and Ownership of Works Generated by AI

·         Is the uncertainty surrounding authorship or ownership of AI-assisted and AI-generated works and other subject matter impacting the development and adoption of AI technologies? If so, how?

A: Yes, with no clarity many opt out of AI use, or use AI irresponsibly. Companies face backlash from creator communities when they are caught using AI, as it is viewed as a cheap and unethical replacement for commissioning original work. In terms of development, this is continuing, but it is done so without any ethical concern to ownership and attribution, meaning that models have already been trained on unauthorized data and will already have issues that only continue being further built upon.

·         Should the Government propose any clarification or modification of the copyright ownership and authorship regimes in light of AI-assisted or AI-generated works? If so, how?

A: Yes, through informed policy analysis and policy making, as well as through addressing creator concerns, and legal battles against AI use of copyrighted materials and intellectual property. The stance of the government should be to protect creators’ rights and ensure that their work cannot be infringed upon or used without authorization. A framework and set of standards for responsible AI use should be constructed and applied to models producing content.

·         Are there approaches in other jurisdictions that could inform a Canadian consideration of this issue?

A: Yes, current events show that across the world, governments are moving to learn more about responsible AI use, and implement better policy and means for regulations. MIT provides a good high-level overview of national movement toward AI policy across key players, and where this got throughout 2023 (https://www.technologyreview.com/2024/01/05/1086203/whats-next-ai-regulation-2024/). Similarly, examining the lawsuits currently in progress against generative AI models such as those created by OpenAI and Midjourney are great points of analysis when it comes to infringement that is being alleged against these companies.

Infringement and Liability regarding AI

Are there concerns about existing legal tests for demonstrating that an AI-generated work infringes copyright (e.g., AI-generated works including complete reproductions or a substantial part of the works that were used in TDM, licensed or otherwise)?

A: Yes, they are lacking and need much better software and competencies behind them. All current systems created to demonstrate AI use are faulty and don’t accurately pick up on AI use; this is both in terms of text and visual content.

·         What are the barriers to determining whether an AI system accessed or copied a specific copyright-protected content when generating an infringing output?

 A: Simply put, the biggest barrier is there is no technical way to identify whether an AI system accessed or copied copyright-protected content from a generated output. This issue is currently exacerbated in creative industries where more often than not AI is replicating exact style and tone from artists or writers based on unauthorized training of their data. A detailed public log from a company/ organization that has documented their training data would only be part of the puzzle in overcoming this larger barrier.

 ·         When commercializing AI applications, what measures are businesses taking to mitigate risks of liability for infringing AI-generated works?

A: Little to no measures as none are currently explicitly required, though they should be.

·         Should there be greater clarity on where liability lies when AI-generated works infringe existing copyright-protected works?

 A: Yes, without this clarity it is nearly impossible to enforce copyright protection or to identify legal cause for copyright infringement. Creators and IP rights holders deserve to know that their work is protected and that they have the ability to challenge situations where their rights are not respected.

 ·         Are there approaches in other jurisdictions that could inform a Canadian consideration of this issue?

|A: Yes, current events show that across the world, governments are moving to learn more about responsible AI use, and implement better policy and means for regulations. MIT provides a good high-level overview of national movement toward AI policy across key players, and where this got throughout 2023 (https://www.technologyreview.com/2024/01/05/1086203/whats-next-ai-regulation-2024/). Similarly, examining the lawsuits currently in progress against generative AI models such as those created by OpenAI and Midjourney are great points of analysis when it comes to infringement that is being alleged against these companies.

Comments and Suggestions

·         Are there any other considerations or elements you wish to share about copyright policy related to AI?

A: Beyond creating a legal framework that will inform both companies and individuals operating within the AI space, or simply using AI, it’s more important than ever to stimulate support for companies focused on creating systems that can work in parallel to legislation, helping all rights holders and stakeholders of original works safe from IP and copyright theft/ infringement. All over Canada, teams like our own at Secur3D, are facing the problem of IP and copyright infringement head on, and working tirelessly to build complementary solutions. Our team, for example, is working on new software and tools which leverages AI to deliver digital content analysis, moderation, and authentication, at scale. These are capabilities intended to streamline the protection and authentication of 3D models across the digital landscape, eliminating IP and copyright infringement before it is able to occur.The reality is, current protection vehicles are antiquated and reactive, and certainly do not take into account the disruptive nature of AI innovation. Without new solutions being able to advance and gain traction, policy, alone, will struggle to meet the speed and volume at which content is being uploaded online. This would render even the best informed framework for IP and copyright protection fairly ill functioning. In the same way that the government has provided specialized support to various sectors when development was necessary, it would be extremely beneficial to approach copyright protection in the same way, and consider the businesses and individuals already working toward critical improvements in the space.

Simon Fraser University

Technical Evidence

Numerous researchers at Simon Fraser University (SFU), across a variety of disciplines, are engaged in multidisciplinary research in artificial intelligence to help create innovative, equitable and novel solutions to the challenges facing contemporary Canadian society across diverse sectors (see https://www.sfu.ca/big-data/using-data/artificial-intelligence-at-sfu.html). Researchers at SFU develop AI systems for a variety of applications, from improved medical diagnostics and decision making to enhanced language training. These researchers use licensed datasets for training AI models, and also create their own training datasets when necessary. When creating their own datasets SFU researchers follow relevant copyright policies as well as research and ethics protocols to ensure data and privacy protection.

SFU’s Copyright Office assists researchers in understanding the copyright implications of their AI research projects. Researchers want to responsibly develop AI tools, which includes ensuring that copyright is considered and respected. Many generative AI applications being developed by SFU researchers do not involve copyright protected works – for example those involved with medical imaging and data pattern correlations. SFU researchers and copyright specialists review dataset licenses to ensure that use of the dataset respects license terms. When potential training dataset content is not licensable the SFU Copyright Office or external experts provide guidance to researchers regarding the application of fair dealing to their curation and use of the contents well as their use of generative AI systems to generate outputs.

SFU is interested in not only mitigating the risk of copyright infringement, but also ensuring transparency and non-bias in training data. Many researchers are concerned about the potential for generated outputs to be infringing or cause other harms and therefore some SFU researchers are working on solutions to attribute, or link, training data to the generated works to provide greater transparency to the user. To do so requires using datasets that come with sufficient metadata describing the ownership of each piece of copyrighted content such that each content element is identifiable. SFU researchers, particularly graduate students, who use generative AI to create new works as part of their research are often very aware of terms of use of different AI systems and tend to use only those systems where they can use the model locally, without contributing their copyrighted works to the larger training set. Transparency requirements for high-risk AI applications (e.g., biometrics, medical devices) and codes of conduct for lower risk AI applications (e.g., image generation systems) in the proposed EU AI Act (European Commission, 2021; see: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=celex%3A52021PC0206) offer a model for possible applicability in the Canadian context. As well, ethical transparency should also ensure that items of knowledge or cultural significance to Indigenous communities are not included in training datasets without consent and participation of those same communities.

While SFU researchers are very involved in the design and training of generative AI models, this does not mean that these same researchers can necessarily claim copyright in the output of their models. SFU advises its researchers that the outputs of their models are likely not protected by copyright. This is because outputs from generative AI systems are not directly attributable to the model’s developers, but instead the outputs are based on deterministic statistical probabilities (Lee et al., 2023; see https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4523551). Consequently, the developers responsible for the AI model are unable to claim any exercise of skill and judgement in each output of a generative AI system. Since an exercise of skill and judgement is a requirement for the adherence of copyright to a work,  the developer cannot claim copyright in the AI model’s outputs (see: CCH Canadian Ltd. v. Law Society of Upper Canada, 2004. Para 16.).

In addition to developing AI systems for a variety of purposes, SFU researchers, educators and students also use AI systems and AI generated content. Students use AI systems such as ChatGPT or Stable Diffusion to assist in the creation of assignments. Staff and faculty use generative AI tools to assist them in developing ideas and content. Note that they are not using AI systems to create their work, but rather using them as a starting point or an aid in creating new works. Educators use AI systems to help them develop student assessment tools such as quizzes – for example, an instructor can feed a chapter from a book into an AI question generator in order to obtain a short multiple-choice quiz based on the content of the chapter. Digital resources that universities subscribe to likely cannot be used in such ways due to license restrictions, but short extracts from print books could conceivably be used.

Universities also utilize regular and generative AI tools to carry out laborious and time-consuming tasks. For example, university instructors will use AI tools to assess student writing for plagiarism and ensure academic integrity. Libraries and archives are employing generative AI tools to create and enrich metadata for their unique collections. These tools are becoming key to supporting a variety of ways in which the University carries out its mission.

RECOMMENDATIONS

Provide clarity around training dataset content by requiring transparency in training dataset metadata for high-risk AI applications and developing transparency codes of conduct for training dataset metadata for lower-risk AI applications with the aim of having sufficient metadata such that each content element is identifiable.

Facilitate the development of protocols created through collaboration between Indigenous communities and researchers to ensure that Indigenous Knowledge and Traditional Cultural Expressions are not captured in training datasets without consent and participation of the relevant Indigenous communities.

Provide clarity regarding copyright implications of using unlicensed data for training generative AI models (see recommendations in Text and Data Mining section).

Text and Data Mining

TDM as an analytical tool involves a non-consumptive use of a work to reveal trends and relationships in the facts and ideas underlying the work’s expressed form. Generative AI is one example of the application of TDM technology that allows users to engage with the facts and ideas presented in works. By non-consumptive we mean the use of a work in a way that does not involve consumption of a work by a human being. When a digital work is mined for data and facts by a computer, there is no human consumption, or enjoyment, of the work in the way that the creator of the work envisioned the work being used. No one writes a novel so that a computer can analyze it – they write it so that a human can read it and engage with it emotionally and intellectually.

Current legislative ambiguity around the copyright implications of engaging in TDM with copyright protected works hinders SFU from confidently providing access to the information and knowledge to develop TDM tools for application in society. Not only do researchers and students need clarity for TDM uses of works, we also require clear language in the Copyright Act clarifying that a contract cannot override exceptions in the Act in order to halt the erosion of the public domain, and to safeguard against restrictive licensing agreements that override fundamental user rights codified in the Copyright Act and clearly expressed in Supreme Court jurisprudence. See, for example, Alberta (Education) v. Canadian Copyright Licensing Agency (Access Copyright), 2012 SCC 37, [2012] 2 SCR 345. Para 22 & CCH Canadian Ltd. v Law Society of Upper Canada, 2004 SCC 13, [2004] 1 SCR 339. Para 12 as examples of the Supreme Court endorsing the view that exceptions are “user rights.”

Twenty-first century Canadian copyright legislation and jurisprudence have given Canadians a balanced copyright regime. This regime recognizes that a copyright owner cannot have exclusive control over all uses of their work (i.e., the doctrine of exhaustion) but instead confers upon users of copyrighted works certain user rights (see: CCH Canadian Ltd. v. Law Society of Upper Canada, 2004. Paras 12-13 & 54.). This balance must be maintained when considering the copyright implications of TDM. Users must not face additional barriers to engaging with the facts and ideas in copyrighted works just because they are using these works for TDM purposes. It is imperative that universities like SFU be able to foster and encourage, through research, the increasingly complex ways that users interact with the ideas and facts contained in corpora of copyrighted works. However, it is important for SFU to make clear that we specifically support applications of non-consumptive use which do not encroach on the original expression of the work by generating copies of existing works. That is, TDM applications that extract and understand the patterns, information and correlations – essentially the facts and ideas behind these works – to generate, or assist in the generation of, new and different works.

Many jurisdictions have enacted provisions in their copyright legislation to support TDM, or have existing systems that provide the flexibility to implement TDM. For example, the USA’s flexible and illustrative fair use provision provides a solid legal basis for the non-consumptive use of copyrighted works. Canada’s fair dealing exception lacks the flexibility of US fair use. Although SFU favours expanding Canada’s fair dealing exception to be illustrative and flexible, SFU and all research institutions would benefit from an explicit exception for non-consumptive uses of copyrighted works along the lines of Japan and Singapore. Japan’s exception has been described as one that “comprehensively allows an exploitation of a work by any means to the extent deemed necessary, if the exploitation is aimed at neither enjoying nor causing another person to enjoy the work, unless such exploitation unreasonably prejudices the interests of the copyright holder” (Ueno, 2021. See https://doi.org/10.1093/grurint/ikaa184).  A similar exception in Canada would meet the desire for a technologically neutral Copyright Act and would ensure that future non-consumptive TDM-like processes would also be covered. Acknowledging the separation of the economic interests vested in the original expression and the user rights to the information behind the work further supports an exception for both commercial and non-commercial applications.

RECOMMENDATIONS

Create a specific exception to infringement that would allow for the copying and use of a work or other subject-matter for the purposes of informational analysis and related purposes. This aligns with recommendation 23 from the 2019 Statutory Review of the Copyright Act (INDU, 2019). A number of Canada’s key trading partners already have a specific exception for TDM, including Japan, Singapore, the United Kingdom and the EU. SFU supports an exception that applies to both commercial and non-commercial research, and which includes both the reproduction right and the communication right.

Introduce an provision in the Copyright Act that prevents contracts from overriding copyright exceptions for non-infringing purposes. This provision should apply to all future and pre-existing contracts. Moreover, this exception would equally apply not only to Canadian law-governed contracts, but also contracts governed by foreign law to avoid situations where the choice of a foreign contract is made to evade the contract override exception. For example, see Singapore Copyright Act 2021, s 188 (https://sso.agc.gov.sg/Acts-Supp/22-2021/Published/).

Allow circumvention of Technological Protection Measures (TPMs) for any non-infringing purpose. This would make the users’ rights in the Copyright Act technologically neutral and contribute to restoring the balance between copyright owners’ rights and users’ rights.

Follow the recommendation of the Standing Committee on Industry, Science and Technology in the 2019 Copyright Act Review and make fair dealing illustrative by adopting an illustrative, rather than exhaustive, list of purposes by including the words “such as” before the list of purposes in s 29. An illustrative list would provide the flexibility to carry out AI related research in a legally secure manner.

Authorship and Ownership of Works Generated by AI

At SFU, the ongoing research and development of generative AI systems is not impacted by uncertainties around the copyright ownership in outputs of AI systems. This is either because the outputs are not copyrightable – such as medical image enhancement or large scale data analysis – or because in many cases the developers are also the ones generating the output as in the case of musical generative AI research projects. The development work is going on irrespective of concerns about copyright ownership of outputs. However, this is for now. According to a Bloomberg report 2023 Generative AI Growth, generative AI is poised to become a $1.3 trillion dollar (USD) market by 2032. Technology with such a large financial impact requires certainty around its various copyright elements. It is clear that the development of an AI algorithm and model is a work of skill and judgement and is therefore protected by copyright as it incorporates an incentive to create – one of the main functions of granting copyright protection. However, as research projects move out of the lab and into the public market the question of who owns the output from generative AI systems must be addressed.

Copyright in Canada protects the expression of human creativity that involves the exercise of skill and judgement. Generative AI makes us consider the characteristics of originality, creativity, data and computation. The outputs of generative AI systems are the result of statistical and routine processes and therefore may not reach the originality bar set out in CCH (see: CCH Canadian Ltd. v. Law Society of Upper Canada). While outputs are created based on human prompts – either textual or visual – it is the complexity of these prompts that will determine if there is any copyright protection to the outputs. In situations where the complexity of the human prompts is considered an exercise of skill and judgement, then the human user of the AI system would own some copyright in the output. However, substantial amounts of the output would still be purely the product of the AI system and should not be protected by copyright.

Copyright legislation also strives to “maintain a balance between the rights of authors and the larger public interest, particularly education, research and access to information” (WIPO Copyright Treaty, 1996). AI processes create works in a faster and more systematic way than human authors. The mass output that AI makes possible can cause an economic reordering – disadvantaging human authors and privileging machine outputs over human creations (Abbott, 2022; see: https://doi.org/10.4337/9781800881907). If protected by the full range of copyright protections, this kind of volume-based output will crowd out human authors and enclose the space of the public domain. 

SFU agrees with the stance taken by the US Copyright Office regarding copyright protection for AI generated works and when and how copyright protection may adhere to works partially generated by generative AI. “If a work’s traditional elements of authorship were produced by a machine, the work lacks human authorship” and is therefore not protected by copyright (see: US Copyright Office, Copyright Registration Guidance for Works Containing AI-Generated Materials. Federal Register, 88 Fed. Reg. 16,190). However, “a human may select or arrange AI-generated material in a sufficiently creative way that ‘the resulting work as a whole constitutes an original work of authorship.’ Or an artist may modify material originally generated by AI technology to such a degree that the modifications meet the standard for copyright protection” (US Copyright Office). 

That is, if a work is solely the output of a generative AI system, SFU supports the position that the work should remain in the public domain. Providing full copyright protection for generative AI outputs with no human creativity involved in their creation threatens copyright’s balance and the value Canada places on human expression as Carys Craig and others have warned (Craig, 2014. See https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3733958). However, if a human selects and arranges AI-generated material in a sufficiently creative way, then those specific creative elements are deserving of copyright protection.

RECOMMENDATIONS

AI authored works that are produced solely by an AI system should not be protected by copyright. But, where human creativity is involved in the creation or adaptation of an AI generated work, then copyright should adhere to those elements that are the product of human skill and judgement.

The Canadian Intellectual Property Office should not accept copyright registrations for works solely created by AI systems. Accepting such works for registration results in confusion in the research community, and in the general public, as it goes against the generally accepted requirements for copyright protection in Canada.

Infringement and Liability regarding AI

In Canada it is the courts who determine if copyright infringement has occurred, and this function is best left for the courts when it comes to generative AI outputs. The courts, on a case by case basis, are well placed to determine if generative AI outputs are reproductions of a substantial part of a work or adaptations of a work, and to determine who would be considered liable (the company providing the service, the programmer or the user). However, there is a problem with identifying the content of datasets. This leads to difficulties in applying the existing tests for copyright infringement since without transparency around the dataset’s content how can a court know if the work is protected by copyright, or even if the work was copied, or if there is a causal connection between a generated output and an original work within the training dataset? Currently it is difficult for a claimant to prove that they are the copyright owner of a work, particularly when attempting to prove infringement in an AI generated work. For example, in the ongoing US litigation Andersen v Stability AI the parties have been told to sort out during discovery which images were part of the training set as it was not immediately obvious if the claimant’s works were in the training dataset (see: https://www.findlaw.com/legalblogs/federal-courts/judge-trims-copyright-lawsuit-against-ai-model-stable-diffusion/). A lack of transparency when it comes to training data is an obstacle for the discovery of whether or not a non-consumptive copy of a specific copyright protected work was used in the AI-training process and subsequently substantially reproduced in an output.

This transparency problem demonstrates the need for developers of AI systems to ensure that their training datasets contain sufficient metadata to identify every work in the content set to promote transparency and to guarantee that rights holders can properly exercise their rights. Along with the need for rich metadata is the need for transparency in how the AI model is trained, and in the intentions and designs behind the algorithm. When AI is operating, only the humans involved in the design of the algorithms for that specific application can be responsible for the AI model. Bias and dominant modes of thought can easily find their way into selection for training datasets used to train AI systems. Therefore, underlying all AI systems must be ethical considerations that promote transparency and trustworthiness (Mehan, 2022) amongst regulators and the general public so that they understand how and why an AI model renders a certain output.  As a public institution committed to the public good, SFU strongly supports ethically centred AI that is developed with the intention of ensuring marginalized communities are not disrespected and again shunted to the side when it comes to the development of AI systems.

Some private generative AI companies that used creative copyright-protected works to train their machines, such as Stability AI, Microsoft and Google, have already taken steps to create tools that allow for creators to opt-out of the inclusion of their work in the companies’ models going forward. (See the following for further information: Stability AI - https://arstechnica.com/information-technology/2022/12/stability-ai-plans-to-let-artists-opt-out-of-stable-diffusion-3-image-training/; Microsoft - https://blogs.bing.com/webmaster/september-2023/Announcing-new-options-for-webmasters-to-control-usage-of-their-content-in-Bing-Chat; Google - https://blog.google/technology/ai/an-update-on-web-publisher-controls/).

It is also possible for creators themselves to implement technological protection measures to prevent their works from being scraped and included in AI training datasets. For example, the Overlai AI photo protection app due for release in December 2023 uses distributed ledger technology (DLT), decentralized storage and advanced watermarks to enable creators to protect their photographs from being added to AI training datasets (https://www.overlai.app).The ability to opt out of contributing to training data however should remain a private ordering matter and not be legislated. Legislating TDM in a way that allows opt-outs could have a number of significant unintended consequences. While attempting to protect creative industries, legislating opt-outs could lead to serious long-term effects on the future reliability of AI machines in high-risk applications such as biometrics, health care or autonomous vehicles where the broadest and most inclusive training datasets are required. For example, see Philipp Hacker’s “A legal framework for AI training data—from first principles to the Artificial Intelligence Act” (2021) for an explanation of the discrimination risks in biased (non-inclusive) training datasets (see: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3556598).

AI authored works that infringe copyright should be removed from distribution and any circulation of the outputs ceased. When addressing where copyright liability lies when a generative AI output is found to be infringing, SFU believes this is best left to the courts on a case by case basis. This is because the liability may well lie with the AI developer, the AI company, the user or along a continuum of all three. For example, an AI developer or AI company may be guilty of infringement if a defect was present upon the AI’s release that allowed for substantial reproduction of existing works in generated outputs. Thus determining liability is best left with the courts.

RECOMMENDATIONS

Retain existing distinctions between commercial and non-commercial statutory damages in section 38.1(1) of the Copyright Act. As noted above, liability is context dependent. Society and the judicial system must always consider the underlying purpose that led to the infringement; infringement that happens as part of a research project at a university is vastly different than infringement that happens as part of a profit-oriented motive.

Comments and Suggestions

As noted earlier, AI possesses the capacity to revolutionize many occupations and alter the work of creators. Similar disruptive innovations, or technologies, such as printing presses, industrial automation, automobiles and the internet, are seen throughout human history. Addressing the resultant innovative disruption by supporting training for new opportunities in jobs related to AI development or by supporting worker retraining through organizations like community colleges, universities and public libraries, should be approached at an economic and society wide level (see: https://www.librarycopyrightalliance.org/wp-content/uploads/2023/06/AI-principles.pdf). Since these disruptions need to be addressed at a societal level, the Copyright Act is not the tool to use to attempt to remediate the effects of these disruptions – for example by implementing a mandatory licensing scheme for AI training datasets. Nor should AI innovation be constrained in Canada by implementing constraining copyright laws with fewer exceptions than other competing jurisdictions, such as the US and Japan, who have more AI innovation friendly copyright exceptions.

Société des auteurs et compositeurs dramatiques – société civile des auteurs multimédia

Preuve de nature technique

N/A

Fouille de textes et de données

La SACD-SCAM est d'avis qu'aucune modification ne doit être apportée à la LDA afin de permettre les activités de fouille de textes et de données. Premièrement, tout comme nul ne songerait à exproprier les propriétaires d'intrants tangibles afin de satisfaire les besoins d'une industrie naissante requérant ces intrants pour mener à bien ses opérations, le gouvernement doit aussi cesser d'entretenir le réflexe d'exproprier les auteurs de leurs droits afin de répondre aux besoins des entreprises ayant besoin de leurs œuvres pour la conduite de leurs opérations.  Dans la mesure où l'IA a besoin des œuvres des auteurs, nourrir l'IA au détriment des auteurs ne peut qu'ultimement mener à l'appauvrissement de ces derniers, de leur création et ultimement de l'IA elle-même. Deuxièmement, les droits d'auteur visent à permettre aux titulaires de droit d'auteur sur des œuvres d'en autoriser ou interdire l'utilisation et donc de négocier librement les conditions auxquelles ils consentent. Il n'y a pas lieu de créer une exception alors qu'il existe déjà un système via des sociétés comme la SACD-SCAM pour obtenir l'accord des ayants droit. Troisièmement, les traités internationaux en matière de droit d'auteur  auquel le Canada est partie imposent le respect du test en trois étapes. L'adoption d'une exception permettant la FTD serait contraire aux engagements du Canada. La SACD-SCAM est donc d'avis que rien ne justifie la création de nouvelles limitations ou exceptions visant la FTD au Canada, du moins à l'égard des œuvres cinématographiques et dramatiques.

Titularité et propriété des œuvres produites par l’IA

La SACD-SCAM invite le gouvernement à faire preuve de la plus grande prudence sur ces questions, les choix retenus pouvant entraîner des répercussions profondes et à long terme pour les auteurs, pour le développement de la culture, elle-même largement nourrie de créations humaines et pour la cohésion de la société dans la mesure où celle-ci s'appuie largement sur les référents culturels que partagent ses membres.

Conférer des droits d'auteur aux productions générées par l'IA qui s'apparentent aux oeuvres de l'esprit pourrait entraîner un chômage technologique chez les créateurs en comparaison duquel celui vécu par les artistes interprètes de l'entre-deux-guerres par suite du perfectionnement et de la démocratisation des technologies de fixation, reproduction et télécommunication des sons et des images pourrait ne faire que pâle figure, surtout comme on semble l'anticiper, l'IA sera en mesure de créer des oeuvres en mesure de concurrencer celles des humains et ce, de façon rapide, massive et à peu de frais.

La SACD-SCAM recommande donc de ne pas modifier la LDA sur cette question et de laisser se poursuivre les réflexions actuellement en cours à ce sujet au sein d'organisations telle que l'OMPI.

Violation et responsabilité en matière d’IA

Le développement de l'IA dans la sphère de création soulève des défis majeurs qui doivent être considérés avec toute l'importance que celui-ci exige.

L'état d'avancement de l'IA et, surtout des réflexions sur ses conséquences en matière de droit d'auteur et, plus largement, de culture et de société, conduisent la SACD-SCAM à recommander au gouvernement la plus grande des prudences quant aux décisions qu'il pourra prendre relativement à l'identification des auteurs dont la création est assistée par l'IA et surtout, de l'attribution de droit sur les productions générées par l'IA sans apport original humain.

Commentaires et suggestions

Que la LDA, malgré les changements incontournables dus aux évolutions technologiques, dont l'IA, conserve sa pertinence et soit en mesure de relever les défis soulevés par ces évolutions.

La SACD-SCAM rappelle que l'adaptation de la LDA doit d'abord viser à permettre aux auteurs de pouvoir continuer à contrôler l'exploitation de leurs œuvres de façon effective dans le contexte des évolutions technologiques et non de réduire, lors de chaque révision, la portée de leurs droits afin de satisfaire les besoins des protagonistes de ces mêmes évolutions technologiques.

Société professionnelle des auteurs et des compositeurs du Québec

Preuve de nature technique

Plusieurs de nos membres utilisent l’IA générative comme outil de création. Nos membres estiment que l’IA générative est et doit demeurer un outil, en ce sens que la créativité humaine doit toujours primer.

Fouille de textes et de données

- Une plus grande clarté permettrait de mieux appréhender le fonctionnement de la FTD à titre de processus technologique, incluant la façon dont les œuvres et autres objets de droit d’auteur sont utilisés.

- Oui, des activités de FTD sont actuellement menées au Canada, afin d’entraîner des modèles algorithmiques. Les activités de développement et d’entraînement de systèmes d’IA sont susceptibles d’impliquer la reproduction de contenus protégés par droit d’auteur (œuvres et autres objets de droit d’auteur tels que des prestations), sans que les titulaires de droits y consentent et reçoivent une juste rétribution. Ceci est évidemment problématique et il importe d’y remédier, par exemple, via l’imposition d’une obligation de transparence.

- Oui. En outre, il est difficile pour les titulaires de droits d’auteur de déterminer quel contenu est utilisé dans le contexte de FTD et quelle est l’ampleur de cette utilisation. Afin de pallier cette lacune, il pourrait être envisagé d’imposer une obligation de transparence auprès des entités développant et entraînant des systèmes d’IA. Cette dernière avenue ne devrait pas entraîner de difficultés particulières, puisque les développeurs et chercheurs du secteur de l’IA générative documentent déjà leurs données d’entraînement, par exemple, par le biais de fiches de données ou « model cards ». Les « model cards » peuvent documenter des informations structurées tels que les noms de domaines où ont été collectés les données d’entraînement. Par exemple, la « model cards » de l’IA GPT-2 d’OpenAI (publiée en 2019) incluait une liste de 1000 des noms de domaine ayant servi de source, ainsi que le nombre de références par nom de domaine. Dans cette liste, on pouvait retrouver des sites illégaux (Pirate Bay), pornographiques (YouPorn) ou d’ayants-droits (Le Monde). En utilisant ces mécanismes, les ayants droit pourraient disposer d’informations essentielles à la gestion de leurs droits d’auteur.

- Diverses licences sont disponibles pour les activités de FTD impliquant l’exercice d’un droit réservé aux titulaires de droit d’auteur. Ces licences peuvent être négociées de gré à gré avec les titulaires de droits d’auteur ou être obtenues par le biais d’une société de gestion collective. Ces licences ne semblent toutefois pas être obtenues par les personnes menant des activités de FTD, en dépit du cadre législatif actuel pourtant clair et des licences disponibles. Ceci crée évidemment un manque à gagner pour les titulaires de droits d’auteur qui peinent à obtenir une juste compensation pour l’utilisation de leurs contenus.

- Nous ne sommes pas favorables à l’adoption d’une exception générale permettant la FTD, laquelle serait prématurée et contraire aux engagements du Canada en vertu de divers traités internationaux, tels que la Convention de Berne, l’ADPIC et l’ACEUM lesquels précisent que toute limitation ou exception à laquelle le Canada entend assujettir un droit d’auteur doit être restreinte à certains cas spéciaux où il n'est pas porté atteinte à l’exploitation normale de l’œuvre, ni causé de préjudice injustifié aux intérêts légitimes de l'auteur.

- Oui, nous le recommandons. Comme exposé ci-dessus, ces développeurs disposent déjà d’outils permettant de documenter les données d’entraînement. L’introduction d’une obligation de transparence ne devrait donc pas entraîner de coûts additionnels pour l’industrie de l’IA.

- Le niveau de redevance doit être déterminé par le marché, la Commission du droit d’auteur ou les tribunaux.

- Comme exposé plus tôt, il n’est pas recommandé d’introduire une exception de FTD au Canada, car ceci est prématuré en plus de nuire aux intérêts des créateurs.

Titularité et propriété des œuvres produites par l’IA

- Nous ne sommes pas au courant de telles incidences.

- Nous ne recommandons pas de protéger des créations « artificielles » dépourvues de créativité humaine. L’IA doit demeurer un outil au service de la créativité humaine : le récent sondage mené auprès de nos membres est clair sur ce point. Le gouvernement devrait suivre attentivement les développements jurisprudentiels en matière d’IA générative et de droit d’auteur et décider si, à la lumière de ces décisions, des amendements devraient être apportés à la Loi sur le droit d’auteur.

- Le Royaume-Uni, l’Irlande et la Nouvelle-Zélande sont souvent cités en exemple sur cette question. Les législations de droit d’auteur de ces pays attribuent en effet la titularité d'œuvres générées par ordinateur à la personne qui a pris les dispositions nécessaires à la création de l’œuvre créée. Nous ne recommandons toutefois pas d’emprunter cette voie, car ces dispositions ont été introduites dans un contexte étranger à l’IA générative. Or, cette technologie soulève des questions bien plus complexes. Au surplus, avant même d’adresser la question de la titularité des « œuvres générées par ordinateur », il convient de statuer sur leur protection.

Violation et responsabilité en matière d’IA

- Oui. Il peut être difficile pour un titulaire de droit d'auteur :

(i) D’identifier le contenu contrefait, ainsi que la ou les personnes responsables de la violation ; et

(ii)  D’établir que la partie qui a violé le droit d'auteur a eu accès à l'œuvre originale, que l'œuvre originale était la source de la copie et qu’une partie importante de l'œuvre a été reproduite.

- La pluralité des intervenants, l’incertitude juridique, le manque de transparence quant aux systèmes de gestion des données, ainsi que l’opacité des systèmes d’IA. Pourtant, nous comprenons que les développeurs et chercheurs du secteur de l’IA documentent leurs données d’entraînement : une plus grande transparence sur ces données auprès des ayants droit est donc techniquement faisable.

- Certaines entreprises obtiennent des licences, alors que d’autres utilisent du contenu libre de droits, incluant des œuvres tombées dans le domaine public.

- Non, la Loi sur le droit d’auteur dispose de mécanismes suffisants pour déterminer la responsabilité en cas de violation de droit d’auteur. Toutefois, le Canada pourrait imposer une obligation de transparence auprès des entités développant et entraînant des systèmes d’IA.

- Oui. Dans son projet de règlement « AI Act », le Parlement européen a introduit une obligation de transparence, de sorte que les entités qui développent des systèmes d’IA devront publier un résumé suffisamment détaillé de leur utilisation de « données d’entraînement protégées par la législation sur le droit d’auteur », ainsi qu’une information appropriée, claire et visible qui distingue le contenu généré de l’original.

Commentaires et suggestions

La Société professionnelle des auteurs et des compositeurs du Québec (« SPACQ ») est une association qui représente les intérêts moraux, économiques et professionnels des auteurs de chansons francophones à travers le Canada et de tous les compositeurs de musique de commande au Québec.

La SPACQ accueille favorablement la consultation publique et voit en cet exercice une volonté du gouvernement de clarifier les incidences de l’IA générative sur le droit d’auteur.

Nous avons récemment lancé un sondage auprès de nos membres, dans le cadre de cette consultation publique. À la lumière des résultats obtenus, nous constatons que la majorité de nos membres ne souhaitent pas freiner l’avancement de l’IA, mais désirent préserver l’équilibre que la Loi sur le droit d’auteur (la « Loi ») sous-tend, en veillant à ce que les intérêts des auteurs et des titulaires de droits d’auteur soient préservés. En effet, notre association voit le potentiel de l’IA : cette technologie, si elle est adéquatement encadrée, pourrait alimenter la créativité, favoriser la découvrabilité de certains contenus et outiller les créateurs dans le respect de leurs droits.

Il est néanmoins essentiel de prendre conscience des impacts négatifs que l’IA peut avoir sur les droits des auteurs. Afin de freiner ces risques, notre principale recommandation est de veiller au respect de la Loi sur le droit d’auteur en évitant l’introduction de toute nouvelle exception à des fins de fouille de textes et de données (« FTD »). Nous recommandons également qu’une obligation de transparence ou de tenue de registre soit imposées aux chercheurs et développeurs de systèmes d’IA générative, dans le contexte de la FTD. Cette obligation ne devrait pas constituer un fardeau particulier, puisque nous comprenons que ces acteurs documentent déjà leurs données d’entraînement.

La consultation publique est accueillie favorablement par notre association, qui voit en cet exercice une volonté du gouvernement de clarifier les incidences de l’IA sur le droit d’auteur. Notre association ne souhaite pas freiner l’avancement de l’IA, mais désire préserver l’équilibre que la Loi sur le droit d’auteur sous-tend, en veillant à ce que les intérêts des auteurs et des titulaires de droits d’auteur soient préservés. En effet, notre association voit le potentiel de l’IA : cette technologie, si elle est adéquatement encadrée, pourrait alimenter la créativité, favoriser la découvrabilité de certains contenus et outiller les créateurs dans le respect de leurs droits. Il est néanmoins essentiel de prendre conscience des impacts négatifs que l’IA peut avoir sur l’ensemble des secteurs, les fondements de notre société, ainsi que sur les droits des auteurs. Afin de freiner ces risques, notre principale recommandation est de veiller au respect de la Loi sur le droit d’auteur en évitant l’introduction de toute nouvelle exception à des fins de fouille de textes et de données (« FTD »). Nous recommandons également l’imposition d’une obligation de transparence auprès des utilisateurs. Spécifiquement, ce cadre devrait obliger la divulgation de toute œuvre utilisée dans le contexte de l’IA. Un tel mécanisme est une action faisable, qui ne pose pas de difficultés techniques et qui jetterait les premières bases de l’édifice, afin d’assurer une rémunération juste et équitable aux auteurs et titulaires de droits d’auteur.

Nous soutenons, par ailleurs, les commentaires soumis par la Coalition pour la diversité des expressions culturelles (CDEC), la SCGC et la SAC.

Society of Composers, Authors and Music Publishers of Canada (SOCAN)

Technical Evidence

1. Introduction.

SOCAN (Society of Composers, Authors and Music Publishers of Canada) is Canada’s largest rights management organization. SOCAN has over 185,000 songwriter, composer, and music publisher members, and licenses tens of thousands of businesses and organizations across Canada. SOCAN issues licences for the performing rights and reproduction rights in musical works. SOCAN collects and distributes royalties to its members and connects more than 4 million creators and publishers worldwide through international rights management organizations with which it has reciprocal agreements. In 2023, SOCAN licensed over 170 billion individual performances of musical works to its licensees.

SOCAN believes that, with proper safeguards and an appropriate copyright framework, artificial intelligence (AI) can support and enhance human creativity in the music industry. Indeed, in its response to the 2021 “Consultation on a Modern Copyright Framework for Artificial Intelligence and the Internet of Things”, SOCAN noted that the full potential of AI, as well as its implications, were only starting to be uncovered.

Since then, however, there have been monumental developments in the field of AI, most notably the rapid development and adoption of generative AI models. While these developments showcase the possibilities of generative AI, they also lay bare the risks that it poses to songwriters, composers, and other creators if left unchecked. Those risks include (i) the use by AI companies of massive amounts of copyright-protected music to program their models, without permission from or payment to rights holders; (ii) AI-generated works that imitate or reproduce those copyright-protected works or substantial portions of them, which not only threaten the livelihoods of Canadian creators and their ability to continue to pursue careers in music, but also risk destroying a nascent market for the licensing of musical works to AI companies before it has a chance to mature; and (iii) a complete lack of transparency by AI companies, which makes it extremely difficult, if not impossible, for rights holders to know whether AI companies have used their music without permission and, if so, to pursue meaningful enforcement measures.

With appropriate transparency, it will be feasible and practical to license the use of music by AI companies, just as SOCAN and other rights holders have done and continue to do for other innovative technologies. The licensing market that is currently developing should be encouraged and promoted. It should not be eliminated either by sanctioning the large-scale unauthorized use of music or by introducing copyright exceptions that would afford preferential treatment to text and data mining activities at the expense of creators.

SOCAN is well-positioned to help develop a licensing scheme for AI. It has operated its public performance business for almost 100 years, licensing virtually all musical works under a blanket licence for use in various new technologies, including Internet streaming. Creating a licensing scheme to cover the musical works that may be used by AI developers, for programming or other purposes, is well within SOCAN’s expertise.

SOCAN urges the Government of Canada, when considering a copyright policy framework for generative AI, to ensure that creators and copyright are respected, and that human expression is incentivized. The preamble to the Copyright Modernization Act, SC 2012, c 20, emphasizes that the Copyright Act supports creativity, culture, and innovation. To promote those values, it is imperative that creators be able to control, and be paid for, the valuable use of their music by AI companies.

2. How does your organization access and collect copyright-protected content, and encode it in training datasets?

Based on public reports, SOCAN understands that several major generative AI models have been programmed on vast numbers of copyright-protected works, obtained either from large-scale text and data mining (TDM) activities or from datasets containing unlicensed works.

For example, from the Music Publishers Canada (MPC) submission, it has been reported that Google’s T5 model and Meta’s LLaMA model were developed using a dataset containing protected content that was scraped from the Internet, including from scribd.com, a subscription-only digital library, and another website that is notorious for e-book piracy [https://www.washingtonpost.com/technology/interactive/2023/ai-chatbot-learning/]. OpenAI, the company behind the leading generative AI model supporting ChatGPT, has acknowledged the use of “large, publicly available datasets that include copyrighted works.” [https://www.uspto.gov/sites/default/files/documents/OpenAI_RFC-84-FR-58141.pdf].

That said, SOCAN is also aware of reports from the MPC submission of generative AI models that have been programmed using licensed content. For example, Meta announced that its AI-powered music generator tool, MusicGen, was developed on “20,000 hours of music, including 10,000 ‘high-quality’ licensed music tracks” and instrumental tracks from stock media libraries [https://techcrunch.com/2023/06/12/meta-open-sources-an-ai-powered-music-generator].

While licensed uses, to date, appear to have used a smaller scale of data than foundational large language models, they nonetheless demonstrate that a market for the licensing of music to AI companies currently exists, and is feasible and practical. SOCAN is well-positioned to foster the development of that market.

3. What measures are taken to mitigate liability risks regarding AI-generated content infringing existing copyright-protected works?

SOCAN is not currently aware of any measures taken, by either AI developers or data providers, to mitigate the risk of liability for infringing existing protected works. In any event, SOCAN believes that it would be a mistake to focus on how to “mitigate” liability. A focus on mitigation tends to frame the issue in a way that fails to respect creators and tacitly accepts AI companies’ non-compliance with established copyright laws and policy. The focus should be on fostering the development of a licensing market and encouraging AI developers to obtain permission before using creators’ works to program their AI models or for other purposes. If that permission is not obtained, then AI developers should be subject to copyright infringement claims like any other user.

Text and Data Mining

1. Licensing is workable, practical, and necessary to follow appropriate copyright principles.

As Canada’s largest and most active licensor of musical works, SOCAN firmly believes that the licensing of musical works to AI companies for TDM or other purposes is not only workable and practical, but the best and most appropriate way to build on the foundational purpose of copyright.

SOCAN is also confident that a licensing market, which is already developing, will flourish in Canada, if given the opportunity to do so. SOCAN therefore urges the Government not to enact any new or modified exceptions for TDM, which would destroy this market and prevent creators from being compensated for valuable uses of their work.

Many of SOCAN’s songwriter and composer members depend entirely on copyright, and their ability to control and be paid for the use of their works, for their livelihoods. Since many members are not also recording or performing artists, they do not have the opportunity to generate income from touring, sponsorships, merchandise, or other such projects.

AI developers benefit from the use of high-quality protected works, including musical works, for the programming of their AI models.

There is no justification for an AI developer to reap the full benefits of a creator’s labour without permission and without providing any remuneration to the creator. That would be contrary to the objectives of the Copyright Act, which includes securing a just reward for the creator and preventing “someone other than the creator from appropriating whatever benefits may be generated.” [Théberge v Galerie d’Art du Petit Champlain inc, 2002 SCC 34 at para 30 <https://canlii.ca/t/51tn#par30>]. Those objectives are especially important because of the unique risks that generative AI models pose to human creators. After using massive amounts of creators’ works, generative AI models can generate outputs that will compete with the works of those very creators. That creates a serious risk that human songwriters and creators will be displaced and forced to pursue other careers, leading inevitably to an erosion of Canadian culture and the industries that support it.

A licensing model for AI will ensure that songwriters, composers, and their music publishers are able to control and be paid for the use of their works by AI companies, in accordance with Canadian copyright law and policy. SOCAN firmly believes that such a licensing model is feasible and practical, regardless of the number of works or the nature of the technology involved. SOCAN has consistently adapted its licensing processes to respond to major technological developments and market disruptions over the years, including the shift to streaming and digital musical consumption. There is no reason for the emergence of generative AI technology to be treated any differently.

SOCAN has the experience and tools necessary to license and administer large catalogues of works for a variety of purposes, including AI-related uses. As already noted, SOCAN connects billions of performances in Canada with millions of rightsholders worldwide. SOCAN grants licences to tens of thousands of users across a wide spectrum of industries and activities. That includes licensing millions of musical works under blanket licences for use in new technologies, such as online music streaming. In 2023 alone, SOCAN has licensed more than 170 billion individual performances of musical works to its licensees. There is nothing unique about generative AI that should preclude SOCAN from developing and offering an appropriate licensing scheme for TDM and other AI-related activities.

2. There should be no exception for TDM activities.

SOCAN urges the Government of Canada not to enact any new or modified copyright exceptions for TDM activities. An exception would wipe out the developing licensing market before it has a chance to mature and flourish. It would also deprive creators of the ability to control and be paid for valuable uses of their works by AI companies. Even if other jurisdictions choose to narrow or limit the scope of copyright protection in relation to TDM activities, Canada should resist the temptation to do the same.

SOCAN similarly urges the Government to reject any proposal that would allow an AI developer to use a creator’s works for TDM activities unless the creator “opts out”. Canadian copyright law is inherently an “opt-in” regime: it requires users to seek authorization from creators before using copyright-protected works. Requiring creators and their representatives to take positive steps to opt out of TDM activities by all AI users, in relation to every webpage or platform on which their works are available, and to monitor all AI platforms for compliance, would be an onerous and unfair task to impose on creators. It may also run afoul of international treaties to which Canada is a signatory, including the Berne Convention. The prejudice is exacerbated by the fact that, once an AI model has “learned” from the works of a creator who did not know they could opt out, or was unable to do so, that learning is extremely difficult, if not impossible, to reverse.

The fact that some AI companies indiscriminately scrape vast amounts of copyrighted works from the Internet, without permission or transparency, is not a reason to reward them with preferential treatment, either by enacting TDM exceptions or adopting an opt-out approach. To the contrary, robust copyright protection is necessary to ensure that users who wield such significant technological power do so in a responsible and ethical way that respects the creators on whose works they depend.

3. Remuneration is best determined in a voluntary licensing market

The appropriate level of remuneration for creators whose works are used to program AI models, or for other AI-related purposes, is best determined in the developing licensing market. A free market voluntary licensing system is the most likely way to allow creators to control the use of their works while requiring interested parties to agree on the terms, including the price, of that use.

Therefore, SOCAN urges the Government to avoid any approach that would limit a creator’s right to control the use of their works by AI companies. For example, the Government should avoid any form of compulsory licensing system, which would deny creators their right to contract freely in the market and, in doing so, to decide whether, how, and by whom their works are used. A compulsory licence would prevent creators from realizing fair value for the use of their works by AI companies. It would also raise concerns under Canada’s international treaty obligations, which require the core reproduction and performing rights in works to be true exclusive rights, not mere rights of remuneration. A free market voluntary licensing system must be protected to ensure that creators and AI companies can negotiate fair terms for the use of copyright-protected works.

4. Transparency is paramount.

Stakeholders generally agree that transparency, record-keeping, and disclosure obligations are important to ensure that creators understand how and when their works are used by AI developers and whether that use has been licensed or not. These requirements will help foster the developing licensing market by incentivizing AI developers to obtain permission before using copyright-protected works for programming or other purposes. They are also necessary to address AI’s “black box” problem, which makes it extremely difficult, if not impossible, for creators to know when their works have been used by AI models or to pursue legal remedies without proper disclosure from AI developers.

For a further discussion of record-keeping and disclosure obligations, please refer to our submission to the Consultation’s question on infringement and liability.

Authorship and Ownership of Works Generated by AI

It is not currently necessary to amend the Copyright Act to address authorship or copyright ownership of AI-generated or AI-assisted content. The Copyright Act is based on a philosophy of providing incentives for human creativity, and SOCAN considers the statute—including, for example, the tying of the general term of copyright ownership to the life of a human author—to be clear that the author of a work must be a human [Copyright Act, RSC 1985, c C-42, s 6].

The authorship and ownership of works that are created with the assistance of AI tools can be determined on the specific facts of each case. In its current form, the Copyright Act is sufficient to allow courts to develop the law by making those determinations. It would be prudent for the Government to avoid amending the Act unless and until the Government determines that judicial decisions, whether in Canada or elsewhere, expose gaps in the law that need to be addressed.

SOCAN’s operations—and, indeed, the music industry more broadly—are premised on the understanding that music creators are individuals, based on the longstanding principle that authors must be human individuals. If this longstanding principle is changed, it could have potentially unintended consequences across an entire music industry which is built on a foundation of human expression.

Infringement and Liability regarding AI

A critical issue in the field of AI is the “black box” problem, meaning there is a lack of visibility into the works used to program an AI model or how those works, once copied, affect the model’s programming.

Due to transparency problems, creators typically have no way to know or detect whether their works have been used to develop an AI model. Even if a creator suspects that their work has been used for programming purposes, it would be extremely difficult, if not impossible, to confirm that use. When large-scale infringements are carried out without the knowledge of creators, it creates a significant windfall for AI developers at the expense of creators.

The black box problem also makes it difficult, if not impossible, to prove that an AI-generated work infringes copyright in a creator’s existing work. Even if the two works are substantially similar, a lack of transparency and record-keeping by an AI company will thwart efforts to prove that the AI model used, and therefore had “access” to, the creator’s work, which is a necessary element of the test for copyright infringement.

In short, without appropriate transparency, disclosure, and record-keeping, it would be difficult, if not impossible, for a creator to know that its rights have been infringed, much less to pursue and obtain any remedy for that infringement.

SOCAN therefore urges the Government to require AI developers, and every person involved in the programming and testing of an AI model, to keep and make readily available detailed and accurate records of the works they have used for that development and how they have used them, including the source of the works and details of any licences authorizing the use of the works and how those works are kept, maintained, or stored by the AI developers. AI developers are best positioned to track that information.

SOCAN believes that, with robust transparency, record-keeping, and disclosure obligations, coupled with established copyright principles, the current Copyright Act will be sufficient to address liability issues specific to AI technology.

Comments and Suggestions

SOCAN encourages the Government to focus on the perspective of creators when considering AI. Specifically, AI developers must seek consent from creators before using their works, compensate creators for that use, and credit creators as co-authors of AI-generated works where applicable. Any copyright policy that sacrifices the interests of creators for technological advancement erodes the very concept of copyright, and the large-scale infringement by AI companies threatens to eliminate longstanding copyright principles altogether if the perspective of creators is neglected. Together, a free market voluntary licensing system and robust transparency obligations will ensure that Canada carries an appropriate copyright policy, with the interests of all parties adequately represented, into the age of generative AI.

Soproq

Preuve de nature technique

La Soproq est une société de gestion collective des droits des producteurs d'enregistrements sonores et de vidéoclips. Nous percevons et distribuons les redevances découlant des droits d’exécution publique, de ceux liés à la reproduction et au régime de copie privée pour les enregistrements sonores et les vidéoclips faisant partie de notre répertoire. Nous négocions également avec les services de musique et émettons des licences générales pour l’utilisation des titres contenus dans notre répertoire. Le répertoire de la Soproq inclut près de 2,5 millions de titres (enregistrements sonores et vidéoclips) appartenant à près de 7000 ayants droit provenant du Québec, du Canada et d’autres pays : entreprises de production, artistes auto-producteurs, distributeurs, associations, sociétés de gestion étrangères, etc.

La Soproq est heureuse de répondre à l’appel lancé par le gouvernement du Canada dans le cadre de la « Consultation sur le droit d'auteur à l'ère de l'intelligence artificielle générative ».

Plusieurs de nos sociétaires utilisent l’IA générative comme outil d’aide à la production et création d’enregistrements sonores et de vidéoclips. Ces derniers considèrent l’IA comme un outil facilitant certains procédés plus techniques mais estiment que l’exercice du talent et du jugement humain demeure toutefois à la base des principes du droit d’auteur. Afin de recevoir leurs redevances, en contrepartie de l’utilisation de leurs contenus protégés par les utilisateurs de musique, les ayants droit doivent déclarer à la Soproq les métadonnées relatives à ces contenus. La réception de ces métadonnées est accompagnée d’une déclaration de la personne qui soumet les données, attestant qu’elle détient les droits applicables sur le contenu soumis.

Actuellement, les métadonnées recueillies des ayants droit ne permettent pas à la Soproq d’identifier les enregistrements sonores ou vidéoclips qui auraient été produits, en partie ou en entièreté, par intelligence artificielle générative. Il serait néanmoins possible de mettre en place rapidement la cueillette de telles informations. La Soproq n’utilise pas les données recueillies à des fins d’entrainement de système d’IA générative.

Fouille de textes et de données

La popularisation de la Fouille de Textes et de Données (FTD) comme moyen principal pour fins d’apprentissage-machine marque une étape significative dans l’avancement de l'intelligence artificielle.  Selon OpenAI, compagnie derrière l’intelligence artificielle générative ChatGPT, leur système utilise une combinaison de trois sources de données pour entrainer l’algorithme : « (1) Information publiquement disponible sur internet, (2) Information acquise sous licences de tiers, et (3) Information fournie par nos utilisateurs ou formateurs humains » (https://help.openai.com/en/articles/7842364-how-chatgpt-and-our-language-models-are-developed ) [notre traduction]. De ces trois sources d’information, seule la deuxième semble respecter le principe du droit d’auteur puisque l’information est acquise sous licence.

D’abord, une distinction importante doit être faite entre « information publiquement disponible » et information libre de droits. Par exemple, un quotidien publiant les articles de ses journalistes sur son site web rend ceux-ci accessibles publiquement, mais ne libère pas nécessairement les droits et privilèges que leur accorde la Loi sur le droit d’auteur. En ce sens, bien que le contenu utilisé pour entraîner les systèmes génératifs soit disponible publiquement sur le web et puisse être moissonné en grande quantité à l’aide de programmes de type web-crawlers, l’utilisation de contenus protégés viole les droits des créateurs lorsque ces contenus ne sont pas adéquatement libérés. De plus, les contenus saisis par les utilisateurs ou les formateurs humains, soit la troisième source de données mentionnée par OpenAI, peuvent eux aussi contenir du contenu protégé, et donc, doivent aussi être dûment libérés et compensés. 

L’ampleur de l'essor de l'IA rappelle la démocratisation du web au début des années 1990, impactant tous les pans de la société. Au même titre, nous constatons déjà qu’il y a un avant et un après ChatGPT. Cependant, le développement de l’IA générative présente une différence importante avec l’explosion du web d’il y a 30 ans. Contrairement aux fondements non lucratifs du développement web, la trajectoire des modèles d'IA générative est en grande partie définie par des acteurs multinationaux en quête de gains financiers. Cette distinction soulève des interrogations significatives quant à l'équité et à la juste distribution des avantages issus de ces percées technologiques.

La protection des contenus protégés par le droit d’auteur est la pierre d’assise de la culture des démocraties modernes. La raison d'être du droit d'auteur est de sauvegarder et de reconnaître la valeur de l'expression humaine. Toute proposition qui sape ou va à l'encontre de cet objectif doit être rejetée. Nul avancement technologique, aussi important soit-il, ne peut se soustraire aux principes fondamentaux du droit d’auteur sous prétexte que la libération des droits du matériel utilisé nuirait aux développements technologiques. Ainsi, l'accès aux contenus protégés par le droit d'auteur par ces systèmes génératifs, et la reproduction des contenus que ces systèmes font, constitue un point sensible pour le milieu culturel et ses représentants. La Loi stipule que « Le droit d’auteur sur l’œuvre comporte le droit exclusif de produire ou reproduire la totalité ou une partie importante de l’œuvre, sous une forme matérielle quelconque » (art 3). La reproduction inhérente à la FTD ne peut être considérée comme une activité accessoire. Au contraire, elle se présente comme la matière première essentielle pour les systèmes d'IA générative. Dans cette optique, l'application d'une exception relative aux reproductions temporaires (art 30.71) ne semble pas appropriée, étant donné la centralité de la reproduction dans le processus technologique. Au contraire, la Loi accorde aux ayants droit un droit de paternité, un droit moral, et un droit exclusif quant à la reproduction de leurs contenus protégés et il est de la responsabilité des acteurs technologiques souhaitant reproduire ces contenus de les respecter.

En ce sens, l’introduction d’une exception à la Loi viendrait non seulement affaiblir les droits actuels des ayants droit, mais viendrait également couper des opportunités financières aux détenteurs de droits en limitant leurs possibilités de monétiser leurs contenus dans un marché libre. De plus, l’ajout d’exceptions à la Loi pour les activités de FTD constituerait un financement détourné des compagnies technologiques de la poche même des créateurs culturels. Rappelons que les entreprises mettant en marché les outils d’IA générative les plus populaires le font en partenariat avec certaines des compagnies les plus lucratives au monde, et qu’ils le font dans un but commercial, et non de façon altruiste.

L’industrie technologique argumentera certainement que de libérer les droits pour tout le contenu utilisé lors de l’apprentissage machine serait irréaliste, et qu’une telle restriction nuirait au développement technologique de leurs services. En ce sens, ils argumenteront sans doute qu’une distinction doit être faite entre le matériel protégé par le droit d’auteur et les données utilisées pour entraîner les algorithmes. Même le gouvernement canadien, dans les documents liés à cette consultation, fait le même raccourci en parlant de « grandes quantités de données, y compris celles extraites de contenus protégés par le droit d'auteur », comme si le contenu protégé perdait son identité aux mains de l’apprentissage-machine et était réduit à de simples données anonymes. Or, les contenus protégés sont la matière première des systèmes d’IA générative, au même titre que les enregistrements sonores sont la matière première des stations de radio. Si ces derniers doivent libérer les droits des contenus protégés qu’ils utilisent, il en va de même pour les compagnies derrière les systèmes d’IA générative. Ainsi, nous recommandons que les plateformes d'IA générative soient tenues de respecter des normes de transparence qui engloberaient la publication de registres renfermant des informations sur les contenus protégés utilisés à des fins d’apprentissage-machine.

Rappelons également que l’argument souvent entendu selon lequel la libération des droits sur le contenu utilisé comporterait un fardeau administratif trop grand est fallacieux. La gestion collective des droits permet aux utilisateurs de contenus protégés de libérer de façon simple et efficace les droits sur des contenus appartenant à un grand nombre d’ayants droit distincts. La gestion collective est déjà largement répandue au Canada et l’efficacité de ce type d’administration de droits n’est plus à démontrer. Dans cette perspective, il est fondamental de reconnaître le rôle crucial de l'autorégulation du marché, favorisant un équilibre entre la libération des droits et la facilité administrative.

Pour l’information publiquement disponible moissonnée sur le web, la Soproq dispose de différents modèles de licences permettant aux utilisateurs de faire la reproduction des contenus protégés faisant partie de son répertoire. Ces modèles peuvent s’adapter à tout développement technologique, y compris l’IA générative. Quant à la reproduction de contenus protégés lorsque de l’information est fournie par les utilisateurs ou formateurs humains, ce sont en plus les principes du régime de copie privée qui devraient s'appliquer.

Au sujet de la question du niveau de rémunération appropriée, il doit être évidemment juste et équitable en regard de l’importance de cette reproduction dans les activités de l’entreprise, étant entendu que ces reproductions sont au cœur des systèmes d’IA générative. Sans contenu pour entrainer les algorithmes, ces systèmes n’existent pas. Dans tous les cas, cette valeur doit être déterminée par le libre marché, en prenant en considération les principes déjà établis par la Commission du droit d’auteur.

Titularité et propriété des œuvres produites par l’IA

La Soproq est une société de gestion collective représentant principalement des producteurs d’enregistrements sonores et de vidéoclips ou des titulaires de droits sur ces contenus protégés.

La Loi définit le producteur initial comme étant « [l]a personne qui effectue les opérations nécessaires à la confection d’une œuvre cinématographique, ou à la première fixation de sons dans le cas d’un enregistrement sonore » (art 2). L’identification de la titularité du producteur initial représente un élément essentiel au travail des sociétés de gestion collective œuvrant dans le milieu de l’enregistrement sonore. Non seulement l’identité du producteur initial permet d’identifier le premier titulaire des droits associés à un contenu protégé, mais celle-ci permet également de déterminer si un enregistrement sonore ou un vidéoclip est éligible à certaines formes de rémunération.

Dans le cas d’enregistrements sonores créés à partir d’IA générative, plusieurs questions se posent : Qui est le producteur initial?  Est-ce que l’action de rédiger un ou plusieurs prompts dans le but de créer un enregistrement sonore constitue, au sens de la Loi, les opérations nécessaires à la première fixation de sons? Et où cette fixation a-t-elle lieu?

Pour répondre à ces questions, il peut être souhaitable de prendre un pas de recul. Dans un article publié dans Computer Law & Security Review en 2010, soit avant l’apparition grand public des IA génératives comme nous les connaissons aujourd’hui, les professeurs de droit à l’université Western Ontario Mark Perry et Thomas Margomi soulignent « qu’il faut prendre en compte si une œuvre générée par ordinateur peut remplir les critères de protection en fonction de l'exigence d'originalité rapportée » (Perry & Margomi 2010, p. 6)[notre traduction]. Les auteurs soulignent que dans son jugement de la cause CCH Canadienne Ltée c. Barreau du Haut-Canada, la Cour suprême du Canada rappelle que, « [p]our être ‘originale’ au sens de la Loi sur le droit d’auteur, une œuvre doit être davantage qu’une copie d’une autre œuvre. Point n’est besoin toutefois qu’elle soit créative, c’est-à-dire novatrice ou unique. L’élément essentiel à la protection de l’expression d’une idée par le droit d’auteur est l’exercice du talent et du jugement.  J’entends par talent le recours aux connaissances personnelles, à une aptitude acquise ou à une compétence issue de l’expérience pour produire l’œuvre.  J’entends par jugement la faculté de discernement ou la capacité de se faire une opinion ou de procéder à une évaluation en comparant différentes options possibles pour produire l’œuvre.  Cet exercice du talent et du jugement implique nécessairement un effort intellectuel.  L’exercice du talent et du jugement que requiert la production de l’œuvre ne doit pas être négligeable au point de pouvoir être assimilé à une entreprise purement mécanique. » (CCH Canadienne Ltée c. Barreau du Haut-Canada, par 16).

À la lumière du jugement de la Cour de la cause CCH Canadienne Ltée c. Barreau du Haut-Canada, il nous apparait que le cadre législatif actuel accorderait à l’utilisateur, et non l’IA, un droit d’auteur sur une œuvre créée par IA générative, dans la mesure où l’utilisateur peut démontrer avoir exercé, au sens de la loi, talent et jugement pour générer le contenu via la rédaction des prompts et un processus décisionnel réfléchit sur le produit sortant. Ainsi, des enregistrements produits par IA de façon purement mécanique (par exemple, à partir d’un autre algorithme) devraient se voir refuser une protection puisqu’ils ne remplissent pas le critère de jugement exigé par la Loi.

En ce sens, nous jugeons que la Loi telle qu’actuellement rédigée est suffisamment robuste pour s’adapter aux nouvelles réalités technologiques, notamment à l’IA générative. Ainsi, nous sommes d’avis que la Loi se doit de rester technologiquement neutre et ne doit pas être modifiée pour s’adapter à la technologie du moment.

Violation et responsabilité en matière d’IA

L’un des enjeux principaux avec les systèmes d’IA générative est la traçabilité des données utilisées lors de l’apprentissage-machine. Cette traçabilité est essentielle d’un point de vue de santé et sécurité publique. On constate déjà comment ces modèles génératifs peuvent être utilisés comme armes de désinformation massive. Cette désinformation pourrait être accentuée si les données utilisées pour entrainer les algorithmes comportent des biais importants, ou pire, des discours haineux.  Mais au-delà de ces questions, la traçabilité des données est également le seul moyen de s’assurer que les compagnies derrière les IA génératives respectent les créateurs du contenu qu’ils utilisent.

Dans un article publié le 27 juin 2023, l’éditeur de la revue scientifique Nature arguait que « les entreprises technologiques doivent formuler des normes industrielles pour le développement responsable des systèmes et outils d'IA, et réaliser des tests de sécurité rigoureux avant la mise en vente des produits. Elles devraient soumettre l'ensemble des données à des organismes de réglementation indépendants capables de les vérifier, tout comme les entreprises pharmaceutiques doivent soumettre les données d'essais cliniques aux autorités médicales avant la mise en vente des médicaments » (Nature, Vol. 618, pp 885-886) [notre traduction].

Nous partageons le désir soulevé par l’équipe éditoriale de la revue scientifique Nature de voir les firmes d’IA divulguer l’intégralité des données utilisées lors de la phase d’apprentissage de leurs systèmes, ainsi que de faire preuve d’une plus grande transparence sur les processus utilisés. Comme mentionné précédemment, la compagnie OpenAI affirme que ses données proviennent de trois sources différentes : « (1) Information publiquement disponible sur internet, (2) Information acquise sous licences de tiers, et (3) Information fournie par nos utilisateurs ou formateurs humains ». Mais sans traçabilité des données, comment s’assurer que les données appartenant à la première et à la troisième catégorie n’auraient pas dû, elles aussi, avoir été acquises sous licences de tiers et être dûment compensées?

Chaque reproduction d’un contenu protégé se doit d’être compensée en vertu du cadre législatif canadien actuel. Les reproductions faites dans le cadre d’un processus d’entrainement de système d’IA générative ne font pas exception à la règle.

À titre comparatif, les radiodiffuseurs ont un devoir de traçabilité et ont le devoir de rapporter l’assemble des enregistrements sonores diffusés afin que les ayants droit soient compensés conformément aux tarifs établis. Les firmes technologiques qui utilisent des contenus protégés devraient être soumises aux mêmes exigences.

De plus, l’IA ne doit pas faire fi du droit de paternité et des droits des créateurs. Le gouvernement doit prendre en considération l'impact de l'IA générative sur les droits moraux des auteurs, y compris l'intégrité de leurs œuvres et les droits adjacents tels que le nom, l'image, les droits de la personnalité et de la publicité. Ainsi, la consultation sur l’IA générative devrait viser à favoriser un marché de licences sain pour les contenus protégés. Les obligations de transparence, incluant la divulgation de registres du contenu protégé utilisé à des fins d’apprentissage-machine, sont cruciales pour permettre à l'IA générative de continuer à innover au sein d'un système de droit d'auteur encourageant la création et assurant une compensation équitable aux créateurs.

Commentaires et suggestions

Dans la dynamique évolutive de l'intelligence artificielle générative, l'équilibre entre la préservation du droit d'auteur et la promotion de l'innovation émerge comme une nécessité impérative. Nous espérons que ces considérations contribueront à guider le gouvernement dans l'élaboration de politiques qui favorisent le progrès technologique, tout en préservant les intérêts des créateurs au sein du Canada.

Nous soutenons que le cadre législatif actuel est suffisamment robuste pour s’assurer de cet équilibre, à condition qu’il soit appliqué tel quel, sans nouvelles exceptions. La montée en puissance de l'IA n’est que le plus récent d’une longue série de cataclysmes technologiques débutant il y une trentaine d’années avec la démocratisation du web, l’iPod, le téléphone intelligent et l’écoute en continu. Au cours de ces trois décennies, multiples exceptions et exemptions ont été introduites à la Loi sur le droit d’auteur, privant les ayants droit d’une rémunération juste et équitable. Nous saluons l’initiative du gouvernement canadien quant à la consultation actuelle menée au sujet du cadre du droit d'auteur et des technologies d'IA. Bien que la conversation sur l'IA soit essentielle, nous ne pouvons pas perdre de vue les amendements nécessaires à la Loi sur le droit d'auteur qui peuvent être promulgués immédiatement, qui sont basés sur le marché et ne nécessitent aucun financement supplémentaire du gouvernement. Ces amendements permettront non seulement de maintenir la neutralité technologique de la Loi, mais favoriseront également l'équité pour des milliers de créateurs de partout au pays. Ainsi, nous réitérons les trois demandes du milieu musical canadien, soit (1) de modifier la définition d’enregistrement sonore de manière à ce que les artistes-interprètes, les producteurs et les maisons de disques puissent toucher une rémunération équitable lorsque leurs enregistrements sonores sont utilisés dans un film, à la télévision ou à même un autre contenu audiovisuel ; (2) d’éliminer l’exemption introduite en 1997 permettant aux radios commerciales de ne pas verser de redevances pour l’exécution publique d’enregistrements sonores sur leurs ondes pour les premiers 1,25M$ de revenus publicitaires ; et (3) d’actualiser le régime de copie privée pour maintenir sa neutralité technologique et ainsi percevoir une redevance à la vente de tous produits pouvant stocker des copies de musique.

En plus du travail de réflexion sur l’IA, la Soproq demande ardemment au gouvernement canadien de corriger ces aberrations dans la Loi. Les revendications de la Soproq s’alignent avec celles des autres intervenants du secteur des droits voisins au Canada.

Finalement, nous tenons également à souligner que la Soproq est membre de la Coalition pour la diversité des expressions culturelles (CDEC). À ce titre, nous avons contribué à l’élaboration des recommandations fournies par la Coalition et nous endossons, en plus des réponses fournies ici, les réponses fournies par la CDEC.

Stephen Spong

Technical Evidence

The speed with which AI has evolved even in the past 12 months shows little sign of abating, even with legal challenges such as recently brought forward by the New York Times against OpenAI attempting to curb this growth. At the time of this writing (January 2024), being overly hasty in coming to conclusions about the current capacity for AI systems could lead to unintended consequences in terms of overly or insufficiently regulated areas. As per the Canadian Association of Research Libraries’ (CARL) submission, current best practices should be guided by the Innovation, Science, and Economic Development Canada’s Voluntary Code of Conduct on the Responsible Development and Management of Advanced Generative AI Systems as well as frameworks established through legislation.

While it is a delicate balance, at this point the restriction of AI should be limited as excessive restriction would likely have a deleterious effect on innovation. It is hoped that as the capacity and uses of AI are better understood, both governments and the courts would be able to further provide guidance through legislative and judicial means to ensure that growth is managed in a responsible manner.

Text and Data Mining

The inclusion of a line of questioning around TDM is curious given the current lack of explicit language in the Copyright Act that permits such activities or provides guidelines for what is legally permissible and - in the case of the Canadian copyright framework - non-infringing. This lack of clarity has diverted researcher focus, time, and effort away from their research topics with an unnecessary infringement risk.

It would be a welcome change to the Copyright Act to include language that provides clarity around TDM. The current framework (or lack thereof) impedes research by requiring a significant amount of copyright due diligence before meaningful work can be done with the raw data. A distinction should be made regarding commercial and non-commercial TDM activities, with the latter including research and educational uses that are non-compensable.

Both the Canadian Association of Research Libraries (CARL) and Canadian Federation of Library Associations (CFLA) submissions include expanded rationales for expanding TDM provisions, highlighting that many key Canadian trading partners including the UK, EU, and Japan already have such provisions and that there has been a commensurate uptick in research activity and outputs. This submission supports that position.

Authorship and Ownership of Works Generated by AI

Because AI-generated works do not involve a human exercise of skill and judgement, they do not merit copyright protection. This is a principle that is by now well-established in case law, starting with CCH Canadian Ltd. v. Law Society of Upper Canada, 2004 SCC 13, [2004] 1 SCR 339.

However, in response to the Government’s question of “propos[ing] any clarification or modification of the copyright ownership and authorship regimes in light of AI-assisted or AI-generated works? If so, how?” it would be worthwhile to amend the Copyright Act to further enshrine the principle of human involvement as a necessary element of authorship.

Infringement and Liability regarding AI

Liability and statutory damages for infringement are already set out in the Copyright Act, with limitations for educational institutions, libraries, archives, and museums.

Comments and Suggestions

N/A

Stratford Intellectual Property - Part of Stratford Group Ltd.

Technical Evidence

N/A

Text and Data Mining

It is our view that companies and organizations that perform TDM activities log all data used. They should not only take note of the extracted data, but also the jurisdiction in which the data is extracted. Compensation models can then be developed to ensure that copyright owners are properly recognized for their original ideas. It could also be envisioned that copyright holders could also include a notice around terms of use of their content for data training.

Given the rise in generative AI platforms from large US-based tech companies such as OpenAI, Microsoft and Google, it is no surprise that Canada has followed suit to embrace the new age of artificial intelligence. For generative AI to operate, it is trained with large amounts of data sets obtained through text and data mining (TDM) activities.

Although there are concerns about these TDM activities involving the use of copyrighted material without consent, there have been benefits to doing so as well. For instance, during the height of the COVID-19 pandemic, the BlueDot program, founded by a professor from the University of Toronto, provided useful information for public health authorities to quickly take action. However, this would have not been possible if there were limits to the use of public, commercial and academic data sets to train this program and map the spread of infectious diseases (Fiil-Flynn et al., 2023). Both companies behind generative AI platforms and authors of copyrighted material should find a middle ground to respect copyright ownership without completely shutting down the endless possibilities for innovation through this new technology.

During the infancy of generative AI, there were not many clear regulations when it comes to TDM activities in many parts of the world. As such, tech companies behind these platforms have found themselves freely scraping the World Wide Web to obtain data sets for AI training, including original copyrighted material from different artists and authors. At the time, no one could seem to stop these companies from doing so – or so they thought. They later found themselves in hot water as multiple class action lawsuits have piled up against them due to copyright infringement (Lutkevich, 2024).

Content creators are also discussing whether to allow their data to be used in TDM.   For example, when Creative Commons polled respondents on whether openly licensed content should be used to train AI models, nearly half the respondents were “it depends” (Vézina, 2021). 

Because of the public outcry against TDM activities, software developers have come up with ways to control the data fed into AI training models. One example is the platform called Glaze, which aims to prevent AI models from learning a particular artist’s distinctive style. On the other hand, there are some companies, such as Bria, a generative AI startup, that has committed to pay royalties to artists when their works are used to train their AI models (Glover, 2023). Having these examples simply show that it is not impossible to find solutions to meet both the needs of the AI and creative industries.

We recommend that companies and organizations that perform TDM activities responsibly log all data used to train their AI models. They should not only take note of the type of extracted data, but also the jurisdiction in which the data is extracted from due to varied copyright considerations. Compensation models can then be developed to ensure that copyright owners are properly recognized for their original ideas. For example, compensation models could be on a pay-per-use basis. 

References:

1) Fiil-Flynn, S. et al. (2022, December 1). Legal reform to enhance global text and data mining research. https://www.science.org/doi/10.1126/science.add6124  

2) Lutkevich, B. (2024, January 2). AI lawsuits explained: Who’s getting sued? https://www.techtarget.com/WhatIs/feature/AI-lawsuits-explained-Whos-getting-sued  

3) Vézina, B et al. (2021, March 4) Should cc-licensed content be used to train ai? It depends. https://creativecommons.org/2021/03/04/should-cc-licensed-content-be-used-to-train-ai-it-depends/

4) Glover, E. (2023, August 23). AI-Generated Content and Copyright Law: What We Know. https://builtin.com/artificial-intelligence/ai-copyright

Authorship and Ownership of Works Generated by AI

Briefly, our view is that they should be a harmonized and aligned system for copyright as it applies to ownership and inventorship of AI generated work globally, like other intellectual property right systems.

Generally, the authorship of copyright on AI-assisted and AI-generated content should be attributed to the individual (or organization) that originally created the content, not to the person who arranged for the work to be created. Content should be clearly attributed if taken from training data. If the AI-system is modifying the data, then citing the source should be done. Clearly mapping use to creation will also support compensation allocation. If the generative AI-tool is inferring or generating a hallucination, it should be marked as such and would support reducing misinformation.  

Authorship of copyrighted works in Canada should be clearly defined in light of AI-assisted and AI-generated works as it has been in the United States. In 2023, the US Copyright Office stated that there should be sufficient human authorship in AI-generated material, but even then, copyright protection will only protect the human-authored aspects of the work. It has also been determined that simply prompting a generative AI platform to produce an output is considered to completely lack human authorship. Furthermore, if applicants will be submitting AI-generated content for copyright registration, they are required to disclose that it has been produced by an AI platform (US Copyright Office, 2023).

One of the infamous cases in the US relating to authorship and AI-generated works is that of the case of Dr. Stephen Thaler and his AI machine called DABUS, which stands for Device for Autonomous Bootstrapping of Unified Sentience. Besides being involved in a case of inventorship in patent applications, DABUS has found its way to be listed as an author for a copyright registration application for an AI-generated image titled “A Recent Entrance to Paradise”. Since the US Copyright Office has already established the guidelines surrounding authorship and copyrights at the time, it was clear to a US Federal Judge that AI-generated images cannot be granted copyright protection due to the lack of human authorship (Growcot, 2023). In a similar manner, the graphic novel by Kristina Kashtanova titled “Zarya of the Dawn” was created using the Midjourney platform, was initially granted copyright registration. However, it has since been lost since it has been found that the generative AI platform was significantly lacking human involvement in its conception (Wolfson, 2023).

On the other hand, across the border in Canada, we find an interesting case where a fully AI-generated work was granted copyright registration (Stephens, 2023), and is considered to be registered with CIPO to this day. This illustrates that there are still gaps in the Canadian intellectual property system, and we believe that it will be best to emulate a similar approach taken by the US Copyright Office when it comes to their guidelines surrounding AI-assisted and AI-generated works. In addition, we recommend adding due diligence efforts to ensure that there is sufficient human authorship in works already granted copyright registrations in Canada, especially those that have been filed shortly after the launch of generative AI platforms.

References:

1) US Copyright Office, Library of Congress. (2023, March 16). Copyright Registration Guidance: Works Containing Material Generated by Artificial Intelligence. https://www.govinfo.gov/content/pkg/FR-2023-03-16/pdf/2023-05321.pdf  

2) Growcoot, M. (2023, August 21). Federal Judge Rules AI Images Cannot be Copyrighted, Contrasts it to Photos. https://petapixel.com/2023/08/21/federal-judge-rules-ai-images-cannot-be-copyrighted-contrasts-it-to-photos/

3) Wolfson, S. (2023, February 27). Zarya of the Dawn: US Copyright Office Affirms Limits on Copyright of AI Outputs. https://creativecommons.org/2023/02/27/zarya-of-the-dawn-us-copyright-office-affirms-limits-on-copyright-of-ai-outputs/

4) Stephens, H. (2023, April 19). Canadian Copyright Registration for my 100 Percent AI-Generated Work. https://hughstephensblog.net/2023/04/19/canadian-copyright-registration-for-my-100-percent-ai-generated-work/

Infringement and Liability regarding AI

Briefly, our view is that organizations and innovative development will benefit from clear direction on parties responsible for infringement and associated liability. Whether it is the companies behind generative AI tools or the end user – defining where the boundaries exist will benefit the community.  AI systems should be liable for infringement of copyrights if not properly attributed or if modified without consent. The law should be clear that although the use of the copyrighted information has been done by an AI system, the developer (human) of the AI system is liable for the copyright infringement even if the machine has evolved to use other content. In the end the accountability and liability around the use of AI remains with the developers. 

Interestingly, whereas in the patent system, it is common for the patent owner to go after the system manufacturer for infringement of its component (instead of the component manufacturer). Within AI-generated works, the opposite seems to be mostly happening – where the generative AI-company is being sued rather than its end user. This makes sense from not only a practical perspective, one entity to go after rather than all its users and it is the generative AI company that used the data to train the models.    

Given the backlash following the rise in AI-generated works, several tech companies have developed measures to mitigate the risks of liability for infringement – not for themselves, but rather for their users.

In September 2023, Microsoft announced that they will provide legal protection for customers who are sued for copyright infringement over content generated by the company’s AI systems – GitHub Copilot and Bing Chat (Criddle, 2023). More recently, OpenAI offered to cover their clients’ legal costs for copyright infringement suits, calling it Copyright Shield. However, there is a caveat as it only applies to users of their business tier, ChatGPT Enterprise, excluding the free version of ChatGPT or ChatGPT+. In addition to OpenAI, companies such as Google, Microsoft and Amazon have also developed similar terms to protect users of infringement with the use of their generative AI platforms (Montgomery, 2023).

Numerous active lawsuits have been filed against companies behind generative AI tools because they have the ability to create unauthorized, derivative outputs of copyrighted works – but should these companies really be held liable, or should it be directed towards the user prompting the tool to generate the derivative works? Although a user directly prompts the generative AI platform, the generative AI platform is not completely out of water. For instance, the class action lawsuit against Stable Diffusion and Midjourney claims that companies should be held vicariously liable for copyright infringement. In the United States, there are 2 conditions in which a third party (i.e., the company) can be held liable for copyright infringement: (1) the third party has the ability to supervise and control acts of the person who committed the direct infringement, and (2) the third party has an “obvious and direct” financial benefit from the infringing activity.” Although companies currently may not have the capacity to scan through the use of their tools and limit infringing activities, we believe that they should have the ability to do so. Furthermore, it does not appear that companies have an “obvious and direct” financial benefit from copyright infringement at first glance. However, it has become more common to encounter generative AI platforms that operate based on a subscription, obviously showing that they have a direct financial benefit from the use of their tool, including when it is used to generate works infringing on copyrighted material.

Similar to Microsoft’s Copilot Copyright Commitment, indemnification should be clear. In this case, if a user of Microsoft’s Copilot prompts the tool to generate outputs with copyrighted material, then Microsoft promises to defend the user and cover any expenses involved in the copyright litigation process (Smith, 2023). This not only helps manage risk to provide users peace of mind, but also pushes a company to ensure that they have the necessary measures in place to respect the copyrights of original works.

References:

1) Criddle, C. (2023, September 7). Microsoft pledges legal protection for AI-generated copyright breaches. https://www.ft.com/content/cd7f5391-bba5-4af1-8309-346eb2eafa02

2) Montgomery, B. (2023, November 6). OpenAI offers to pay for CHATGPT customers' copyright lawsuits. https://www.theguardian.com/technology/2023/nov/06/openai-chatgpt-customers-copyright-lawsuits

3) Wolfson, S. (2023, March 24). Style, Copyright, and Generative AI Part 2: Vicarious Liability. https://creativecommons.org/2023/03/24/style-copyright-and-generative-ai-part-2-vicarious-liability/

4) Smith, B. (2023, September 7). Microsoft announces new Copilot Copyright Commitment for customers. https://blogs.microsoft.com/on-the-issues/2023/09/07/copilot-copyright-commitment-ai-legal-concerns/

Comments and Suggestions

N/A

T

Vandana Taxali

Technical Evidence

As a lawyer specializing in intellectual property, technology, and the arts/entertainment sectors, my work primarily revolves around representing and collaborating with individuals and organizations within the arts and creative industries. I am also the founder of Artcryption, an art+tech platform for artists and creators.  I had the opportunity to engage with artists and gather insights during the week-long digiArt Art + Tech conference which I helped organize and was held on November 24, 2023 until Nov 30th. The valuable feedback and perspectives shared by these artists have significantly informed my views on the matter.

During the conference, a comprehensive poll revealed that a minority of artists acknowledged the integration of AI into their artistic practices, while the majority did not incorporate AI technologies into their creative processes. It became evident that artists who did engage with AI often used it for generating or assisting in the creation of their artworks. Notably, a common sentiment among these artists was a lack of awareness or clarity regarding the training datasets employed in AI models, raising concerns about the source and usage of these datasets.

Nevertheless, it is essential to highlight that a substantial portion of the artistic community expressed a strong desire for more explicit and well-defined copyright protection of their art and copyright protected works in AI models. These artists voiced their opposition to the incorporation of AI in AI models, emphasizing the need for clearer rules to safeguard their creative works in the evolving landscape of AI-generated art. These insights underscore the importance of addressing copyright concerns and providing guidance to artists as they navigate the intersection of AI and the arts.

Text and Data Mining

Transparency and clarity concerning the use of copyright-protected works are crucial in the evolving landscape of AI and creative industries. There is a growing consensus for increased transparency and well-defined guidelines for attribution and compensation.  AI developers should indeed maintain records and disclose their utilization of copyright-protected works in accordance with the provisions set forth in the Copyright Act.

However, it's worth noting that during our survey of artists, many respondents expressed a lack of awareness regarding how to grant permission or monetize the use of their works in AI models. Interestingly, a significant portion of artists did not express a strong desire to do so. Therefore, it is imperative that mechanisms are in place to simplify the process for artists who wish to provide licenses for their work, ensuring accessibility and ease of use.  There is a growing consensus for increased transparency and well-defined guidelines for attribution and compensation. 

Within the legal realm, the discourse surrounding the choice between opt-in and opt-out consent models transcends mere privacy considerations and delves into the core of artists' autonomy over their creative works. Opt-in models not only uphold ethical principles but also empower artists by enabling them to make conscious and informed choices regarding how their creations are utilized as per the Copyright Act. 

Furthermore, the notion of collective societies emerges as a potential avenue for compensation rights holders for the use of their work in AI models.  However, a better approach would be to provide the artists/creator with complete autonomy and control over the amount of fee, and conditions for usage of their work.  A variation of the Creative Commons licenses for the use of artists works in AI may provide artists a greater efficiency in providing permissions and licenses.  However, it is imperative that an artist/creator/copyright owner in the creative industry have complete control over which option they prefer

There is a need for support  including legal defense and advocacy for the creative industry, in navigating the complex terrain of copyright in the AI era. Many corporations controlling AI have legal defence funds making it difficult for an artist or creator to protect their rights.  A collective approach not only ensures that artists' rights are safeguarded but also facilitates a unified voice in negotiating fair compensation and protection.

The US Copyright Office has previously ruled out copyright registered to an AI as the author instead of a human  in 2022 noting it “lacks the human authorship necessary to support a copyright claim”.

Authorship and Ownership of Works Generated by AI

Under the Copyright Act, copyright owners possess the inherent right to attribution, credit, and the exclusive permission for the use of their intellectual creations. This fundamental principle should extend seamlessly to the realm of AI-assisted and AI-generated works. Copyright owners must retain the unequivocal ability to grant authorization for the utilization of their works within AI systems, while also delineating the specific conditions and parameters governing such usage.

It is imperative for the government to play a role in ensuring clarity and transparency when AI models incorporate copyright-protected works with the appropriate permissions. This entails clearly indicating whether a work is being licensed within an AI model and enforcing the obligation to provide due credit and attribution, contingent upon the permissions granted by the artist or creator.

I concur with the approach adopted by the United Kingdom, which involves the creation of a code of practice for the use of AI systems. Such a framework enhances clarity for both copyright owners and users, fostering a balanced environment where authorship and ownership of AI models are well-defined and respected. This approach aligns with the evolving landscape of AI technology and copyright in the modern era, where clear guidelines are essential to address the intricacies of authorship and ownership in AI models.

There are a also number of legal risks in the use of AI for the creative industries and include the following:

Copyright Infringement:

The use of AI in creative industries may involve the incorporation of copyrighted materials without proper authorization, potentially leading to copyright infringement.

Lack of Transparency:

The opacity of AI-generated creative processes can create challenges in identifying the origin of content, raising issues of transparency and attribution.

Privacy Violations:

AI systems may collect and process personal data, risking privacy breaches and violations of data protection laws if not managed appropriately.

Ownership and Authorship Ambiguity:

Description: Determining ownership and authorship of AI-generated artworks or content can be complex, leading to disputes over intellectual property rights.

Ethical Concerns:

The use of AI in creative works may raise ethical questions, such as the potential for bias in AI-generated content or the use of AI to replicate an artist's style without consent.

Data Security:

The handling and storage of large datasets for AI training can pose security risks if not adequately protected, leading to data breaches and legal repercussions.

Regulatory Compliance:

Description: Compliance with evolving regulations related to AI, copyright, data privacy, and consumer protection is essential, with non-compliance carrying legal penalties.

Algorithmic Accountability:

Lack of accountability in AI algorithms can result in unintended consequences, discrimination, or bias, leading to legal challenges and liabilities.

Licensing and Permissions:

The use of third-party datasets or copyrighted materials in AI applications requires proper licensing and permissions, with failure to do so risking legal action.

Liability for AI-generated Content:

Determining liability for AI-generated content, especially in cases of harm or misinformation, poses legal challenges that need resolution.

These legal risks underscore the complexity of integrating AI into the creative industries, necessitating clear legal frameworks and responsible practices to mitigate potential legal issues.

Infringement and Liability regarding AI

The existing legal criteria for determining copyright infringement may prove inadequate in cases where the authorship of AI-generated works is ambiguous, particularly when multiple pre-existing works are utilized or amalgamated. To address this issue and enhance transparency, it is essential that when users create an output using AI, they are provided with information about which copyright-protected works contributed to the output.

It's worth noting that many artists are reluctant to embrace AI due to concerns about the use of copyrighted works without permission. This hesitancy reflects a broader sentiment within the artistic community, emphasizing the need for greater clarity in defining liability when AI-generated works potentially infringe upon copyright.  Anti-AI theft technologies such as Nightshade and Glaze are emerging as protective measures and should be explored provided their use is lawful and ethical. Nightshade employs data poisoning to introduce errors into AI training datasets, and Glaze masks artwork to prevent scraping.

The Copyright Board in the US has provided guidance on whether AI assisted-works can get a copyright registration.  Further, current developing cases of various class actions in the creative industries also holds significant promise in offering guidance.

It is crucial to strike a balance where the rights of copyright owners are preserved and not diminished in any way under the Copyright Act. A comprehensive and evolving legal framework will help navigate the intricate terrain of copyright in the age of AI, ensuring that authorship, ownership, and protection remain robust and equitable.

The US Copyright Office Guidance: Work Containing Material Generated by Artificial Intelligence,  March 16, 2023 is extremely helpful and should serve as guidance for the Canadian government.

- The US copyright reviewed a graphic novel in February 2023 comprised of human authored text combined with images generated by the AI service Midjourney constituted a copyrightable work, but that the individual images themselves could not be protected by copyright.

- “In the Office’s view, it is well established that copyright can protect only material that is the product of human creativity. Most fundamentally, the term “author,” which is used in both the Constitution and the Copyright Act, excludes non-humans”.

- “a human may select or arrange AI-generated material in a sufficiently creative way “that the resulting work as a whole constitutes an original work of authorship”.

- In previous case law, a monkey cannot register a copyright in photos it captures with a camera because the Copyright Act refers to an author’s “children,” “widow,” “grandchildren,” and “widower,”— terms that “all imply humanity and necessarily exclude animals.”

- Artist can modify AI generated works as long as such modifications meet the standard for copyright protection

- In the case of works containing AI-generated material, the Office will consider whether the AI contributions are the result of “mechanical reproduction” or instead of an author’s “own original mental conception, to which [the author] gave visible form.” 24 The answer will depend on the circumstances, particularly how the AI tool operates and how it was used to create the final work.25 This is necessarily a case-by case inquiry.

- Prompts are not considered copyright protectible as prompts as the “traditional elements of authorship” are determined and executed by the technology – not the human user.  Users do not exercise ultimate creative control over how such systems interpret prompts and generate material. Instead, these prompts function more like instructions to a commissioned artist— they identify what the prompter wishes to have depicted, but the machine determines how those instructions are implemented in its output.

Comments and Suggestions

UNESCO, Recommendations, 2021 are also insightful and can provide guidance to the government. 

The General Conference of the United Nations Educational, Scientific and Cultural Organization

Recommendations:

“that globally accepted ethical standards for AI technologies, in full respect of international law, in particular human rights law, can play a key role in developing AI-related norms across the globe.”

“considers ethics as a dynamic basis for the normative evaluation and guidance of AI technologies, referring to human dignity, well-being and the prevention of harm as a compass and as rooted in the ethics of science and technology.”

“approaches AI systems as systems which have the capacity to process data and information in a way that resembles intelligent behaviour, and typically includes aspects of reasoning, learning, perception, prediction, planning or control.”

“provides ethical guidance to all AI actors, including the public and private sectors, by providing a basis for an ethical impact assessment of AI systems throughout their life cycle.”

Principles:

Proportionality and Do No Harm

Safety and security

Fairness and non-discrimination

Sustainability

Right to Privacy, and Data Protection

Human oversight and determination

Transparency and explainability

Responsibility and accountability

Awareness and literacy

Multi-stakeholder and adaptive governance

TECHNATION Canada

Technical Evidence

How does your organization access and collect copyright-protected content, and encode it in training datasets?

Data that is publicly available online is needed to develop safe and performant AI that is unbiased. Much of the data needed for AI development generally is collected from publicly available sources online. It is necessary to train large-scale AI on vast, broad and varied data sets to ensure correctly functioning, safe and unbiased AI. Using the public internet as a source of information is necessary to achieve this scale.

How does your organization use training datasets to develop AI systems?

N/A for trade association

In your area of knowledge or organization, what measures are taken to mitigate liability risks regarding AI-generated content infringing existing copyright-protected works?

N/A for trade association

How do businesses and consumers use AI systems and AI-assisted and AI-generated content in your area of knowledge, work, or organization?

AI and AI-generated outputs are used across all areas of industry and society. It will be used in all areas of industry, manufacturing, including material development, health, including drug discovery, etc. Conversational interfaces (chatbots) will be used by people to access many different AI services. Applications will be mostly unrelated to the interests of rightsholders.

Text and Data Mining

What would more clarity around copyright and TDM in Canada mean for the AI industry and the creative industry?

More clarity that ensures that TDM is not considered a copyright infringement would enable greater confidence in the development, use and investment in AI development. It would also enable content creators to understand how their works may be used for AI training.

Are TDM activities being conducted in Canada? Why or why not?

AI development will thrive in jurisdictions that support responsible AI development. Training will occur predominantly in jurisdictions such as the US and Japan where there is more clarity.

Legislators and policymakers must not overlook that TDM applies to data analysis more generally; for example, national security techniques may require analysis of large volumes of data, which may include photographs and text documents.

We assume that these are conducted in Canada. However, the legality of such methods also risks being put in question if the law is amended to make TDM a copyright infringement. AI use and data analysis techniques that are critical to many use cases would benefit from amendments to the law to clarify that TDM is not a copyright infringement.  

Are rights holders facing challenges in licensing their works for TDM activities? If so, what is the nature and extent of those challenges?

Licensing for TDM is not required because TDM is not a copyright infringement. While there is no explicit TDM exception in Canada, other exceptions and limits on copyright exist that permit TDM for commercial purposes. This does not limit the ability of rightsholders to build revenue models around their works that grant access to works that are not publicly available. 

What kind of copyright licenses for TDM activities are available, and do these licenses meet the needs of those conducting TDM activities?

Licences are not required for TDM – we are not aware of any such licences. 

If the Government were to amend the Act to clarify the scope of permissible TDM activities, what should be its scope and safeguards? What would be the expected impact of such an exception on your industry and activities?

Clarity could be improved by an express provision stating that performing TDM is not a copyright infringement. The right to read should be equivalent to the right to mine.

Should there be any obligations on AI developers to keep records of or disclose what copyright-protected content was used in the training of AI systems?

No

What level of remuneration would be appropriate for the use of a given work in TDM activities?

Remuneration is not required. If a person has legal access to a work, they should not be required to pay to read or learn from the work.

Are there TDM approaches in other jurisdictions that could inform a Canadian consideration of this issue?

Japan

Authorship and Ownership of Works Generated by AI

Is the uncertainty surrounding authorship or ownership of AI-assisted and AI-generated works and other subject matter impacting the development and adoption of AI technologies? If so, how?

The use of AI as a tool to create new artistic works should not prevent a person from being entitled to own the copyright in a work that they have created. The law permits a person who authors a work, even when using AI, to own the copyright work, and there is no need to amend the law in this regard. 

Should the Government propose any clarification or modification of the copyright ownership and authorship regimes in light of AI-assisted or AI-generated works? If so, how?

No legislation change is needed. 

Are there approaches in other jurisdictions that could inform a Canadian consideration of this issue?

No. The UK uniquely has a provision for protecting computer-generated works where there is no human author; this approach is not recommended or needed. A human will be creatively involved in the creation of a new work – it is unlikely that outputs will be solely generated by a computer.

Infringement and Liability regarding AI

Are there concerns about existing legal tests for demonstrating that an AI-generated work infringes copyright (e.g., AI-generated works including complete reproductions or a substantial part of the works that were used in TDM, licensed or otherwise)?

No – the existing law is sufficient

What are the barriers to determining whether an AI system accessed or copied a specific copyright-protected content when generating an infringing output?

No response

When commercialising AI applications, what measures are businesses taking to mitigate risks of liability for infringing AI-generated works?

It is not likely that AI will output something that infringes copyright, unless the user uses the model to produce a copy of a work. AI developers and providers should not be liable for such use. However AI developers can take steps in the design of AI systems to mitigate the risk of a user using the model to create a copy. 

Should there be greater clarity on where liability lies when AI-generated works infringe existing copyright-protected works?

No – the existing law is sufficient 

Are there approaches in other jurisdictions that could inform a Canadian consideration of this issue?

No response

Comments and Suggestions

N/A

Telecommunications Service Providers - Rogers, Cogeco, Québecor, Bell

Technical Evidence

N/A

Text and Data Mining

N/A

Authorship and Ownership of Works Generated by AI

Should the Government propose any clarification or modification of the copyright ownership and authorship regimes in light of AI-assisted or AI-generated works? If so, how?

Existing Canadian copyright law principles offer protection to works that are the product of skill and judgment by a natural person and achieve a threshold level of originality.

In the context of generative AI, attribution would likely correspond to the natural person who provides instructions to the generative AI system (but not the creator of the AI system itself), for example in the form of a text or visual prompt, and/or who modifies, edits, arranges, or compiles the AI-generated work. However, it should be noted that with the advent of sophisticated generative AI technologies such as ChatGPT and Dall-E, there are, and will continue to be, instances where humans may not contribute sufficient skill and judgment to meet the necessary threshold for authorship/protection. An example of such an occurrence would be where an AI generates work based solely on a single prompt without any further human input, either before or after the work is AI-generated.

Maintaining the existing framework adheres to the fundamental principle of technological neutrality, which provides that the Copyright Act should not be interpreted or applied to favour or discriminate against any particular form of technology.

The government should proceed with care when considering modifications to the existing framework as the practical consequences, and potential unintended consequences, of changing the law with respect to authorship and ownership are not yet well understood.

Are there approaches in other jurisdictions that could inform a Canadian consideration of this issue?

The above noted approach of attributing copyright authorship to the person who makes arrangements for the creation of the work is in line with what has already been adopted in other jurisdictions, for example the United Kingdom or Australia.

Infringement and Liability regarding AI

Should there be greater clarity on where liability lies when AI-generated works infringe existing copyright-protected works?

Given the prevalent role that intermediaries such as telecommunications service providers (“TSPs”) in providing the passive infrastructure that certain AI technology operates on, we feel it is important that the existing, safe harbour protection for Internet intermediaries remain in full effect.

Under the Copyright Act, safe harbour provisions shield intermediaries from obligations or liability in connection with alleged or proven copyright infringement where the intermediaries only provide the technical means by which others infringe. Where TSPs are acting as mere conduits or passive carriers in this way, they must not be required to monitor or be otherwise engaged with the activities of generative AI users or providers to support the protection or enforcement of private rights. The Supreme Court of Canada has already held that Internet Service Providers do not authorize infringement by merely providing connectivity to their users. This stance makes sense because TSPs provide the technical means for communicating and storing content, but their role does not change based on the nature of an end-user’s activity online, even if said activity is infringing or AI-related.

Enforcement obligations on TSPs that would impose monitoring, notification or enforcement obligations on TSPs for AI-related activities would run contrary to the existing structure of Canada’s copyright regime and the direction of the Supreme Court’s previous decisions, which carefully balance protections for rightsholders against the burdens on third party stakeholders. To make such changes would put the vibrant online marketplace and digital services sector in Canada at risk.

The adoption of a framework  that excludes TSPs from the existing copyright safe harbour provisions or imposes obligations on TSPs would increase the cost of the provision of internet services to Canadians, raise concerns about the neutrality of the Internet in Canada, and introduce a disincentive to the massive infrastructure investments necessary to bring the benefits of 5G networks (and future 6G networks) to Canadians and Canada’s economy.  Furthermore, given that AI outputs may well engender innumerable, highly contentious claims of copyright infringement, anything less than an absolute safe harbour for TSPs would threaten an untenable amount of cost and intervention, both of which would have serious negative impacts on Canadians who use digital networks.

Comments and Suggestions

N/A

Annex

Annex: Detailed Questions