05/06/2024 - On 23rd May 2024, the European Data Protection Board (EDPB) released its first report on the work being undertaken by the ChatGPT Taskforce (ChatGPT TF). In this post we highlight the ChatGPT TF’s most pertinent views on OpenAI’s popular LLM. Although still preliminary, the possible permutations of the ChatGPT TF’s findings are interesting for any individual or organisation that utilizes LLMs such as ChatGPT during the course of their work.
Consumer-facing LLMs burst onto the scene at the end of 2022, led by OpenAI’s ChatGPT service. In response, several European Supervisory Authorities initiated data protection investigations against OpenAI. At the time, OpenAI did not have an establishment within the European Union (a fact which has since changed). As a result, the normal coordination procedures provided by the GDPR’s One-Stop-Shop mechanism couldn’t apply. This meant that coordination between Supervisory Authorities in their respective investigations against OpenAI would be challenging. To combat this, the EDPB established the ChatGPT TF to foster cooperation and promote information exchange on possible enforcement actions regarding the processing of personal data in the context of ChatGPT.
As explained in the report, investigations conducted by the SAs comprising the ChatGPT TF are still ongoing. This report only represents preliminary views on certain aspects of the investigation.
Below, we have highlighted three portions of the ChatGPT TF’s findings that are of interest to any organisation using LLMs such as ChatGPT in the course of their work.
Lawfulness:
According to the ChatGPT TF, When evaluating the lawfulness of OpenAI’s data processing, it’s useful to distinguish different stages within their processing operation:
Regarding the lawfulness of the first three stages identified above (i-iii), OpenAI relies on their legitimate interest as a data controller as a legal basis. However, considering the large amounts of personal data processed via web scraping and the possibility of processing special categories of personal data, this processing activity presents a peculiar risk to the rights and freedoms of European data subjects. Without going into specifics, the ChatGPT TF has warned that OpenAI will likely need to demonstrate strong technical measures and precise collection criteria (preventing the collection of special categories of personal data) in order to rely on their legitimate interest as a legal basis.
OpenAI also indicated their legitimate interest as the legal basis for stages (iv) & (v) identified above. Prompts refer to the input an end user makes when interacting with ChatGPT (e.g. file uploads, user feedback etc.) Output refers to the AI generated response from the LLM. OpenAI have stated that prompts may also be used to train the model. The ChatGPT TF have indicated that clear and demonstratable notification to end users that their prompts may be used for the purpose of further training the model could be decisive in the balancing of interests (as is necessary when relying on legitimate interest as a legal basis.)
Fairness:
The ChatGPT TF has stated clearly that there can be no risk transfer from OpenAI to data subjects. By data subjects, the ChatGPT TF is referring specifically to end users providing the model with prompts. Regarding ChatGPT, this ensures the responsibility for demonstrating GDPR compliance remains with OpenAI. They would be unable to transfer this responsibility via their Terms and Conditions to the effect that “data subjects are responsible for their chat inputs.”
Transparency:
The ChatGPT TF draws a distinction between OpenAI’s transparency obligations regarding personal data collected via scraping versus personal data provided via user prompts. Due to the sheer volume and complexity of personal data collected via web scraping, it would be impractical and impossible to notify each data subject. Therefore, based on article 14(5)(b) GDPR, it’s possible that OpenAI may be exempt from their GDPR disclosure obligations for that data specifically. However, this exemption does not apply to personal data collected directly from the data subject in their interactions with ChatGPT (prompts). For personal data collected via prompts, OpenAI remains obliged to notify each data subject about what personal data is collected about them and for what purpose.
Data Accuracy:
The ChatGPT TF concedes that the output data from ChatGPT does not necessarily need to be factually accurate. Such an enforcement would limit its capability due to its probabilistic nature. What’s important is the provision, by OpenAI, of proper information on the probabilistic output creation mechanisms and their limited level of reliability. Even if the generated output is not factually accurate, OpenAI will satisfy their data accuracy obligation by making data subjects aware that such output cannot be relied upon as factual.
Although still preliminary, the possible permutations of the ChatGPT TF’s findings are interesting for any individual or organisation that utilizes LLMs such as ChatGPT during the course of their work. As organisations scramble to figure out their data protection obligations regarding LLMs, more guidance from the EDPB will be welcome. Meanwhile, if your organisation has a policy for the deployment of LLMs, it is at least prudent to take into account the aforementioned topics of fairness, transparency and data accuracy.
At Considerati, we have ample experience in the world of data protection and AI regulation. If you need an LLM policy or if you’re curious about what legal or ethical effects the use of AI could have on your business, please reach out to us.
Our services ContactRecente blogs
With great regularity, you hear the term Human-Centric AI (HCAI) or Human-Centered AI (HCAI) being used. During these discussions, it becomes apparent…