06/11/2024 – As generative AI continues to advance, it brings complex data privacy challenges, with growing risks of data breaches and unintentional information exposure. From inadvertent data sharing through prompt leaks to vulnerabilities within the models and training data, navigating the AI landscape can be increasingly difficult for organisations. This blog will focus on the most commonly used generative AI models by organizations—large language models (“LLMs”) like ChatGPT and Microsoft Copilot—and explore practical steps to mitigate these risks while ensuring compliance with applicable data protection regulations.
In the past few years, the world has come to see significant advancements in AI through generative technology. This enables machines to produce original content across various formats, including text, images, and audio. Central to this evolution are LLMs, sophisticated algorithms that learn from vast datasets to generate human-like language. Notable examples include OpenAI’s ChatGPT, Google’s Bard, and Microsoft’s Copilot which are widely used for applications ranging from customer support to content generation.
Yet, despite this transformative potential, generative AI models also bring unique data privacy challenges. Key concerns arise from issues such as prompt leakage, which can inadvertently expose sensitive information, and model inversion attacks, where malicious actors may exploit the model to extract confidential data. Additionally, residual data retention becomes particularly problematic when organisations rely on these tools to handle private documents and proprietary information.
As organisations continue to embrace generative AI, understanding these challenges is crucial for ensuring compliance with data protection regulations, such as the EU GDPR. Prompt leakage, model inversion, and residual data retention are just some examples of how LLMs can unintentionally expose personal data.
Prompt leakage refers to the unintended disclosure of (sensitive) information when interacting with an LLM. For instance, an employee might enter confidential or client-specific information to seek assistance with summarising a document or generating ideas. If the model unintentionally retains fragments of these prompts, it can expose sensitive data in later interactions, posing security risks and potential violations of confidentiality.
Model inversion is a method that allows an individual to reconstruct original input data by exploiting a trained AI model’s response. For example, if an organisation develops its own LLM, and the model was trained on anonymous health records or financial transactions, a malicious actor could potentially reconstruct data points about specific individuals, therefore re-identifying them.
Residual data retention refers to the unintentional storage of data remnants after the processing is complete. In the context of generative AI, this means fragments of ordinary or sensitive data might remain within the model even after deletion or anonymisation attempts. For example, a model trained on sensitive documents may “remember” certain information, potentially resurfacing this data when responding to unrelated prompts. This can lead to major compliance challenges, as it complicates organisations’ abilities to meet data minimisation and erasure obligations. This is even more difficult if an organisation uses a proprietary LLM such as ChatGPT, where stored data would not be easily accessed.
All together, these are just some examples of vulnerabilities that raise the likelihood of data breaches, placing organisations using LLMs at increased risk. Addressing them proactively is essential to uphold data security and maintain regulatory compliance when it comes to generative AI applications.
In response to the risks associated with LLMs, some organisations, like Samsung, have resorted to outright bans on generative AI tools following incidents of leaked sensitive information. This approach is not too far off from the actions of some European privacy watchdogs, such as Italy’s data protection authority (“DPA”). Italy was the first European country to block ChatGPT in April last year over GDPR compliance concerns, only to rescind the block a month later after ChatGPT’s creator OpenAI reportedly addressed the concerns.
Such an outright ban to AI tools, including LLMs, may seem drastic, but it underscores the perceived risks tied to the use of AI in corporate settings. Nonetheless, a complete ban on such tools may stifle productivity and limit potential innovation. This should in turn prompt organisations to adopt a balanced approach instead. This could include for example:
All of these controls should serve as practical alternatives to a downright AI ban, allowing organisations to leverage the benefits of generative AI while minimising the associated risks – to the extent that that is feasible.
If you have questions about the responsible use of AI, including your organisation’s use of LLMs, or require assistance in drafting responsible AI use policies and providing comprehensive training, Considerati is here to help. Do not hesitate to connect with us to ensure your organisation embraces AI responsibly and remains compliant with the applicable regulations.
Our services ContactOur blogs
On October 4th, the Court of Justice of the European Union (CJEU) issued a ruling in the long-awaited case between the Royal Dutch Lawn Tennis Association…