Anonymisation

Anonymising data before ChatGPT: methods, limits and good practice

Is replacing a name with a tag enough to “anonymise” it? Not under the GDPR. This guide distinguishes anonymisation from pseudonymisation, lists the data to mask and gives an honest comparison of the methods, from manual review to a tool installed on the computer.

Published on Updated on 9 min readBy the Tacite-IA editorial team

Contents
  1. Why anonymise before sending text to an AI
  2. Anonymisation or pseudonymisation: what the GDPR says
  3. What to mask before using ChatGPT
  4. Four methods and their limits
  5. How Tacite-IA does it, on the computer
  6. Good practice to remember
  7. Frequently asked questions
  8. Sources

“I took the name out, so it’s anonymous.” The sentence is reassuring, but rarely accurate. Before sending a customer letter, a payroll spreadsheet or meeting notes to ChatGPT or another assistant, many employees remove whatever looks sensitive to them. That is a good reflex. But you still need to know what the law means by “anonymous”, which data to mask, and what each method really protects. This guide sets out the facts, with sources, without promising more than the tools can deliver.

Why anonymise before sending text to an AI

An online AI assistant processes the text you type and the files you attach on its provider’s servers. Depending on the plan and the settings, this content is kept for a while and, for consumer services, may be used to improve the models. On the page dedicated to its consumer services, OpenAI itself asks users not to enter sensitive information that they would not want reviewed or used.6 We set out these rules, plan by plan, in our article on what happens to personal data typed into ChatGPT.

For an organisation, the issue is first and foremost a legal one. The GDPR requires personal data to be “adequate, relevant and limited to what is necessary” in relation to the purpose (data minimisation principle, Article 5).1 Yet most tasks given to an AI, such as rewording, translating, summarising or structuring, do not need to know who the person concerned is. France’s data protection authority, the CNIL, also recommends governing the use of generative AI and defining the data that must not be entered into it.7

The message is sent as it is: the name, IBAN and health information reach the vendor, which processes and retains them under the rules of the plan being used.

The same message, sent as it is or after the data has been replaced on the computer. Fictitious data.

Anonymisation or pseudonymisation: what the GDPR says

The two words are often used interchangeably. Yet they do not cover the same thing, and the difference has direct legal consequences.

Pseudonymisation: data that remains personal

The GDPR defines pseudonymisation as processing that means the data can no longer be attributed to a specific person without the use of additional information, provided that this information is kept separately and protected by technical and organisational measures (Article 4, point 5).1 Replacing “Claire Dupont” with “[NAME_1]” while keeping the mapping table somewhere typically falls into this category.

Recital 26 is explicit: pseudonymised data, which could be attributed to a person by the use of additional information, remains information about an identifiable person.1 The GDPR therefore continues to apply. The European Data Protection Board (EDPB) has devoted guidelines to the subject, adopted on 16 January 2025 in a version submitted for public consultation, which describe the value of pseudonymisation as a protective measure.3 The CNIL sums it up unambiguously: pseudonymised data remains personal data, and the operation is reversible, unlike anonymisation.2

Anonymisation: an irreversible transformation

Conversely, the same Recital 26 states that the regulation does not apply to anonymous information, that is, information which does not relate to an identified or identifiable person, or data rendered anonymous in such a way that the person is no longer identifiable.1 To judge this, account must be taken of all the means reasonably likely to be used to identify the person, taking into consideration in particular the cost, the time required and the available technology.1

For the CNIL, anonymisation must make any identification of the person impossible in practice, by any means whatsoever, and irreversibly.2 In their opinion of 10 April 2014 on anonymisation techniques, the European data protection authorities propose three criteria for assessing it: singling out (can a person still be isolated?), linkability (can records concerning them be linked?) and inference (can information about them be deduced?).4

What to mask before using ChatGPT

The list depends on the line of work, but some categories come up everywhere.

  • Identity and contact details: first names and surnames, postal addresses, email addresses, phone numbers, customer account identifiers.
  • National identifiers, starting with the French social security number (NIR): the CNIL points out that it may only be used in very specific cases, most often in connection with social protection.5 The same goes for passport or identity card numbers.
  • Bank details: IBANs, card numbers, direct debit references.
  • Special categories of data under Article 9 of the GDPR: health, political opinions, religious beliefs, trade union membership, sexual orientation, ethnic origin.1 A period of sick leave mentioned in an HR email is one of them.
  • Technical secrets: passwords, API keys, access tokens, connection strings. These are not personal data, but a leak can open access to entire systems. Developers are particularly exposed with coding assistants (see our guide to protecting secrets in Claude Code and Cursor).
  • Context. This is the most common trap. “The finance director of our Lyon subsidiary, on leave since March” identifies a person without naming her. A job title, a town, a date and an event are often enough.
hr-message.txtAnonymised version

Can you draft a polite reply to Ms Sophie Lambert[NAME_1]?

She is asking for her expenses to be refunded to IBAN FR76 3000 6000 0112 3456 7890 189[IBAN_1].

Her social security number: 2 84 05 75 112 345 28[NIR_1], reachable on 06 45 87 12 90[PHONE_1].

4 items replaced on the device
Identifying data is replaced with consistent tags; the request still makes sense to the assistant. Fictitious data.

Four methods and their limits

There are several ways to prepare a text before sending it. None is perfect, and the right choice depends on the volume, the type of data and the level of risk.

1. The manual method

Reread the text and replace each piece of data by hand with a neutral mention. It costs nothing, and an attentive human is good at spotting identifying context, which an automatic tool does poorly. Its limits are well known: it takes time, becomes unrealistic for a file of several hundred lines, and relies entirely on the vigilance of the person at the moment they are in a hurry.

2. Online anonymisation tools

Some websites let you paste a text to get an “anonymised” version of it. The paradox is obvious: to protect data from one online service, you first hand it over to another online service, often with no contract and no information about retention. Before using one, you need to check who publishes it, where the data is processed and what happens to it, exactly as for the AI assistant itself.

3. A tool installed on the computer

A browser extension or a local program analyses the text on the computer, before it leaves. Its main advantage: it steps in at the precise moment of copying and pasting or adding a file, without depending on anyone’s memory. Its limits: it is good at spotting structured data (IBANs, social security numbers, email addresses, card numbers, API keys), but much less so free-form context; and it only protects the tools it actually covers.

4. A locally installed AI

Running a language model on your own machine to spot sensitive data, or to do without an online service altogether, has become possible. The data does not leave the computer. But a model’s output is not deterministic: it may miss a piece of data from one run to the next, depending on its settings, without it always being possible to explain why. It also requires a sufficiently powerful machine, and installing these tools deserves the same scrutiny as any other software.

CriterionManualOnline toolLocal toolLocal AI
Data stays on the computerYesNoYesYes
Works at the moment of copy and pasteNoNoYesNo
Also handles attached filesPartlyPartlyYesPartly
Explainable, reproducible resultPartlyPartlyYesNo
Spots indirect data (context)YesPartlyPartlyPartly
Does not depend on each person’s vigilanceNoNoYesNo
Indicative comparison of the four methods, drawn up by the editorial team. “Partly” indicates a result that depends on the tool or the person.

In practice, these methods are combined. France’s national cybersecurity agency (ANSSI) generally recommends a cautious approach when deploying generative AI and particular attention to the data transmitted.8 A local tool for structured data, a human reread for context and a written rule to settle doubtful cases form a coherent whole.

How Tacite-IA does it, on the computer

Tacite-IA belongs to the third family. The analysis is carried out entirely on the computer, without AI: precise rules and check-digit calculations (for an IBAN, a social security number, a bank card) detect the data, which makes every alert explainable. No message or file is sent to our servers for analysis.

  • For messages, in Chrome and Edge, with ChatGPT, Claude, Copilot, Gemini and other assistants: before sending, a window shows the data detected and offers to send the version in which it is replaced with tags.
  • For files (Word, Excel, PowerPoint, PDF or text) dropped into the assistant: the content is read on the computer, and an anonymised text copy is sent instead of the original. Scanned PDFs, or PDFs with encoded fonts, cannot be read: their content cannot be checked, and Tacite-IA says so.
  • In coding assistants such as Claude Code or Cursor: messages and files read by the assistant are checked, and an anonymised copy is offered.
customer-export.xlsxAnonymised version

customer ; email ; IBAN ; balance due

Mr Marc Petit[NAME_1] ; m.petit@exemple.fr[EMAIL_1] ; FR76 1027 8060 4100 0204 5590 174[IBAN_1] ; €3,200

Ms Inès Roche[NAME_2] ; i.roche@exemple.fr[EMAIL_2] ; FR76 3000 4000 3123 4567 8901 429[IBAN_2] ; €870

6 items replaced on the device
An attached file handled like a message: the columns needed for the calculation remain, the identities go out as tags. Names are detected thanks to the title in front of them (Mr, Ms). Fictitious data.

The limits, stated plainly: like any rule-based tool, Tacite-IA is good at spotting structured data and confidentiality markings, but does not understand identifying context expressed in free form. Names are only detected automatically when preceded by a title (Mr, Ms, Dr…): a name on its own, for example in a spreadsheet column, is not. Text sent with tags generally remains pseudonymised, not anonymous, within the meaning of the GDPR. And desktop apps, such as the ChatGPT desktop app, are not covered: Tacite-IA flags them in the console but cannot control them.

Good practice to remember

  1. Ask whether the data is neededA rewording or a translation works perfectly well without the real name.
  2. Replace, do not just deleteConsistent tags ([NAME_1], [IBAN_1]) keep the text understandable.
  3. Think about filesAn attached export or contract often contains more than the message itself.
  4. Reread for contextJob title, town, date and event can be enough to identify someone.
  5. Keep the mapping on the computerThe table linking tags to values never leaves with the text.
  6. Write the rule downA policy that says what to mask and with which tool.
Six simple habits to include in your AI use policy.

These habits belong in a written rule: our AI policy template includes an article dedicated to prohibited data and anonymisation. For sectors bound by professional secrecy, the requirements go further: see our guides on accountants and professional secrecy and on health data.

Frequently asked questions

Is replacing names with tags enough to anonymise data?

Generally, no. If the mapping between the tags and the real values exists somewhere, or if the context makes it possible to find the person, this is pseudonymisation: under Recital 26 of the GDPR, the data remains personal. It is nevertheless a good minimisation measure before sending text to an AI.

What is the difference between anonymisation and pseudonymisation?

Pseudonymisation prevents data from being attributed to a person without additional information kept separately (Article 4 of the GDPR): the data remains subject to the regulation. Anonymisation makes any re-identification impossible in practice, irreversibly: the GDPR no longer applies.

What data should never be entered into ChatGPT?

In an organisation, you should at the very least remove names and contact details, identifiers such as the social security number, bank details, health data and other sensitive data under Article 9 of the GDPR, as well as passwords and access keys. Context (job title, place, date) can also be enough to identify someone.

Are online anonymisation tools safe?

They require you to hand the data to a third-party service before it is even sent to the AI. Before using one, check its publisher, where the data is processed and how long it is kept. A tool that works on the computer avoids this step.

Is a local AI a good solution for anonymising data?

It keeps the data on the computer, which is a real advantage. But its results are not deterministic: a piece of data may be missed from one run to the next with no explanation. It also requires a powerful machine and scrutiny of the installation.

Sources

  1. Regulation (EU) 2016/679 on the protection of natural persons with regard to the processing of personal data (General Data Protection Regulation). EUR-Lex. Accessed on 5 October 2026.
  2. L’anonymisation de données personnelles. CNIL (French data protection authority), 19 May 2020. In French. Accessed on 8 October 2026.
  3. Guidelines 01/2025 on Pseudonymisation (version for public consultation). European Data Protection Board (EDPB), adopted on 16 January 2025. Accessed on 8 October 2026.
  4. Opinion 05/2014 on Anonymisation Techniques (WP216). Article 29 Data Protection Working Party, 10 April 2014. Accessed on 8 October 2026.
  5. Qui peut me demander mon numéro de sécurité sociale (NIR) ?. CNIL (French data protection authority). In French. Accessed on 8 October 2026.
  6. How OpenAI handles data in consumer services. OpenAI Help Center. Accessed on 5 October 2026.
  7. Les questions-réponses de la CNIL sur l’utilisation d’un système d’IA générative. CNIL (French data protection authority), 18 July 2024. In French. Accessed on 5 October 2026.
  8. Recommandations de sécurité pour un système d’IA générative. ANSSI (France’s national cybersecurity agency), 29 April 2024. In French. Accessed on 5 October 2026.
One measure among others

Govern AI use without slowing your teams down

Tacite-IA detects sensitive data in messages and files before they are sent to AI assistants in Chrome and Edge, and in coding assistants (Claude Code, Cursor, Windsurf, Codex, Gemini CLI, Copilot CLI). Analysis runs 100% on the device, with no AI, and comes with an admin console and an audit mode. Desktop apps (the ChatGPT desktop app, the Chat tab in Claude Desktop, Copilot in Windows) are not covered.

€6 excl. VAT per user per month billed annually, €8 excl. VAT billed monthly. 3-month trial.

Further reading