Pseudonymize your documents, and include the values in your AI deliverablesPseudonymize your documents consistently, and restore (de-pseudonymize) the original values in your AI assistants' outputs

Marvin Systems CEO

When an AI tool claims to be secure and GDPR-compliant, three arguments almost always come up to reassure users. They are (sometimes) true, they are useful, but they don’t quite cover what we think they cover.
Let’s take five minutes to scrutinize these three arguments and see where the blind spots lie. I find this an enlightening exercise to share because it shifts the conversation to the only arena where we truly remain in control.
This is the most common promise, and it has merit. But it answers a question about what happens after processing, not during it.
Whether or not my data is reused to train future versions of the AI model I’m using is a valid concern. But between the moment I send my request and the moment I receive the response, my data is indeed transmitted, read, and processed by the model.
It passes through the provider’s servers, is loaded into memory, and is processed by the inference infrastructure. The “no training” promise pertains to what happens afterward. It says nothing about what happens during the process.
It is a subtle but fundamental distinction: one cannot expose data to a system and simultaneously consider that it has not been exposed.
Local hosting is a strong selling point, and it addresses legitimate concerns: applicable jurisdiction and technological sovereignty. But it is important to distinguish clearly between what it does and does not cover.
First, a nuance that is often overlooked: hosting in France and a French hosting provider are not synonymous. A foreign company may very well have data centers located on French territory. The data is then stored in France, but the operator remains subject to the laws of its country of origin, which potentially places it under extraterritorial regimes such as the U.S. Cloud Act. “Hosted in France” therefore does not automatically mean “sovereign.”
A second, even more fundamental distinction: hosting and processing are not the same thing.
Hosting refers to where data is stored at rest. Processing refers to where it is actually processed by the model that analyzes it. However, in the vast majority of current AI tools, the model that performs the intellectual work - synthesis, analysis, generation - is not hosted by the provider you pay. It is operated by a third-party provider, often American (OpenAI, Anthropic) or Chinese (DeepSeek), accessible via an API.
In practical terms, this means that requests pass through this third-party provider’s servers during processing, regardless of the country where they are stored before and after.
“Hosted in France” does not mean “processed in France.” And it is precisely the processing stage that raises the most pressing privacy concerns.
Certification is a positive sign. It attests that the vendor has implemented security processes, that they have been audited, and that they are maintained. No reputable vendor should be without it.
That said, it’s worth looking closely at what a certification does and does not attest to. ISO 27001 certifies that an organization has implemented an information security management system: documented processes, governance, and regular risk analysis. This is valuable, but it says nothing, in and of itself, about the location of the hosting, the exact scope of the data covered, or GDPR compliance, which falls under a separate legal framework.
SOC 2, HDS, and ISO 27701 address yet other issues. Each certification has its own scope, and confusing the scope of a certification with a blanket guarantee leads to a false sense of security.
Above all, a certification applies only to the holder, within the scope they directly control. It does not automatically extend to their technical subcontractors.
When a certified software vendor sends a request to an AI model operated by a third party, what happens next falls outside the scope of its certification.
This is what is known as the supply chain. The publisher can follow best practices. The template provider can, in theory, do the same. But “in theory” warrants an important caveat: the GDPR is supposed to apply to any provider offering a service to users in the European Union. In practice, inspections are not systematic, and a provider whose primary business is not focused on the EU does not have the same incentives to invest in compliance as a European player exposed to the risk of inspection.
Overall compliance is the sum of the entire chain, and this chain involves players with very uneven levels of actual legal exposure.
Taken separately, each addresses a legitimate concern. Taken together, they reveal a broader blind spot: they all focus on what happens after the data leaves the trusted perimeter. They describe best practices downstream. But they say nothing about what goes out, or in what form.
Yet that is precisely the only variable you have complete control over.
It is not possible to audit in real time the servers of a model provider located on the other side of the world. Nor is it possible to personally guarantee the compliance of an entire technical subcontracting chain.
On the other hand, it is possible to prioritize solutions that minimize this subcontracting chain and offer an infrastructure designed to limit risks - a zero-data-retention policy, for example. And above all, it is possible to decide what leaves your trusted environment, and in what form.
This shift in perspective changes the very nature of the problem: we stop looking downstream for an absolute guarantee that does not exist, and instead take control upstream over what is actually within our control.
In everyday language, the terms “anonymization” and “pseudonymization” are often used interchangeably. Many publishers perpetuate this confusion. However, these two processes are legally and technically distinct.
Anonymization is irreversible. Once processed, the data can no longer be linked to a person or entity. Under the GDPR (Recital 26), such data falls outside the scope of the regulation. This is the appropriate standard when a document must be permanently removed from the system: archiving, sharing for training purposes, or publishing a dataset.
Pseudonymization is reversible. Identifying elements are replaced by consistent pseudonyms, but a mapping table allows the original values to be reassociated. It is this mechanism that enables the use of an AI tool on a confidential document: the model receives a protected version, produces a usable response, and the original values are reinserted on the business application side.
A tool that offers only one of the two modes forces compromises you shouldn’t have to make. That’s why Marvin Systems offers both: depending on what you’re processing and for what purpose, the trade-offs are different.
None of the above detracts from the validity of the initial arguments. Choosing a certified editor hosted in Europe by a European hosting provider, one that commits to not reusing data for training purposes, remains a sound practice. These criteria are not wrong. They are simply insufficient on their own.
Mastering privacy starts with what you decide to share, and only then with the guarantees offered by those to whom you share it.
This reversal of priorities changes a lot about how we build our daily practices with AI.
This is a subject I’m constantly learning about, and everyone’s experiences add nuances that I wouldn’t have noticed on my own. If your experience contradicts this framework, if you see any blind spots in this reasoning, or if there are situations worth sharing, please feel free to email us and tell us about your approach.