Skip to content

GDPR and LLM APIs: What Can You Send to OpenAI and Claude?

GDPR and LLM APIs in practice: what you can send to OpenAI and Claude, which DPAs you need, EU data residency options and how to pseudonymize data.

By

Freelance full-stack developer

Published
Reading time
13 min
In this post9

Yes, you can send personal data to OpenAI and Claude under GDPR, provided four things are in place: a clear purpose and legal basis, a data processing agreement (DPA) with the provider, a valid mechanism for transferring data to the US, and a habit of sending only what the task needs. GDPR and LLM APIs are compatible. The way many prototypes call those APIs usually isn't.

I'm a freelance developer based in Denmark, not a lawyer. This is the technical side I work through when adding AI features to apps with European users, not legal advice. Provider terms change often: everything here was checked in October 2026, so check again before you go live.

The short answer: what you can send

What you can send to an LLM API under GDPR, and what it takes
Can you send it?What it takes
Data with no personal information, e.g. product copy or internal docsYesNormal confidentiality. GDPR doesn't apply
Business data with names and emails, e.g. support tickets or CRM notesYes, with the basics coveredPurpose, DPA, transfer mechanism and data minimization
Data you process for your own customers, e.g. in a SaaSYes, if your customers have approved itThe LLM provider must be on your sub-processor list
Pseudonymized dataYesStill personal data, but lower risk
Special category data, e.g. health, ethnicity or union membershipRarelyAn Article 9 condition, an impact assessment (DPIA) and ideally EU processing with zero retention
National ID numbersAvoid itMany EU countries add national rules. Strip them before the call

My rule of thumb: never send more than the specific task needs, and remove direct identifiers (name, email, phone number, ID number) before data leaves your server. It won't solve everything, but it moves most AI features from risky to manageable.

If you're getting an AI-built prototype ready for real users, start with my guide to taking a vibe coded app to production. This post covers one part of that: the data your app sends to a language model.

What OpenAI and Anthropic do with API data

Start by separating the API from the chat apps. API calls fall under the providers' business terms. The consumer versions of ChatGPT and Claude have different terms, and conversations there may be used for training depending on the user's settings. For many companies, the biggest GDPR exposure isn't the product at all. It's an employee pasting a customer email into a personal chatbot account.

Here's how the APIs work as of October 2026:

  • Training: OpenAI's data controls documentation says data sent to the API isn't used to train its models unless you opt in. Anthropic says the same about its API and other commercial products in its privacy center.
  • Abuse monitoring: OpenAI keeps abuse monitoring logs for up to 30 days. Anthropic's privacy center says API inputs and outputs are deleted within 30 days by default, with exceptions such as usage policy violations.
  • Storage you control: Some features keep data longer. OpenAI's Responses API stores responses for 30 days by default, and uploaded files stay until you delete them or they hit an expiry you set. Anthropic has a similar exception for its Files API.
  • Zero data retention: Both providers offer arrangements where content isn't kept for abuse monitoring at all. You have to apply for it or agree on it separately.

So "they don't train on our data" is only half the answer. Your data still sits with a US provider for a while, and your contracts and privacy notice need to say so. You can often reduce it yourself: set store to false in OpenAI's Responses API, delete uploaded files once they've been used, and don't keep conversation history you don't need.

DPAs and transfers to the US

When your app sends personal data to an LLM, you're usually the controller and the provider is your processor. Article 28 GDPR requires a data processing agreement between the two of you. Anthropic incorporates its DPA, including the EU Standard Contractual Clauses, automatically into its commercial terms. OpenAI also offers a DPA for API customers, but check how it has been executed for your account and keep a copy on file.

Both providers are US companies, so you also need a lawful basis for the transfer out of the EU:

  • EU-US Data Privacy Framework. The European Commission adopted an adequacy decision on 10 July 2023 covering US companies certified under the framework. Check the official DPF list to confirm your provider is on it.
  • Standard Contractual Clauses (SCCs). If the provider isn't certified, or you want a fallback, the EU's standard clauses do the job, usually as part of the DPA.

Its predecessor, Privacy Shield, was struck down by the EU Court of Justice in 2020, so keeping SCCs as a backup is sensible even when the provider is certified.

If you're outside the EU: GDPR applies to any company offering services to people in the EU, so a US or UK startup with European customers is in scope. UK GDPR is close in substance but has its own transfer mechanisms.

If you run a SaaS, the provider is your sub-processor

If your customers put their own customers' data into your product, you're the processor and the LLM provider becomes your sub-processor. Your customers need to have authorized that, usually through the sub-processor list in your DPA. Depending on your terms, you either notify them in advance or get their written approval before a new AI feature touches their data. What else a SaaS needs beyond the code is in my honest answer on building a SaaS with AI alone.

Keeping data in the EU: provider by provider

GDPR doesn't require data to stay in the EU, as long as the transfer has a lawful basis. But plenty of B2B buyers, especially in the public sector, healthcare and finance, write EU processing into their vendor requirements. Then the choice of provider is about more than model quality.

EU data residency options for major LLM APIs (October 2026)
EU processing?Watch out for
OpenAI APIYes, for approved customersRequires approval and a modified retention amendment. Region is set per project
Azure OpenAI (Microsoft Foundry)Yes, with Data Zone EU or a European regionGlobal deployments can process data in any Azure region
Claude via Anthropic APINoInference can only be set to US or global
Claude via AWS BedrockYes, with EU inference profilesRegion depends on the endpoint and profile you call
Claude via Google CloudYes, for selected models and regionsRegion depends on the endpoint. Check model availability

Three details catch people out:

  • Azure Global isn't EU. Microsoft's deployment types documentation is explicit: Global deployments may process data in any Azure region, even if your resource sits in Sweden Central. For processing in Europe, pick Data Zone EU or a regional deployment. Global is the default and the cheapest, which is exactly why teams end up there without noticing.
  • Anthropic's own API has no EU region. According to Anthropic's data residency docs, the inference_geo setting only accepts US or global, and data at rest is stored in the US. To run Claude in Europe, go through AWS Bedrock or Google Cloud, where the region is set by the endpoint you call. The cloud provider is then your processor, so your DPA needs to be with AWS or Google.
  • OpenAI's European processing needs an agreement. Your organization has to be approved for abuse monitoring controls and sign a modified retention amendment. You then choose the region when you create a new project.

EU processing has trade-offs. At Microsoft, new models arrive in Global first, then Data Zone, then regional deployments, and regional options can cost more. Factor that into your LLM API cost estimate.

Finally, an EU region at a US provider is still a US provider. For most companies, EU processing plus a DPA is enough. If a customer insists on no US company at all, only a European provider or a self-hosted model meets that.

Send less: data minimization and pseudonymization

The most effective safeguard is sending less personal data. It's also where many AI-built apps go wrong: they send whole database rows or conversations because that was the easiest prompt to write.

In practice:

  • Send fields, not records. If the model is categorizing a support ticket, it needs the message text, not the customer's name, address and order history.
  • Replace identifiers before the call. Swap names, emails and phone numbers for placeholders like CUSTOMER_1, then put the real values back when the response arrives. This happens on your server, so the model never sees them.
  • Clean free text. Emails, phone numbers and ID numbers can be caught with fixed patterns. Names in free text are harder and need a dedicated entity recognition step.
  • Watch your own logs. If your app stores prompts and responses for debugging, your logs now contain personal data. An LLM observability tool that receives those calls is another processor.
  • Keep the system prompt free of personal data. It goes out with every single call, so it should hold instructions, not customer records.

Pseudonymized data that you can re-identify is still personal data under GDPR (Recital 26). It lowers the risk, but you still need the agreements and legal basis. True anonymization, where nobody can trace the person again, is hard with free text. A message from "the owner of the only bakery in a village of 800 people" points to one person even with the name removed.

A 7-step process for a GDPR-ready LLM feature

  1. Map the data flow. Write down which fields are sent, from where in the app, to which provider and region. Include logs, vector databases and monitoring tools, which people tend to forget.
  2. Pin down purpose and legal basis. Why are you sending the data, and which Article 6 basis applies? Check at the same time whether special category data under Article 9 could show up, including in free text your users type.
  3. Pick provider and region based on requirements. If your customers need processing in Europe, choose a setup that delivers it from the start. Switching later costs more.
  4. Get the paperwork in place. DPA, transfer mechanism, your Article 30 record of processing and, if you run a SaaS, your sub-processor list.
  5. Build minimization into the code. Filtering and pseudonymization before the call, storage turned off where possible, and automatic deletion of files and saved conversations. I cover the integration itself in how to integrate an LLM API into an existing app.
  6. Check if you need a DPIA. Article 35 requires a data protection impact assessment when processing is likely to be high risk, for example large-scale special category data or systematic evaluation of people. An expert report on privacy risks and mitigations for LLMs, published through the EDPB's support pool of experts, has a risk method you can borrow.
  7. Tell your users. Update your privacy notice with the purpose, the recipient and the transfer to the US. If the AI feature makes decisions about people without a human involved, read up on the rules for automated decision-making in Article 22.

When not to send personal data to an LLM

Sometimes the right answer is not to:

  • When the task works without it. Product copy, summaries of internal documents and classification of generic requests rarely need personal data. Design the feature so it never receives any.
  • When the data is sensitive and can't be cleaned. Patient records, cases involving children and HR files don't belong with a general-purpose model. A model you host yourself, a European provider with the right agreements, or no AI at all are the realistic options.
  • When you can't explain it to the user. If you can't describe the data flow in two plain sentences in your privacy notice, it's too complicated.

A website chatbot is a classic example. Visitors type names, order numbers and sometimes health details, even when you ask them not to. Plan for that before you decide to build the chatbot yourself or buy one.

You don't always need a developer like me, either. A small internal tool with no personal data is something you can set up yourself. If your question is purely legal, such as whether your legal basis holds up, talk to a privacy lawyer.

Next steps

Before your LLM feature goes live

  • You know exactly which fields are sent to the model, and why.
  • There's a DPA with the provider, and the US transfer has a valid mechanism.
  • The region matches your customers' requirements, and no Azure deployment is set to Global by accident.
  • Names, emails, phone numbers and ID numbers are removed or replaced before the call.
  • Provider-side storage is off or limited, and uploaded files get deleted.
  • Your own logs don't keep prompts containing personal data longer than needed.
  • Your record of processing, privacy notice and sub-processor list are up to date.
  • You've decided whether a DPIA is needed.

If you want help building an AI feature with a data flow that holds up, see how I work on custom web apps and platforms. Take the legal assessment to a qualified advisor, especially if sensitive data is involved.

Frequently asked questions

Can my team paste customer data into ChatGPT?

Only on a business account and with clear rules. Free and personal accounts run on consumer terms, where conversations may be used for training, and there's no DPA with your company. The business versions of ChatGPT and Claude come with DPAs and don't train on your data by default. Write down which types of data are allowed, and keep special category data and ID numbers out regardless of the account.

Do embeddings and vector databases count as personal data?

Yes, treat them as personal data. The source text goes to the provider to be turned into embeddings, and most vector databases store the original text chunks alongside the vectors. When someone asks to be erased, you need to find and delete their data there as well. That's much easier if every chunk is tagged with the customer or user it came from.

Is a self-hosted open-source model the GDPR-safe option?

It can be, because the data never leaves your environment and the US transfer question disappears. In exchange, you take on security, hosting, updates and hardware, and smaller open models are usually less capable than the big ones. GDPR still applies in full to the data you process. It tends to pay off only for sensitive data or strict customer requirements.

How do I handle a deletion request when I use an LLM API?

Delete the data in your own systems first: database, logs, vector store and saved conversations. Abuse monitoring data at the provider expires automatically after the retention period, but anything you stored through the API yourself, such as files or saved responses, has to be deleted actively. Since API data isn't used for training, the user's information doesn't live inside the model.

Does the EU AI Act change any of this?

It adds rules on top of GDPR rather than replacing them. For most business apps that call an LLM API, the practical part is transparency, for example telling people when they're talking to a chatbot rather than a human. Higher-risk uses, such as screening job applicants, come with much heavier obligations. Your GDPR duties stay the same either way.