Academy
September 23, 2026
LLM Data Privacy: Questions Teams Must Ask
A plain-English LLM data privacy checklist covering data paths, retention, access, regional processing and supplier evidence.
- llm data privacy
- privacy
- procurement

LLM data privacy is not answered by asking whether a model provider “trains on our data”. That is one important question, but it sits inside a larger request path: who sends the prompt, where it is processed, what logs remain, which tools receive the output, who can inspect it, and how the system behaves when a user asks for deletion.
The useful aim is a map that a privacy lead and an engineer can both read. This guide gives you the questions, a small data-flow table and a way to turn a provider's marketing claim into evidence a team can rely on.
Draw the request path before filling in a questionnaire
Take one real feature, such as a support assistant that drafts a reply. Start at the browser, then trace the request through your application, retrieval layer, model gateway, provider, logging system, analytics tools and human review queue. Mark where personal data, customer confidential information and secrets might appear. The surprising parts are often outside the model call: a debug log, a tracing product, a feedback widget, or a tool invoked after generation.
| Stage | Data that may appear | Evidence to collect |
|---|---|---|
| Client to application | User question, identity, attachments | Transport controls and consent context |
| Application to retrieval | Query, tenant, permission filter | Access-control design and corpus scope |
| Gateway to provider | Prompt, system instructions, tool schema | Provider terms and endpoint settings |
| Output and telemetry | Response, tokens, trace IDs, feedback | Retention period and reader roles |
| Human escalation | Full case context | Reviewer policy and audit record |
This map is more valuable than a generic diagram because it names the system owners. A provider cannot answer for a tracing vendor you added later, and a privacy notice cannot substitute for a permission check in a retrieval query.
Ask precise questions about use and retention
For every external service, ask whether API inputs and outputs are used for training, how long abuse-monitoring or operational logs are retained, whether the answer differs by endpoint, and how an approved retention reduction works. Request the current documentation and contract language, then record the exact configuration you use. “Enterprise customers are protected” is not a configuration detail.
For example, the OpenAI data controls documentation explains that handling can vary by endpoint and that some capabilities are incompatible with zero-data-retention arrangements. The practical lesson is broader than one provider: a single vendor-wide answer is not enough. Check the API, region, feature and background mode that your application actually invokes.
Do not promise users “no data is stored” unless you can account for application logs, backups, provider controls and incident records. More often the accurate statement is that data is minimised, protected, retained for a defined purpose and deleted according to documented processes.
Minimise before the model call
The safest personal data is data that never enters the request. Redact obvious identifiers where the task does not need them, replace account numbers with scoped references, keep secrets out of prompts, and retrieve only documents the user can access. For a support classifier, the full email signature may add no value. For a summariser, a customer name may be unnecessary if the job is to extract actions.
Minimisation should not become blind redaction. A system that removes the information needed to make a correct decision can create a different harm. Document why each field is sent, who benefits and how the model's output is verified. The NIST AI RMF is helpful here because privacy-enhanced handling is one of several trustworthiness characteristics, alongside reliability, security and accountability.
Protect access around retrieval and tools
A language model does not understand your authorisation model merely because you described it in prose. Apply tenant and role filters before retrieved text is supplied to the model. Validate tool calls server-side. Keep privileged operations outside a model's default reach, and require confirmation for actions with an external effect.
This guards against both accidental disclosure and adversarial influence. Prompt injection detection explains why untrusted content should not gain authority just because the model has read it. A privacy control that relies on the model faithfully repeating “do not reveal private data” is not a control.
Make rights and incidents operational
Someone should be able to answer: where would we search for this user's data; which logs are in scope; who can approve a deletion; what is excluded for security or legal reasons; and how long will the change take? Build those answers from the data map, not from a one-off incident scramble.
Run a tabletop exercise with one fictional but realistic request: a customer asks for an account export after using an AI feature. Follow the identifiers through application data, provider records, analytics and human-review notes. The exercise exposes missing ownership quickly, without using real personal information in a test.
Next step
Pick one LLM-backed feature and create its first data-flow record this week. Pair it with an LLM audit log so operations can investigate access without retaining raw content by default, and use the AI guardrails guide to cover the input and output risks beside privacy.
Frequently asked questions
Are API prompts always used to train an LLM?
No, but the answer depends on the provider, product tier, endpoint and agreed settings. Obtain the current documentation and contract terms for the exact services you use, then verify your deployed configuration.
Is redaction enough for LLM data privacy?
No. Redaction reduces exposure but does not replace access controls, retention decisions, supplier due diligence, logging discipline or a process for user rights and incidents.
Can we send customer data to an LLM provider?
That is a governance and legal question that depends on the data, purpose, jurisdiction, provider terms and controls. Start with minimisation and a documented data path, then involve the appropriate privacy and legal owners.
What should go in an LLM privacy review?
Include the feature purpose, data categories, every system in the request path, access controls, provider settings, retention, geography, incident process, user-rights process and the accountable owner.