You’re using AI. So is your competitor. So is your supplier, your client, and the contractor you onboarded last quarter. What’s less certain is whether any of you know exactly what’s happening to your data once it leaves your systems.
As artificial intelligence becomes embedded in business operations, the questions that matter are about custody. Where does your data go? Who can access it? Is it being used to train the model you’re paying to use? And when regulations evolve, who carries the liability?
This guide addresses those questions directly, and closes with a due-diligence checklist your team can use today.
What Happens to Your Data Depends on Which Product You’re Using
A primary concern for leveraging AI is the extent to which your data – particularly proprietary or client-sensitive information – is exposed to third-party AI models. The answer depends almost entirely on which tier of a provider’s product you’re using and what your contract actually says.
OpenAI explicitly states that business data submitted through ChatGPT Enterprise, Business, Edu, and Healthcare is not used for training their models by default. Data sent via their API is also generally not used for training. OpenAI’s data retention policy for API usage defaults to 30 days for abuse monitoring, but clients can request Zero Data Retention (ZDR) for eligible endpoints.
Microsoft Azure OpenAI Service ensures that customer data remains within the Azure environment and is not used to train their models. Customers can select specific regions for data storage, which helps address data residency requirements.
AWS Bedrock guarantees that customer data is never shared with model providers and is not used to train foundation models. When fine-tuning models on Bedrock, private copies are created, ensuring proprietary data remains isolated.
Anthropic’s Claude follows a similar pattern. Consumer-tier products may use conversation data to improve models unless users opt out. Their enterprise and API commercial terms, however, include provisions that explicitly prohibit training on client data, consistent with the other providers above.
The common thread: consumer products carry more risk than enterprise agreements. The critical step is confirming which tier you’re actually on, and what the contract says, not what the marketing page implies.
Cross-Border Data Transfers Are the Hidden Compliance Risk
For businesses operating in regulated industries or across multiple jurisdictions, the question of where data is processed matters as much as where it’s stored.
Data sovereignty means that data is subject to the laws of the nation where it’s collected or processed. This becomes complex when AI models are hosted in a different country from where the data originates. A common situation with cloud-based AI services.
Azure and AWS both offer regional endpoints that allow businesses to keep data within specific geographical boundaries (the EU, US, UK, and others) to comply with local data residency laws. However, even with data stored in a specific region, processing by an AI model running in another jurisdiction can trigger cross-border data transfer obligations.
Mechanisms such as Standard Contractual Clauses (SCCs) are commonly used to facilitate legal data transfers across borders. But their validity is subject to ongoing legal scrutiny and evolving international agreements. SCCs are not a set-and-forget solution, they require active monitoring.
Your Privacy Policy May No Longer Reflect Reality
Many privacy policies were drafted before generative AI became a standard business tool. A policy that was accurate 18 months ago may now misrepresent how your organisation actually handles data, not because of bad intent, but because the technology moved faster than the documentation.
If an AI provider’s data practices, such as using data for model training or retaining it beyond expected timeframes, contradict your published privacy policy, that policy becomes a liability rather than a safeguard. Regulators under frameworks like GDPR and CCPA don’t accept “we didn’t know” as a defence.
Privacy policies need to be treated as living documents. Schedule regular reviews that specifically account for your current AI tool stack, and update them whenever you onboard a new provider or change your usage tier with an existing one.
Governance as a Practice
Effective AI data governance requires more than a policy document. It requires consistent operational practice across four areas.
Governance means establishing clear, documented procedures for how AI data is collected, stored, processed, and deleted, and who is responsible for each stage.
Access control means ensuring that only authorised personnel and systems can interact with sensitive data used by AI models. This includes reviewing which employees have access to AI tools that connect to sensitive systems.
Auditability means maintaining comprehensive logs of all data interactions with AI systems. Without this, accountability is impossible and incident response becomes guesswork.
Downstream processing means understanding how data processed by one AI model might flow to subsequent models or connected services. Many AI platforms integrate with third-party tools, each integration is a potential data pathway that needs to be mapped.
AI Due-Diligence Checklist
The following questions should be asked of every AI provider before onboarding, and revisited at each contract renewal.
| Checklist Item | Description |
|---|---|
| Data captured | What exact data classes are captured, including prompts, files, attachments, connectors, logs, and metadata? |
| Retention and training | Does the specific plan or API endpoint use zero retention and no training? |
| Subprocessors and providers | Which subprocessors and model providers can receive the data? |
| Data storage and processing | Where is data stored, processed, and backed up? |
| Cross-border transfers | Can data leave your jurisdiction, and under what transfer mechanism? |
| Data deletion | Is deleted data actually deleted, and on what timeline? |
The Work Starts Before You Need It
The organisations that navigate AI data governance well won’t be the ones with the most sophisticated legal teams. They’ll be the ones that built their governance frameworks before they needed them.
The checklist above isn’t a one-time exercise but a repeating process. AI providers update their terms of service. Regulations evolve. New subprocessors appear in contracts that previously didn’t include them. The questions you ask today need to become the questions you ask every time you onboard a new tool, renew a contract, or expand into a new jurisdiction.
The risks outlined in this guide – data used for model training without consent, cross-border transfers that breach local regulations, privacy policies that no longer reflect actual practice – are predictable. Predictable risks can be managed. But only if the system to manage them is already in place when the situation arises.
Start with the checklist. Build it into your procurement and governance process. Revisit your privacy policy against your current tool stack. Know exactly where your data goes, under what terms, and with what protections in place.
That’s the work. And the time to do it is now.