AI safety: what information should you never share with AI?

AI safety: what information should you never share with AI?
Table of content

Careful control over what you enter is essential when using consumer versions of language models: information submitted to public AI tools leaves your local environment. Entering confidential information into AI risks a data leak, the disclosure of trade secrets and a breach of data protection law. Treat every interaction with consumer AI as though you were publishing it publicly.

What should you never paste into AI?

Never enter the following categories of information into publicly available AI tools:

  • personal data – your own or other people’s names, addresses, National Insurance numbers, phone numbers or email addresses, including those of customers and employees;
  • medical records – patient histories, test results and diagnoses, which are subject to strict legal requirements such as the UK GDPR and HIPAA, where applicable;
  • confidential business information – unpublished financial reports, marketing strategies, customer databases and plans for mergers, acquisitions or redundancies;
  • trade secrets – software source code, product formulas, proprietary algorithms and internal technical documentation;
  • passwords and credentials – API keys, login details, authentication tokens, payment card details and encryption keys;
  • copyrighted material – complete books, paywalled academic articles or scripts that you are not licensed to share;
  • illegal content – material promoting violence or hatred, or instructions for unlawful activities.

What happens to information you enter into AI?

Information sent to cloud service providers can pass through several stages of processing and analysis. Its lifecycle does not necessarily end when an answer is generated.

Model training

Information entered into standard consumer versions of generative AI tools is often used to improve their models. This means it could influence a neural network’s weights and, in theory, be reproduced in a response to another user in the future.

Human review and moderation

Alongside automated processing, AI systems use safety filters. If an algorithm flags a prompt as suspicious, for example because it may breach the terms of service, the conversation may be referred for human review. The provider’s staff or external reviewers, sometimes called data labelers, may then see what was entered while working to improve moderation.

Data storage and backups

AI providers store conversation records on their servers to maintain their services, diagnose faults and preserve context for future sessions. Deleting a chat from your account rarely means it is immediately removed from every server. Copies may remain in the provider’s backup systems for a period of time, increasing the risk that confidential information could be exposed if its cloud infrastructure is compromised.

Personal and business accounts: how does protection differ?

The level of protection an AI service offers varies considerably by subscription and how it is deployed.

Free and paid standard personal accounts on many popular services may use submitted information for model training by default. If confidential material is shared, each session therefore carries a risk of exposing it beyond your control.

Business and enterprise environments, such as Enterprise accounts and services deployed in private cloud environments through providers’ APIs on Microsoft Azure, AWS or Google Cloud, typically offer stricter safeguards. Providers may commit to standards such as ISO 27001 and SOC 2, compliance with the UK GDPR, and contractual terms in SLAs and DPAs stating that customer data will not be used to train public foundation models. Rules about what can be entered may be less restrictive in these environments, but risk management policies and access controls are still essential.

What can you safely paste into AI?

Despite these restrictions, there is plenty of material you can generally enter into AI tools. Examples include:

  • publicly available material – open-access articles, blog posts, public market reports, legislation and website content, such as Wikipedia pages;
  • non-confidential information – general instructions, questions about theoretical concepts, grammar and spelling rules, mathematical equations and word definitions;
  • anonymised extracts from documents – letter templates, sample contracts or code snippets with identifying variables, keys, customer names and details of internal architecture completely removed;
  • your own creative drafts – draft emails, social media posts, presentation outlines and articles that contain no sensitive business information;
  • open datasets – synthetic data or statistics from public repositories used to practise analysis and formatting.

How can you anonymise messages and data before sending them to AI?

Before sending confidential documents to a language model, prepare them by removing sensitive information. This lets you work with the text while reducing the risk of disclosing confidential data.

Replacing identifiers with placeholders

Here, tokenisation means replacing sensitive details with artificial placeholders. For example, you could replace “John Smith” with “[Customer_1]” or “Acme Ltd” with “[Organisation_A]”. This preserves the document’s structure, grammar and context for the language model while making it harder to identify the original person or organisation.

Masking sensitive information

Masking hides critical strings of characters directly, often by replacing them with symbols, as in a card number shown as 4532 **** **** ****. Generalising figures is another useful approach: instead of giving an exact monthly salary of £2,800, you could write “a monthly salary of £2,500–£3,500”. This helps reduce the risk of re-identification.

Automating the process with DLP tools and RegEx

Removing information manually from long documents makes it easy to miss something. Professional teams use scripts based on regular expressions (RegEx) or DLP (Data Loss Prevention) systems. These tools can scan a prompt before it is sent and detect, block or replace strings matching formats such as National Insurance numbers, VAT numbers, email addresses and bank account numbers.

How can you stop AI providers training models on your data?

Configuring the tool itself can also reduce the risks of sharing information with AI. Most providers offer a way to opt out of having your content used for training.

  • ChatGPT (OpenAI) – in your account’s Settings, open Data controls and turn off “Improve the model for everyone”. This stops new conversations from being used to improve the model without removing them from your visible chat history. OpenAI may still retain certain logs temporarily for safety and abuse monitoring; check its current retention policy for the feature you use.
  • Gemini (Google) – turn off “Gemini Apps Activity” in your privacy settings. New conversations will not be saved to your Google Account or used to train models through that setting. Google may still retain them for up to 72 hours for service safety and security.
  • Claude (Anthropic) – check the privacy settings for your consumer account to see whether your conversations can be used to improve the service. To opt out, turn off “Help improve Claude”. This differs from Anthropic’s earlier approach, when it said it did not use consumer conversations for training by default.

Remember that changing privacy settings or opting out of training generally affects future use only. It does not necessarily undo any use of information you submitted before making the change.

Summary

  • Treat interactions with consumer AI tools as potentially public. Do not enter personal data, medical records, passwords or trade secrets.
  • Submitted information may be used for model training, reviewed during moderation and stored on external servers or in backups.
  • Before using AI with documents, remove identifiers, mask sensitive strings and consider automated DLP tools.
  • Enterprise accounts and dedicated business cloud environments can offer stronger safeguards and contractual commitments not to use customer data for training.
  • Privacy settings and training opt-outs generally apply to future submissions; they do not necessarily reverse earlier use of your data.

Leave a Reply

Your email address will not be published. Required fields are marked *

Blog

Recent articles

29.09.2026 Copywriting
28.09.2026 Tips & Curiosities
26.09.2026 Copywriting
25.09.2026 Copywriting
24.09.2026 SEO
23.09.2026 Copywriting
19.09.2026 Copywriting
18.09.2026 SEO
16.09.2026 Content Marketing

Professional business content

Order texts

Build a career with Content Writer

Career

A practical
course
in copywriting