What data cannot be transferred to artificial intelligence: how to use AI without leaking personal and work information
Category
Online.ua Guide
Publication date

What data cannot be transferred to artificial intelligence: how to use AI without leaking personal and work information

What data cannot be transferred to artificial intelligence: how to use AI without leaking personal and work information

It is convenient to ask artificial intelligence to shorten a document, correct a letter, summarize a contract, analyze a table, find errors in the code, or prepare a response to a client. It is in such everyday work tasks that the risk most often arises: a person uploads a complete file to a chatbot, although one paragraph without names, numbers, addresses, amounts, passport data, and details of the transaction is enough for the response.

The problem is not the use of AI itself, but what data is included in the request. In 2025, an analysis of real requests to AI tools, reported by Tele2, showed that sensitive corporate data was present in more than one in five downloaded files. Among them were program code, customer information and internal company documents.

Safe use of AI starts with a simple rule: before sending a text or file, you need to remove everything that is not necessary for the task. A chatbot does not need real names, phone numbers, addresses, bank details, passport numbers, medical data, internal prices, passwords, access tokens, or a complete customer database to edit the style of a letter or create a short summary.

Why uploading files to AI can be a risk

When you paste text into a chatbot or upload a document, you are transferring information to an external digital service. What happens next depends on the specific tool: its privacy policy, account settings, plan type, data storage, administrator access, logging, contractors, integrations, and the ability to use data to improve models.

The UK's National Cyber Security Centre advises against including sensitive information in queries to public large language models and against sending queries that would create problems if publicly disclosed. The NCSC also stresses that you should understand the service's terms of use and privacy policy before working with sensitive information.

For a private user, the risk may look like a leak of a passport number, phone number, address, medical diagnosis, or financial data. For a company, it may look like a disclosure of a customer base, commercial offer, contract, code, internal instructions, strategy, financial plan, or details of negotiations.

A single file may seem secure. But a few small fragments can easily add up to a complete picture: who the client is, what the price is, what the terms are, who makes the decisions, what are the weak points in the contract, what technologies the company uses, and where its data is stored.

What search intent does this topic have?

The intent is mixed: informational, instructional, and security. The user wants to understand what data cannot be given to AI, whether it is dangerous to upload documents to ChatGPT or other chatbots, how to anonymize a file, what distinguishes a public service from a corporate solution, how to work with work materials without leaking, and what mistakes employees most often make.

Therefore, the answer should be practical: not to intimidate AI, but to provide a clear system for verifying data before sending it.

What data should not be transmitted to AI?

The riskiest data is that which could identify a person, give access to money or accounts, reveal a trade secret, violate a non-disclosure agreement, or create legal consequences.

You should not submit to public AI services:

  • passport data, TIN, document numbers;

  • residential addresses, telephone numbers, personal e-mails;

  • bank cards, accounts, IBAN, financial statements;

  • passwords, one-time codes, tokens, API keys;

  • medical tests, diagnoses, medical histories;

  • data of children, relatives, clients, patients, employees;

  • full contracts with real parties, amounts and conditions;

  • client databases, CRM exports, contact lists;

  • commercial offers with prices, margins, discounts;

  • internal policies, instructions, reports, sales plans;

  • closed product code;

  • access keys, configuration files, logs with server addresses;

  • correspondence with lawyers, banks, partners or government agencies;

  • documents under NDA or marked “confidential”.

OWASP identifies the disclosure of sensitive information as a key risk for LLM applications. This information includes personal data, financial details, medical records, confidential business data, security credentials, and legal documents.

Why a “regular document” can also be dangerous

People are often afraid to share passwords or credit card details with AI, but they still upload contracts, business letters, resumes, memos, spreadsheets, meeting minutes, and draft presentations. These files often contain more sensitive information than they realize.

For example, a typical letter to a client might include:

  • the name and position of a specific person;

  • company name;

  • the essence of the problem;

  • internal price;

  • deadline;

  • discount conditions;

  • weak point in negotiations;

  • details of the future deal.

For the task “make the letter more polite”, the chatbot does not need the real name of the client, the full amount, the contract number and the delivery address. You can replace them with neutral designations: “Client A”, “Company B”, “contract amount”, “delivery date”, “document number”.

Tele2 describes this very problem in its own recommendations: a user wants to correct the style of a letter or shorten a document, but copies the entire text, and along with it, names, phone numbers, addresses, financial data, and other information that is not needed for the task enters the system.

Public chatbot and enterprise AI: what's the difference?

A public AI service is a tool that the user uses independently, often from a personal account. The company may not be able to see what files employees upload there, whether data learning is disabled, where history is stored, who has access to the chat, and what third-party extensions are connected.

An enterprise solution usually has more control: user administration, roles, logging, access restrictions, storage policies, data processing agreement, settings for model training, integration with internal systems, and rules for who can work with what data.

OpenAI, for example, notes that by default it does not use data from ChatGPT Enterprise, ChatGPT Business, ChatGPT Edu, ChatGPT for Healthcare, ChatGPT for Teachers, and API Platform — including input and output — to train or improve models. The company also notes that data is encrypted at rest and in transit.

This doesn’t mean that any enterprise AI is automatically secure for everything. The organization still needs to determine what categories of data are allowed to be processed, who has access, what integrations are enabled, how long queries are stored, whether logs are kept, how deletion works, and what happens to files after a task is complete.

“Will my data be used to train the model?” is a valid question, but not the only one.

Many users boil down risk to one question: will the model train on my document? In fact, this is just one possible scenario. Even if the service does not use the data for training, it may be temporarily stored for service operation, security, logging, support, auditing, or account functions.

OpenAI notes in its help materials that in ChatGPT Business, ChatGPT Enterprise, ChatGPT Edu, and API Platform, input and output are not used to train models by default; for personal workspaces, users can disable the use of new conversations to improve models.

Therefore, before working with sensitive materials, you need to check a broader list:

  • whether the content is used for learning;

  • is it possible to turn off training;

  • how long are requests stored;

  • Are chats available to administrators?

  • is there a history of conversations;

  • whether data is transferred to third parties;

  • are there any extensions, plugins, connectors;

  • Is it possible to delete the file and chat?

  • Are there contractual guarantees for the business?

  • Does the service meet the requirements of your organization?

How the “publicity test” works

Before uploading a document to AI, ask yourself: would it be acceptable for an outsider to see this text? If the answer is no, the file should be shortened, anonymized, or not transmitted at all.

The publicness test doesn’t mean that the data will actually become public. It helps you quickly assess the consequences. If a document’s leak could harm a person, company, customer, deal, reputation, finances, or security, that document is not suitable for recklessly uploading to a public chatbot.

Practical formula:

  1. What exactly do I want to get from AI?

  2. What is the minimum fragment needed for this problem?

  3. What data can be removed or replaced?

  4. Does the text contain information from other people?

  5. Do I have the right to transfer this document to a third-party service?

  6. What happens if someone outside my organization sees this snippet?

If in doubt at any stage, it is better to prepare a safe version of the file.

How to anonymize a document before AI

Anonymization is not simply replacing a name with initials. The goal is to remove or mask all fragments that can directly or indirectly identify a person, company, transaction, system, or internal process.

An example of a safer approach:

It was: “Olexander Kovalenko, Financial Director of Alfa LLC, requests a 12% discount under contract No. 45/26 for the supply of equipment by October 20.”

It became: “Client A, a representative of Company B, requests an additional discount under the contract for the supply of equipment by a certain date.”

For style editing, this is enough. AI can make text more polite, structured, and understandable without real names, amounts, numbers, and dates.

What should be replaced:

  • Full name → “Client A”, “Employee B”;

  • company names → “Company X”, “Partner Y”;

  • phones → “[phone]”;

  • addresses → “[address]”;

  • amounts → “[amount]” or conditional range;

  • contract numbers → “[contract number]”;

  • dates → “[date]”, if the exact date is not needed;

  • bank details → “[bank details removed]”;

  • medical data → generalized description without identification;

  • code with secrets → code without tokens, keys, and internal URLs.

The European Commission explains the principle of data minimization as follows: personal data should be adequate, relevant and limited to what is necessary for a specific purpose; by default, companies should process only the data that is necessary for the intended purpose and provide access to it to a limited number of individuals.

Personal data: why you need to be especially careful with it

In Ukraine, personal data is protected by law. The Law “On Personal Data Protection” regulates legal relations related to the protection and processing of personal data and is aimed at protecting the right to privacy. It also states that processing of data about an individual that constitutes confidential information is not allowed without their consent, except in cases specified by law.

For an AI user, this means a simple thing: if you have someone else’s data in a file, you shouldn’t automatically assume that you can transfer it to any online service. This is especially true for HR documents, resumes, medical records, client applications, contracts, document scans, children’s data, financial questionnaires, and internal correspondence.

A safer approach is to work with a generalized version. For example, instead of an employee's actual resume, submit a structure without the name, phone number, address, place of residence, photo, links to profiles, and previous employers unless this information is needed for a specific job.

Working documents: what is especially dangerous to download

In companies, the biggest risk is “shadow AI use”: employees use personal accounts to get work done faster, but don’t check whether security policies allow it. This is how commercial proposals, client letters, contracts, codes, investor presentations, and internal spreadsheets end up in chatbots.

The riskiest work files:

  • contracts to be signed;

  • legal opinions;

  • commercial offers;

  • financial models;

  • sales tables;

  • client bases;

  • personnel documents;

  • internal presentations;

  • product codes;

  • server configurations;

  • log files with tokens;

  • materials under NDA;

  • documents regarding defense, security, medicine, banks, and government systems.

CERT-EU warns that the use of freely available closed AI models can pose risks to sensitive data in queries: users may inadvertently enter confidential or personally identifiable information, and without proper anonymization and protection, such data may be misused or disclosed through a leak.

Code, API keys, and technical files

Developers often use AI for bug detection, refactoring, documentation, and test generation. This is useful, but it is the technical files that can contain the most critical secrets: access keys, tokens, internal server addresses, bucket names, accounts, database structure, CI/CD configurations, and closed product logic.

Before sending the code, you need to remove:

  • API keys;

  • access tokens;

  • private keys;

  • passwords;

  • connection strings;

  • internal URLs;

  • IP addresses;

  • customer names;

  • secrets from .env;

  • private repositories;

  • data from production logs;

  • commercially sensitive algorithms.

It is better to pose an AI problem not through a full project, but through a minimal reproducible example. For example: “Here is a simplified function without real data. Explain why it returns an error.” This approach reduces the risk of leakage and often gives a more accurate answer.

OWASP recommends that LLM applications implement data sanitization, strong input validation, least privilege access control, data source restrictions, tokenization, and redaction of sensitive content before processing.

Separate risk: prompt injection in files and web pages

Not all risks involve the user handing over secrets. There is also prompt injection, a situation where a malicious instruction is hidden in a file, web page, email, image, or other external content that the AI is analyzing. The model may perceive this instruction as part of the task and behave unpredictably.

OWASP explains that indirect prompt injection occurs when an LLM receives data from external sources, such as websites or files, and that content contains instructions that change the behavior of the model. The consequences can include the disclosure of sensitive information, manipulation of responses, unauthorized access to functions, or execution of actions on connected systems.

For an ordinary user, this means: you should not thoughtlessly give AI access to mail, disk, CRM, calendar, repository, or messenger with broad rights. If the tool has agent functions — it can send emails, modify files, create tasks, make database queries — you should restrict rights and require confirmation before important actions.

How companies can organize the safe use of AI

A complete ban on AI often doesn't work: employees will still look for ways to speed up their work. It's better to provide the right tools, rules, and training.

A minimum policy for a company should answer the following questions:

  • which AI services are allowed;

  • what data can be transmitted;

  • what data is prohibited;

  • how to anonymize documents;

  • who coordinates work with client data;

  • is it possible to download files;

  • Are extensions and plugins allowed?

  • who has access to chat history;

  • how requests are saved and deleted;

  • how to check AI answers;

  • who is responsible for an error in the generated text;

  • What to do if you accidentally download a sensitive file.

The NIST AI Risk Management Framework is designed for voluntary use by organizations to help them consider trust, risk, assessment, and governance when developing, using, and evaluating AI systems. In 2024, NIST also released a profile for generative AI that helps organizations identify specific risks to such systems and actions to manage them.

For businesses, it is useful not to simply write “do not transfer confidential data,” but to give examples: what to do with a contract, client letter, resume, spreadsheet, code, financial model, medical document, or internal presentation.

A practical checklist before uploading to AI

Before each request, go through a short check:

  1. Does AI need the entire file?

  2. Can you give a short excerpt?

  3. Does the text contain personal data?

  4. Is there data from other people?

  5. Is there financial information?

  6. Are there commercial terms, prices, margins, discounts?

  7. Are there passwords, tokens, keys, internal links?

  8. Is there medical, legal, or banking data?

  9. Is there an NDA or confidentiality agreement?

  10. Does the company allow the use of this AI service?

  11. Is the use of data for training turned off, if necessary?

  12. Is it possible to complete the task using an impersonal example?

If even one point is questionable, you need to edit the document to a secure version or use a corporate tool with proper terms.

How to correctly formulate a request without disclosing data

The best query to an AI is one that has enough context to respond, but no unnecessary sensitive details.

Instead of: “Here is a contract with LLC “Romashka” for 3.8 million UAH, check the risks.”

Better: “Here is an impersonal fragment of the supply contract. Check for logical risks, unclear wording, conflicting terms, and places where you should contact a lawyer.”

Instead of: “Here is my child’s medical test with full name and date of birth, explain the result.”

Better: “Explain in general terms what these numbers might mean without making a diagnosis. I will see a doctor for a medical decision.”

Instead of: “Here is the CRM export with customers, find the segments.”

Better: “Here is an impersonal table without names, phone numbers, email addresses, and company names. Analyze the structure of segments by conditional parameters.”

This approach retains the benefits of AI but reduces the amount of data transmitted.

What to do if a sensitive file has already been downloaded

If you accidentally upload a document containing personal, financial, or work information to an AI, don't just close the tab and forget about it. You need to act based on the level of risk.

Procedure:

  1. Delete the file and chat if the service allows it.

  2. Check your settings for using data for training.

  3. Change passwords, tokens, or API keys if they were included in the request.

  4. Notify the responsible person in the company if this is a working document.

  5. Record exactly what was transmitted.

  6. Assess whether customer, employee, partner, or patient data was there.

  7. Check the service's data deletion and retention policy.

  8. Do not upload the same file again.

  9. Prepare an impersonal version for further work.

  10. If the risk is high, involve a cybersecurity professional or lawyer.

If the request included access keys, they should be considered compromised. It is not enough to delete the chat - the keys should be revoked and new ones issued.

Common user mistakes

The first mistake is to upload the entire document instead of the desired excerpt. For most tasks, a portion of the text is sufficient.

The second mistake is to leave names, phone numbers, addresses, amounts, details, and document numbers where they do not affect the answer.

The third mistake is using a personal account for work documents. This creates an uncontrolled leakage channel for the company.

The fourth mistake is to think that “it’s just a draft.” Drafts often contain the most valuable details: the parties’ positions, weaknesses, internal comments, edits, prices, and risks.

The fifth mistake is to pass code along with secrets. An API key in a code snippet can be more dangerous than the code itself.

The sixth mistake is to mindlessly connect AI to email, disk, or CRM. The broader the rights of the tool, the greater the consequences of the mistake.

The seventh mistake is not checking the answer. AI can make mistakes, invent facts, misinterpret a document, or suggest legally risky wording.

The eighth mistake is not reading the service policy. Different tools have different rules for storage, learning, deletion, data access, and use of third-party providers.

Conclusion

AI can be a useful assistant for texts, spreadsheets, code, document analysis, and everyday work tasks. But it cannot be used as a bottomless box, where everything is uploaded without preparation. The most common risk arises not from a sophisticated hacker attack, but from a common user action: copying an entire file without removing unnecessary data.

Safe practices are simple: before sending, keep only what is necessary for the task, remove personal and financial data, replace real names and companies with pseudonyms, do not transfer passwords, keys, customer databases, internal documents and materials under NDA. Use approved tools with proper access, storage and model training rules for working with corporate information.

The main test before the request: would it be a shame if this fragment were seen by an outsider? If so, the document should be depersonalized, shortened, or not transmitted to the public AI service at all.

FAQ

What data cannot be transferred to artificial intelligence?

Do not share passport data, TIN, bank details, passwords, API keys, medical information, addresses, phone numbers, children's data, customer databases, contracts with real parties, internal documents, commercial terms, closed code and materials under NDA. OWASP categorizes such information as sensitive information, the disclosure of which may create privacy, financial, legal and business risks.

Is it safe to upload documents to ChatGPT or another chatbot?

This depends on the account type, settings, service policies and document content. For public LLMs, the NCSC advises not to include sensitive information in requests or send anything that would cause problems if disclosed.

Are my queries used to train AI?

Terms vary by service. OpenAI notes that business products like ChatGPT Business, Enterprise, Edu, and API do not use input and output to train models by default, and personal workspace users can disable the use of new conversations to improve models.

Is it enough to simply remove the name from the document?

No. You need to remove all identifiers: phone numbers, addresses, mail addresses, document numbers, details, exact dates, company names, contract numbers, amounts, unique circumstances, and other details by which a person or transaction can be identified.

Can AI transfer working documents?

Only within company policy and with confidentiality in mind. If the document contains customer data, pricing, contracts, internal comments, code, business strategy, or employee personal data, it must be depersonalized or an approved corporate AI tool must be used.

Why is it dangerous to upload a customer database?

A customer database contains personal data, commercial information, interaction history, financial potential and business structure. Transferring it to a third-party service without proper grounds can create a risk of leakage, breach of contracts, reputational damage and problems with personal data protection.

Can I give the AI code?

You can work with simplified snippets without secrets. Before sending, you need to remove API keys, tokens, passwords, private URLs, IP addresses, internal system names, client data, production logs, and commercially sensitive logic.

What is prompt injection?

Prompt injection is an attempt to change the behavior of an AI through special instructions in a query or external content. OWASP explains that indirect prompt injection can be hidden in files, websites, or images and lead to the disclosure of sensitive information, manipulation of responses, or unauthorized actions on connected systems.

How to anonymize a file before uploading it to AI?

Leave only the information necessary for the task. Replace real names with “Client A”, companies with “Company X”, amounts with “[amount]”, phone numbers and addresses with labels, remove details, document numbers, passwords, tokens, internal links and unnecessary details.

What should I do if I accidentally downloaded a confidential file?

Delete the chat and file if possible, check privacy settings, change passwords and revoke tokens if they were in the document, notify a responsible person in the company and record what data was transferred. If the risk is high, it is worth involving a cybersecurity specialist or lawyer.


By staying online, you consent to the use of cookies files, which help us make your stay here even better 

Based on your browser and language settings, you might prefer the English version of our website. Would you like to switch?