# Shadow AI: why your staff paste company data

> Shadow AI is not recklessness and not a training problem. People paste because the assistant cannot see the systems the answer lives in, and pasting is the only bridge they have.

**TL;DR** Surveys keep finding the same thing: most employees paste work data into AI tools, a meaningful slice of it is sensitive, and most organizations cannot see any of it. Banning the tools does not work because the tools are useful. The cheaper fix is to remove the reason for the paste by letting the assistant read the systems the answer already lives in, under the same access the person already has.

Every few months another survey lands with the same shadow AI finding. Most employees are putting work data into AI tools. A meaningful fraction of what they paste is sensitive. Most organizations have no visibility into any of it.

The numbers vary by study and by methodology, and they have moved quickly. Cyberhaven's [2026 AI Adoption and Risk Report](https://www.cyberhaven.com/blog/sensitive-data-flowing-into-ai-tools) found that 39.7% of AI interactions involve sensitive data, and that the average employee puts sensitive data into an AI tool once every three days. The same research puts a majority of some assistants' usage on personal accounts rather than company ones, which is the part that makes it invisible.

IBM's [2025 Cost of a Data Breach report](https://www.ibm.com/think/x-force/2025-cost-of-a-data-breach-navigating-ai) puts the consequence in money: 20% of the organizations studied reported a breach involving shadow AI, adding roughly $670,000 to the average cost. The finding worth sitting with is a different one from the same report: of the organizations that suffered an AI-related breach, 97% said they lacked proper AI access controls.

The response is usually a policy, a training module and a block list. None of which works well, because all three are aimed at the wrong thing.

## Shadow AI is not recklessness

Watch what actually happens.

Someone in support has a customer complaining that an invoice is wrong. The answer requires the customer's order history, the invoice, and the notes from the last conversation. Those live in three systems. The assistant on their second monitor is very good at reading three things and explaining the discrepancy.

So they copy the order history. They copy the invoice. They copy the notes. They paste.

That person is not confused about whether customer data is sensitive. They know. They are making a choice between doing their job well and following a policy, and they are making it under time pressure with a customer waiting. They chose the customer. Most people do.

The cause is mechanical, not cultural: **the assistant cannot see the systems the answer lives in, and pasting is the only bridge available.**

This is why the three standard responses underperform. A policy tells people not to use the bridge without giving them another one. Training makes them feel worse about using it. A block list moves them to a personal device where you cannot see the bridge at all. Notice that the block list makes your actual position worse, because your view of the behavior disappears while the behavior does not.

## What it costs, in the order people care about

**The data is gone in a way you cannot characterize.** Not necessarily leaked, but out of your control and outside your retention, deletion and residency arrangements. The important word is *characterize*. If somebody asks what left, the honest answer is that you do not know, and "we do not know" is the answer that turns a small incident into a large one.

**The audit answer does not exist.** Auditors are asking about AI access now, and the question has a shape: which systems could an assistant reach, who could ask it to, and what did it do. If access happened by copy and paste, there is no log to produce. Not an incomplete log. None.

**Offboarding silently misses it.** Your leaver checklist covers the accounts IT knows about. It does not cover the assistant somebody connected to the CRM themselves, because nobody wrote it down. The person leaves, the account is disabled, and the assistant connection is still there because it was never on the list.

**The work is duplicated badly.** The quieter cost. Three people build three private workflows for the same job, each with their own pasting habits, none reviewable, none shareable, and all of them lost when those people change teams.

## The fix is to remove the reason

If people paste because the assistant cannot see the systems, then the durable fix is to let it see them, under conditions you set.

That is a different project from a policy. Three properties make it work:

**Connect the systems once, at the company level.** Not each person connecting each app under their own account, which is how you get twenty connections nobody has a list of. The person who owns the helpdesk connects the helpdesk. Once.

**Let the access follow the person, not the connection.** This is the part that decides whether the whole thing is safe. If access is a property of the connection, then whoever set it up decided what everyone gets, and the support agent and the finance manager see the same tools. If access resolves against the person the assistant is acting for, then the connection is shared while the reach is not, and the support agent sees support tools.

**Keep the log.** Every tool call recorded, naming the operation, the specific account it reached, and the outcome. This is the artifact that makes the difference between "we allow AI access" and "we can tell you what the AI did", and it is the one an auditor will ask for.

With those three in place, the support agent in the example asks their question and the assistant reads the three systems directly. The customer data never leaves the boundary. The activity is logged. And when that person leaves, removing their account removes their reach, because their reach was never anything but their account.

## What the sanctioned path has to beat

There is a version of this project that fails, and it fails for a predictable
reason: the approved tool is worse than the unapproved one, so people keep using
the unapproved one and now you have paid for both.

The bar is set by what people are doing today, which is pasting into a very good
general assistant. Three things have to hold for the sanctioned path to win.

**It has to be the client they already like.** A separate portal with its own
login is a tool people visit when they remember to. The assistant they already
have open all day is where the work happens. If the approved path means changing
which application people use, the approval is theoretical.

**Adding the next system has to be cheap.** The first three connections will not
cover everything. If system four takes a quarter and a project plan, the gap gets
filled by pasting again, and the habit you were replacing quietly reasserts
itself. Connecting one more system should be an afternoon for the person who
administers it.

**It has to be faster than the tab.** The reason someone pastes an order history
is that reading it in the assistant beats reading it in the CRM. If the governed
version is slower or returns less, they will keep the tab. This is mostly a
matter of connecting the systems people actually use rather than the ones on the
architecture diagram.

None of that is about governance, which is the point. The governance is what you
want; the usefulness is what buys you the right to have it. A project that leads
with the controls and treats the usefulness as a detail produces a tool with
excellent logging that nobody uses, and the pasting continues where you cannot
see it.

## What to do on Monday

**Find out what is already happening.** Ask, in a way that does not punish the answer, which AI tools people are using and what for. The list of systems they name is your connection priority list, already ranked by demand, at no cost.

**Connect the top three.** Not everything. The three systems that came up most often, connected once, shared with the teams that need them. Most of the pasting volume tends to sit in a small number of places.

**Set the access rules before you announce it.** Which roles reach which systems, and which operations they can run. Doing this after rollout means widening from a default, which is harder than narrowing from a considered starting point.

**Look at the log after a week.** It will show you what people are actually doing, which is reliably different from what they said in step one. That difference is where the next three connections come from.

## The framing that gets it funded

Shadow AI is usually presented to a budget holder as a risk item, and it competes badly in that category against risks with more dramatic failure modes.

It goes better framed as what it is. Your staff have found a tool that makes them measurably better at their jobs. They are using it in the only way currently available to them, which happens to be the worst way. The project is not to stop them. It is to give them a better way and get the logging as a side effect.

That version gets approved, because it is true and because it does not require anyone to believe their colleagues are careless.

---

The security lead's version of this problem, which is a different conversation with different evidence, is in [the shadow AI problem, for the person who has to answer for it](/blog/shadow-ai-security-lead/). What the access rules look like in practice is in [designing the access rules](/blog/least-privilege-for-ai-agents/), and how to choose between the products that do this is in [MCP gateways compared](/blog/mcp-gateway-comparison/). If you already know which systems you would connect first, the [connector catalog](/connectors/) will tell you what each one exposes.

## FAQ

### How common is this, actually?

Consistently high across independent studies. Research on enterprise clipboard activity has found a large majority of employees pasting work data into AI tools, with roughly a tenth of that content classified as sensitive. Separately, around two thirds of organizations report suspecting or confirming the use of AI tools they have not sanctioned. The numbers move between studies; the shape does not.

### Would blocking the tools solve it?

Rarely, and it tends to make visibility worse. Blocking moves the behavior to personal devices and personal accounts, where you can no longer see it at all. The activity continues because it is genuinely useful; what disappears is your view of it.

### Is this not just a training problem?

Training helps at the margins and does not address the cause. The person pasting a customer record is not confused about whether customer records are sensitive. They are choosing between a useful answer and a policy, with no third option available. Training raises the discomfort of the choice without changing what the choices are.

### What does it cost when this goes wrong?

Published breach research has put the incremental cost of incidents involving unsanctioned AI in the region of several hundred thousand dollars on top of an already expensive average, affecting roughly one organization in five. Treat the specific figures as indicative rather than precise, but the direction is well established.

### What is the smallest useful first step?

Find out what people are already using and why. Not to enforce anything, but because the systems they name are the shortlist for what to connect first. The tools people paste into are a free, accurate survey of where the answers they need actually live.
