Shadow AI data risk: what your team pastes into ChatGPT
A 175-person healthcare support firm in Hyderabad answered their client’s security questionnaire with one word. None. Two weeks of browser data said fourteen.

Shadow AI data risk is a copy-and-paste problem, not a hacking one. Your people are pasting customer records, contract text and client documents into AI chat boxes, mostly on personal logins, and most company controls are watching file uploads instead. Blocking one website moves that traffic somewhere else. It does not stop it.
The figure worth taking to your next budget conversation comes from IBM’s Cost of a Data Breach Report 2025. One in five breached organisations had a security incident involving shadow AI, high levels of it added around USD 670,000 to the average breach cost, and 97 per cent of the organisations that suffered an AI-related incident had no AI access controls in place at all.
Indian law already put this duty on you. A browser tab is just where it now gets tested.
11:40 AM on a Tuesday in May, and the question was number 34 on a vendor security questionnaire.
List every generative AI service used by your staff in the course of delivering our work. The client was a United States health payer. The firm answering was a 175-person medical coding and claims support company in Hyderabad, eleven years old, profitable, careful. Their operations head had typed one word into the box.
None.
He believed it when he wrote it. That is the part I keep coming back to.
The block that was not blocking anything
Their reason for writing None was real. Fourteen months earlier they had blocked the main AI chat website at the firewall. Somebody had read an article, raised it in a management meeting, and the request went to their managed firewall vendor. One domain, one rule, done in an afternoon.
We ran a two-week visibility exercise before answering the questionnaire properly. Browser-level telemetry on 40 laptops, no blocking, nothing announced beyond what the acceptable use policy already said. It was meant to confirm what everyone expected.
Fourteen distinct AI destinations came back.
The blocked one was not the busiest. What people had drifted to was a mixture of a free summarising site with an unfamiliar registrar, two browser extensions that offered rewriting inside any text box, a model router nobody in the management team had heard of, and one very ordinary alternative that simply was not on the block list because it had not existed when the rule was written.
Some of it was on personal phones over the guest wifi, which the firewall does not see in any useful way. Most of it was not. Most of it was on company laptops, in a company browser, on a personal login.
Nobody was hiding. Two of the medical coders showed me their workflow without being asked, because they were proud of it. Paste a discharge summary, get a suggested code set and a plain-English rationale back, check it, book it. It had taken their per-file handling time down by a real margin. From their seat, they had solved a productivity problem their employer had been complaining about for two quarters.
Arre, and they were right. That is why banning is such a weak instrument here. You are not fighting laziness. You are fighting a tool that genuinely works, in the hands of people being measured on throughput.
The monitoring was watching the wrong verb
The second surprise was worse, because they had paid for the thing that should have caught it.
They ran an endpoint agent on every machine. It watched file uploads, USB writes and email attachments, and it watched them properly. In the same fortnight it correctly flagged a coder attaching a spreadsheet to a personal webmail draft, and that alert worked exactly as designed.
A paste into a browser text box is not a file. There is no attachment, no upload event, no filename, no size. The clipboard goes into a form field, the form field goes out over normal encrypted web traffic to a legitimate domain that resolves cleanly and has a valid certificate, and every layer of the stack sees a person visiting a website. This is not a product defect. Most data loss tooling was built around files because for twenty-five years the data moved as files.
So the discharge summaries left as text. Names, ages, dates of service, diagnoses. Not one of them registered anywhere as an event.
An employee signed into a company AI account leaves you a record you can produce during an audit, a retention setting you control, and a contract that says your data is not used to train anything. The same employee signed into their own free account on the same laptop leaves you none of those, and the account survives their resignation. The tool is rarely the risk. The identity attached to it usually is.
What each control actually watches
We put this table on a whiteboard in their operations room, and I have redrawn it for eight clients since. It is deliberately unflattering to controls that companies think of as finished.
| Control | What it sees | What it misses | Effort to close the gap |
|---|---|---|---|
| Domain or DNS block on one AI site | Direct visits to that one address from the office network | Every other AI site, browser extensions, personal phones, mobile data, home wifi | Already done, and worth roughly what it cost |
| File-upload and USB monitoring | Attachments, uploads, removable media, print jobs | Clipboard pastes into a browser, which is how prompts are actually fed | Configuration change, not a new purchase, on most modern suites |
| Browser-level paste inspection | The text itself, before it leaves the machine, with a warn or block choice | Anything on a device you do not manage | Two to three weeks of tuning, mostly to stop false alarms |
| Corporate account enforcement on AI tools | Who used what, retention, training exclusion, an exportable log | Tools your staff have not told you about yet | Licence cost plus a sanctioned-tool decision somebody has to own |
| Browser extension inventory | What is reading and rewriting inside every page your staff open | Extensions installed on personal profiles | An afternoon, and it is the most under-run check on this list |
| Device enrolment and certified wipe on exit | Which machine is whose, and whether it came back clean | Data already sent somewhere else | Part of how you buy and retire laptops, not a security project |
Row three is the one people assume needs new software. If you already hold Microsoft 365 at the higher tiers, endpoint data loss prevention can warn or block a user who pastes sensitive text into a third-party AI site in the browser, which Microsoft documents in its Purview guidance for AI apps. Plenty of Indian companies are paying for that licence and have never switched the policy on.
The bit I got wrong
I want to be honest about a mistake, because the mistake is the lesson.
I recommended that firewall block. Not in a formal engagement, in a twenty-minute phone call in early 2025, when their operations head asked whether he should be worried about staff using AI. Block the main one, write a one-page acceptable use note, revisit in a quarter. Then I closed the call and moved on to a client with a live audit date.
Nobody revisited it. Neither did I. And the block did something worse than nothing, because it produced a feeling of closure. For fourteen months, every time the subject came up in that building, somebody said it is blocked, and the conversation ended there. I had handed them a sentence that stopped people asking the question.
Bas, that is the shape of the error, and it is not really about AI. A control you cannot observe is a belief, not a control. I now refuse to recommend a block without a way of seeing what moved after it. If the client does not want the seeing part, I would rather they had no rule and knew they had no rule.
Where Indian law actually sits on this
The instinct in most rooms is to reach for the penalty number. I would put it second, because for a firm like this one it is the less urgent of the two exposures.
The first exposure is the contract. That questionnaire was a condition of a services agreement with a client who could terminate on a finding. A false answer to question 34 was a bigger problem in May 2026 than any regulator was, and that is true for most Indian companies delivering work for larger customers. Your client’s security team will find this before the Data Protection Board does.
The second is the statute, and it is not new law. The Digital Personal Data Protection Act, 2023 makes you responsible for reasonable security safeguards over personal data you hold, whoever moved it and however sincerely they meant well. Pasting a customer record into an AI chat box is processing personal data and disclosing it to a third party. The Act’s Schedule prices a safeguards failure at up to ₹250 crore, the operating detail sits in the Digital Personal Data Protection Rules, 2025, and substantive obligations commence on 13 May 2027.
Two points people get backwards. Section 16 does not currently forbid sending personal data outside India, because the negative list it contemplates has not been notified, so the country your AI vendor runs in is not by itself the problem. And Section 7(i) treats employment as a legitimate use, which means you do not need an employee’s consent to watch data movement on a company device for employment purposes, including protecting the company from loss and keeping trade secrets confidential. Section 7 is a closed list, so scope the watching to those purposes, write the scope down, and say it out loud at induction rather than burying it in an appointment letter.
You are not deciding whether to allow AI. Your staff decided that eighteen months ago. You are deciding whether the version they use is one you can see, log and produce on request, or one that lives on a personal login you will lose access to the day they resign.
Shadow AI data risk, handled without buying anything
Most of what went wrong in Hyderabad cost nothing to fix. I want to keep that separate from the part where we sell something, because people in my line of work blur the two.
Sanction one tool, properly, on company accounts. This is the step firms skip and it is the one that does the most. A permitted, logged, contractually protected option removes the reason to go looking. Ban everything and you get fourteen destinations and no visibility, which is where they started.
Run the browser extension inventory this week. It takes an afternoon and it is consistently the ugliest finding. An extension with permission to read and change data on all sites is sitting inside every page your finance and coding teams open.
Turn on the paste policy you may already own. Warn first, do not block first. A warning that names the data type teaches people something. A silent block teaches them to use their phone.
Ask your team what they use, in a way that is safe to answer. We have seen more useful disclosure come out of one meeting where nobody got in trouble than out of any technical scan.
Write the questionnaire answer accurately, even when it is uncomfortable. Naming three tools and the controls around them reads better to a client’s security team than None, and it survives their verification.
Now the part where we sell something. Scoping the monitoring so the console shows a handful of meaningful events a week instead of noise, tuning the paste rules against your actual document types, and having someone read the output every morning, that is work and it does not do itself. That is what Secure Data Guard is. The device half matters too, and it is the one people forget: a laptop that goes home with a resigning employee carries a browser profile still signed into whatever accounts they created on it. Our device lifecycle and Device-as-a-Service work exists so that machine comes back, gets wiped to a certificate, and stops being your problem. Buying 50 or more laptops this year, ask us about DaaS before you raise the purchase order, because retirement is easier to fix at the buying stage than at the leaving stage.
Where we lose, and I would rather write it here than argue it in a proposal. If you are under 25 people, hold no customer data beyond invoice addresses, and answer to no client security team, you do not need us for this. Do the extension inventory yourself, pick one paid AI account for the company, and spend the money on something that grows revenue. Forcing data protection tooling onto a business that does not need it makes us the vendor we claim to dislike.
What to take away
- The leak is a paste, not an upload. File-based monitoring is watching the wrong action. Clipboard text into a browser form leaves no attachment, no filename and no event.
- Blocking one site relocates the traffic. One block plus no visibility produced fourteen destinations and a firm that sincerely believed the answer was none.
- The identity is the risk, not the tool. A company account gives you logs, retention and a contract. A personal login gives you nothing and outlives the employee.
- Your client will find this before the regulator does. A vendor security questionnaire is the enforcement mechanism that actually bites in 2026.
- You are probably licensed for the fix already. Browser paste inspection sits inside suites Indian firms are paying for and have never configured.
Four terms in this piece, in plain English
- Shadow AI
- Staff using AI tools the company has not approved and usually does not know about, most often on their own free accounts. It is the AI version of the unapproved app, with the difference that the company data goes into the tool rather than just sitting near it.
- DLP, or data loss prevention
- Software on company laptops that watches where information goes. Email attachments, USB copies, cloud uploads, and on the newer setups, text pasted into a browser. It records the who, the what and the when, and it can block the movements you tell it to block.
- DPDP Act
- India’s Digital Personal Data Protection Act, 2023, the country’s main privacy law. It makes the company holding personal data responsible for protecting it, gives individuals rights over their own data, and sets penalties running to ₹250 crore for a failure of reasonable security safeguards.
- DaaS, or Device-as-a-Service
- Renting laptops instead of buying them. One monthly cost per device covers supply, support, replacement and a certified wipe when the machine comes back. It matters here because a returned laptop carries browser profiles still logged into accounts nobody inventoried.
Questions the operations head asked me that fortnight
Is using ChatGPT at work actually illegal in India?
No. India has no statute banning workplace use of generative AI, and the DPDP Act does not name AI tools at all. What it does is make you responsible for personal data you hold, including when an employee discloses it to a third party through a browser. So the legality question is the wrong one. The question that decides your exposure is whether you can show, if asked by a client or the Data Protection Board, what personal data left your building, to which service, on whose account, and under what retention terms. A company on a sanctioned tool with logs can answer that. A company relying on a firewall rule from last year cannot.
We blocked the AI sites on our firewall. Are we covered?
Almost certainly not, and the block may be making things harder to see. A domain block stops direct visits from your office network to the addresses on the list, which is a small share of how people reach these tools. It does nothing about the sites that were not on the list, browser extensions that rewrite text inside any page, personal phones on mobile data, or a laptop used at home. In the Hyderabad exercise the blocked site was not even the busiest destination. Worse, the block creates a settled belief in the management team that the matter is handled, which is why nobody looked again for fourteen months. Pair any block with visibility, or skip the block and keep the honesty.
Can we monitor what staff paste into AI tools without their consent?
On company-owned devices, for employment purposes, yes, within limits. Section 7(i) of the DPDP Act 2023 treats employment as a legitimate use, so consent is not the basis you rely on, and its examples cover safeguarding the employer from loss or liability, prevention of corporate espionage, and confidentiality of trade secrets and intellectual property. Section 7 is a closed list, so the further the monitoring drifts from those purposes, the weaker the position gets. Scope it to company devices and to the data categories that would cause real loss, write that scope down, and tell people about it at induction in plain words. Telling them also changes behaviour more than the alerts do.
What is the cheapest first step for a 100-person company?
Two things, in this order. Run a browser extension inventory across your fleet, because it takes an afternoon and it is where the ugliest finding usually is. Then pick one AI tool, buy the business tier of it, and put every employee on a company login. Those two together remove most of the reason for shadow use and give you a log you can show a client. After that, check whether your existing Microsoft or endpoint licence already includes browser paste inspection, because a large number of Indian mid-market firms are paying for it and have never turned it on.
Still deciding? These are the pages we send next.
- How employees steal company data: 5 methods most Indian companies cannot detect
- New employee data protection: what the first 30 days must cover
- The USB data theft risk your endpoint policy is probably ignoring
- DPDP compliance checklist: 10 things to fix before May 2027
- DLP for small business in India: the starter guide for companies under 200 people
- Secure Data Guard, our data protection practice
The Hyderabad firm resubmitted question 34 in June. Three tools named, company accounts, retention settings attached, paste policy in warn mode. The client’s security lead wrote back with one line, and the operations head forwarded it to me without comment. You are the first vendor this year who did not write none. Theek hai. That reply is worth more than the fourteen months of quiet.
A two-week read-only visibility exercise on a sample of your laptops. No blocking, nothing switched off. You get the list of AI destinations, the browser extensions nobody inventoried, and a plain answer you can put on a client questionnaire.
Section numbering, legitimate uses, the cross-border position and commencement dates follow the Digital Personal Data Protection Act, 2023 and the Digital Personal Data Protection Rules, 2025 as they stood at the time of writing. All of that can change, and whether Section 7(i) covers a specific monitoring programme depends on its scope. Confirm before relying on this. Research figures are attributed to their publishers. The Hyderabad engagement is one anonymised client matter with identifying details changed. A practitioner note on operating practice, not legal advice.






