In April 2023, Samsung’s semiconductor division discovered that its own engineers had leaked confidential company data to ChatGPT three separate times in less than 20 days. One engineer pasted proprietary source code into the chatbot to ask it to find and fix a bug. A second uploaded confidential chip test sequences closely guarded because optimizing them is worth real money, and asked ChatGPT to improve them. A third converted a recording of an internal meeting into text and fed it to the tool to help draft a presentation. None of these employees intended to cause a breach. They were trying to work faster.
Samsung’s response was immediate: a company-wide ban on generative AI tools, a 1024-byte cap on prompts while it built an internal alternative, and a warning to staff that anything typed into ChatGPT is transmitted to and stored on servers the company doesn’t control, with no way to get it back. Samsung wasn’t alone. Months earlier, Amazon’s legal team had already warned employees after spotting ChatGPT outputs that echoed internal Amazon material closely enough to raise concern, and told staff not to paste confidential information into the tool under any circumstances.
These two cases are now the most-cited examples of a problem that has only grown since: shadow AI risk, the exposure created by unsanctioned AI tools in the workplace, AI applications that IT and security teams never approved, never reviewed, and often don’t know exist.
Shadow AI Risk Is Bigger Than a Policy Gap
“Shadow IT” used to mean an employee signing up for an unapproved file-sharing tool. Shadow AI is a different order of problem, because the tools in question are specifically designed to ingest whatever you give them, and many general-purpose AI products retain and can train on user input unless an organization has negotiated enterprise terms that say otherwise.
Recent research from data-loss prevention firm Cyberhaven puts real numbers on how common this has become: 39.7% of AI interactions monitored across its customer base involve sensitive data, and the average employee submits sensitive information to an AI tool roughly once every three days. Just as significant is how employees access these tools. Cyberhaven found a majority of interactions with Claude (58.2%) and Perplexity (60.9%) happen through personal, non-corporate accounts, with ChatGPT (32.3%) and Gemini (24.9%) not far behind. A personal account bypasses every control a security team relies on: single sign-on, centralized logging, data retention policies, and any contractual guarantee about how submitted data is used. From the organization’s point of view, that data simply leaves the building.
What’s going in is not limited to source code. Cyberhaven’s data shows sales and go-to-market material (nearly 30% of what sales teams share), R&D and product content, and regulated data such as health records all showing up in prompts. Crucially, the researchers frame this as “rational employee behavior” rather than negligence: people are using the fastest tool available to do their jobs. Governance simply hasn’t kept pace with adoption.
Why This Is More Than a Confidentiality Problem
Unsanctioned AI tools in the workplace create a specific kind of exposure that traditional data-loss controls weren’t built to catch. A file uploaded to an unauthorized cloud drive is still, at least, a file, sitting somewhere identifiable, that a security team can eventually find and revoke access to. Text or audio pasted into a public AI chatbot is different: it can be retained, logged, used to train future model behavior, or exposed if the vendor itself suffers a breach, and an organization frequently has no visibility into which of those things happened.
That distinction matters even more for a specific category of material: recordings of meetings, voicemails, interview clips, and internal video, the exact kind of content employees increasingly feed into AI tools for transcription, summarization, or translation. This is a genuinely emerging concern, and worth stating plainly: the raw audio and video circulating through unsanctioned AI tools is the same category of material that voice-cloning and deepfake-generation tools are built to consume. A leaked recording of an executive’s voice, once outside the organization’s control, doesn’t need to be “hacked” again to become dangerous. It only needs to end up somewhere an attacker can reach it. Security teams evaluating shadow AI risk should treat any recording, transcript, or video pushed through an unapproved tool as a potential input to the same fraud pipeline covered in deepfake CEO fraud and AI voice cloning cases, not a separate problem.
How to Reduce AI Data Exposure
A ban rarely works on its own. Samsung’s own experience shows a blanket “no generative AI” policy tends to push usage further underground rather than eliminate it. A workable approach combines sanctioned alternatives with clear rules and real oversight.
Publish an AI acceptable use policy, and make it specific. Vague guidance (“use AI responsibly”) doesn’t give employees a usable line to work from. Name what categories of data can never be pasted into a public AI tool, source code, customer records, financial data, health information, unreleased product material, and say so explicitly.
Provision an approved, enterprise-grade alternative. Employees adopt shadow AI tools because sanctioned options are slower, more limited, or don’t exist. An enterprise AI subscription with contractual data protections and training opt-out gives people a legitimate tool that meets the same need.
Deploy DLP and CASB tooling that covers AI destinations. Traditional data-loss prevention rules were written for email and file transfer, not for a browser tab pasting text into a chatbot. Modern DLP and cloud access security broker (CASB) tools can specifically flag or block sensitive data moving into AI tool domains, sanctioned or not.
Require enterprise accounts, not personal logins, for any approved AI tool. This alone closes the visibility gap Cyberhaven’s research identifies: enterprise accounts bring SSO, logging, and retention controls that personal accounts never will.
Extend AI governance to vendor and vendor risk review. Any AI tool an employee or a business unit wants to adopt, sanctioned or not, should go through the same third-party risk review as any other software touching company data: what happens to submitted data, is it used for training, where is it stored, can it be deleted.
Train employees on the specific risk, not just the rule. The Samsung and Amazon cases are useful precisely because they show well-intentioned employees causing real exposure. Awareness training that walks through what actually happened, and why it mattered, tends to land better than a policy memo alone.
Where AI Governance Law Currently Stands
Shadow AI risk isn’t just an internal control problem; it increasingly intersects with active or emerging legal obligations, though the regulatory picture differs sharply by region.
European Union. The EU AI Act is the most comprehensive framework in force, using a tiered risk model: unacceptable-risk practices are banned outright, high-risk systems (including those used in employment, essential services, and law enforcement) carry data governance, documentation, and human-oversight obligations, and general-purpose AI model providers must document training data and meet transparency requirements. Obligations are phasing in on a staggered timeline through 2027, with prohibited-practice bans and GPAI rules already in effect.
United Kingdom. The UK has not passed a standalone AI law, instead applying existing UK GDPR obligations to AI use through Information Commissioner’s Office (ICO) guidance. That guidance requires organizations to run Data Protection Impact Assessments for AI systems touching personal data, ensure transparency and statistical accuracy, and address algorithmic fairness, all under the UK’s existing “pro-innovation,” principles-based approach to AI regulation. Employees pasting personal data into a public AI tool can itself trigger a UK GDPR obligation.
United States. There is still no comprehensive federal AI law, so obligations come from a fragmented and fast-growing set of state statutes. The NIST AI Risk Management Framework has become the de facto national standard for demonstrating due diligence, and some states now reward it directly: Texas’s TRAIGA law, effective January 2026, grants an enforcement safe harbor to organizations that can show substantial alignment with the NIST AI RMF. California’s SB 53 and AB 2013 (also effective January 2026) add publication, incident-reporting, and training-data transparency duties for AI developers, and Colorado has replaced its original AI Act with a narrower automated-decision law taking effect in 2027.
Canada. The federal Artificial Intelligence and Data Act (AIDA), part of Bill C-27, died when Parliament was prorogued in January 2025 and will not return in its original form. In its place, AI use in Canada is currently governed by existing law: PIPEDA at the federal level (which limits repurposing personal data collected for one purpose, such as training a model, without fresh consent), Quebec’s Law 25, which requires notice and a right to human review for automated decisions, and sector-specific rules such as OSFI Guideline E-23 for financial institutions.
The throughline across all four jurisdictions is the same: regulators increasingly treat what data goes into an AI system, and what happens to it afterward, as a governance question with legal weight, not just a technical or IT policy matter.
Building Governance That Keeps Up
Shadow AI risk exists because individual tools move faster than organizational policy. Closing that gap takes more than a memo: it takes a documented AI acceptable use policy, a real gap analysis against frameworks like the NIST AI RMF, and a governance structure that can be shown to regulators, auditors, and insurers alike.
This is exactly the work CyberClan’s Governance, Risk, and Compliance team does every day. Built on established frameworks including NIST CSF, ISO standards, and the Secure Controls Framework, CyberClan’s GRC service creates policies tailored to how your organization actually operates, not generic templates, backed by gap analyses of what’s currently in place and audit-ready documentation your team can stand behind.
Not sure where shadow AI fits into your current risk posture?
Talk to CyberClan about a governance, risk, and compliance engagement that brings AI use inside your organization’s existing controls, before an unsanctioned tool becomes your next data exposure.
Sources: Tom’s Hardware on the Samsung ChatGPT leak incidents, Forbes on Samsung’s generative AI ban, OECD.AI incident record on Amazon’s ChatGPT warning, Cyberhaven 2026 AI data exposure research, EU AI Act high-level summary, UK ICO guidance on AI and data protection, US state AI law compliance guide, Canadian AI regulation status post-AIDA.


