DLP for the GenAI Era: Stopping Sensitive Data from Leaking into AI Tools

In 2023, engineers at Samsung pasted confidential source code and internal meeting notes into ChatGPT to speed up their debugging work. Within weeks, the company had banned generative AI tools company-wide. It was one of the first widely reported cases of its kind, but going by recent numbers, it was nowhere close to the last. 

Sensitive data flowing into AI tools has quietly turned into one of the more pressing headaches for security teams, and most of it isn’t caused by hackers at all. It’s caused by regular employees just trying to get through their to-do list a little faster.

This growing gap between how people actually work and how companies try to protect their data is exactly why DLP for the GenAI era has become such an urgent conversation.

The Quiet Risk Sitting Inside Every AI Prompt

The scale of this is larger than most executives assume. According to Cyberhaven’s analysis of enterprise usage, roughly 8.6 percent of employees have pasted company data into ChatGPT, and about 4.7 percent of that involved confidential material. A separate 2025 report from LayerX Security paints an even sharper picture: 45 percent of enterprise employees now use generative AI tools at work, and 77 percent of those users admitted to copying and pasting data straight into their prompts. Roughly a fifth of those pastes contained personal or payment information.

What makes this harder to manage is where most of that activity happens. LayerX found that 82 percent of these pastes came from personal, unmanaged AI accounts rather than sanctioned enterprise tools. 

None of this points to malicious intent. It points to convenience quietly winning out over caution.

Why Traditional DLP Wasn’t Built for This

Traditional data loss prevention tools grew up in a different environment altogether, one built around email attachments, USB drives, and file transfers to recognizable third parties. Those systems are reasonably good at flagging a spreadsheet leaving through an outbound email. They were never designed to interpret a free-flowing sentence typed into a chat window.

That’s the real distinction behind Traditional DLP vs AI DLP

Older systems lean on fixed patterns and known file formats, while GenAI risk shows up as unstructured, conversational text that looks different every single time someone types it. 

A rule built to catch “account number” inside a document may not catch that same number typed casually into a chatbot as part of an unrelated question. Gartner has flagged this shift too, predicting that more than 40 percent of AI-related data breaches will stem from improper cross-border use of generative AI by 2027, largely because oversight hasn’t caught up with adoption. 

What DLP for the GenAI Era Actually Looks Like

Modern DLP built with generative AI in mind needs to behave differently on a few fronts.

It has to read context, not just match keywords. 

Rather than scanning only for fixed patterns, it should recognize when sensitive information, in whatever shape it takes, is about to head toward an external AI service.

It needs real visibility into AI usage. 

Plenty of organizations still don’t have a clear picture of which AI tools their teams are relying on. This is where shadow AI comes in, essentially shadow IT’s newer cousin, and Gartner expects it to become a serious problem, projecting that around 40 percent of enterprises will face a security or compliance incident tied to shadow AI by 2030.

It should act in real time. 

Catching sensitive data before it leaves the browser beats discovering the leak months later during a compliance review.

It must still let people work. 

Nobody benefits from restrictions so heavy that employees quietly route around them. Effective DLP protects data while still letting the AI tools do their job.

Common Ways Sensitive Data Slips Out

A handful of patterns show up again and again across industries:

  • Employees pasting internal reports into AI tools to reformat or shorten them
  • Developers sharing proprietary code to debug faster, echoing the Samsung episode
  • Law firms and healthcare staff feeding client or patient details into public chatbots, occasionally triggering GDPR or HIPAA reviews
  • Support teams entering customer records into AI assistants for quicker responses
  • Marketing and HR teams uploading internal decks or resumes for summarizing

A related 2025 study from Harmonic Security found sensitive information tucked into more than 4 percent of prompts and roughly 20 percent of file uploads sent to AI tools, numbers that line up closely with what other researchers are seeing.

The Numbers Keep Climbing, Not Levelling Off

What’s notable is that this isn’t slowing down as awareness grows, it’s expanding. Cyberhaven’s own tracking found that the share of sensitive corporate data entering AI tools jumped from 10.7 percent to 27.4 percent within a single year, while total data volume flowing into these tools rose by close to 485 percent year over year. Separately, a Gartner survey found that 69 percent of cybersecurity leaders either suspect or have confirmed that employees are using AI tools without approval.

Put plainly, the adoption curve for AI at work has moved far faster than the guardrails meant to sit around it.

Building a Practical Defense

None of this calls for banning AI outright, that rarely works and usually just pushes usage further underground. A layered response tends to hold up better, built around a few core pieces:

  • Visibility first. Businesses can’t secure activity they can’t see, so mapping out which AI tools are actually in use across teams is the natural starting point.
  • Plain-language policies. Guidelines work far better when they’re written the way people actually talk, rather than buried inside a compliance PDF nobody opens.
  • Technical safeguards. Browser-level monitoring, prompt inspection, and automatic redaction of sensitive fields can stop a leak before it ever leaves the network.
  • Ongoing training. A short, clear explanation of what counts as sensitive information often prevents more incidents than any single piece of software.

Put together, these layers don’t just plug gaps individually, they reinforce one another, which is usually what makes the difference between a policy that looks good on paper and one that actually holds up in practice.

Where This Leaves Businesses

As generative AI settles into daily workflows rather than staying a novelty, the question companies are asking has shifted from “should employees use AI” to “how do we make sure they’re using it safely.” That shift touches data governance, staff training, and the systems quietly running in the background of everyday operations.

For organizations working through this, need high-level AI Workforce Security since the majority of leaks trace back to ordinary, well-intentioned use rather than anything unusual or deliberate.

Generative AI isn’t retreating, and neither is the exposure it brings if left unchecked. The organizations that stay ahead won’t be the ones locking AI tools out entirely, but the ones building sensible boundaries around how their people use them day to day. DLP for the GenAI era isn’t a single fix applied once; it’s a habit that has to grow alongside the technology itself.