

An executive whitepaper on how AI-Native DLP differs from legacy DLP, and what it means for enterprise security strategy.
Download WhitepaperThree separate incidents. Three Samsung engineers. Each one pasted proprietary semiconductor source code into ChatGPT in 2023. The first submitted buggy database code for debugging. The second uploaded equipment code for optimization. The third asked ChatGPT to generate meeting minutes. Samsung responded by restricting prompts to 1,024 bytes and initiating disciplinary review.
These weren't malicious insiders. They were solving problems faster using the most convenient tool available.
Developers use ChatGPT to debug logic errors, generate code snippets, and decode cryptic error messages. Federal contractors and defense organizations have moved to block ChatGPT access specifically because proprietary code and potentially classified information can leave through prompts. The problem isn't intent. It's habit.
A government contractor in Australia uploaded a spreadsheet containing personal information from 3,000 flood victims into ChatGPT while reviewing disaster recovery applications. Names, contact details, health information — all of it. The contractor was trying to move faster through a backlog.
Healthcare workers have followed the same pattern, entering patient information into public AI tools to summarize notes and assist with documentation. When protected health information reaches systems without Business Associate Agreements, the result is direct HIPAA exposure. The intent is efficiency. The consequence is a compliance incident.
OpenAI actively promotes uploading Excel files directly to ChatGPT for tasks like identifying variance drivers, checking anomalies, and summarizing trends. Finance teams have taken that capability and applied it broadly — analyzing spreadsheets, fixing formulas, generating variance commentary. Financial analysts build budget trackers, compare regional performance, and visualize revenue by product line.
Each upload carries something that shouldn't leave the organization. Deal structures. Client account details. Proprietary financial models. The upload itself is the exposure.
Cyberhaven data shows 32.3% of ChatGPT usage occurs through personal accounts. Employees use personal accounts for convenience, to work around restrictions, or simply out of habit. Personal ChatGPT accounts sit entirely outside enterprise controls. Conversation history is stored by default. Prompts are used for model training unless the user explicitly opts out.
No corporate policy reaches that account. No DLP rule applies to it. And the data flowing through it is the same data flowing through everything else.

Traditional DLP monitors predictable channels: email attachments, file transfers, network traffic, and sanctioned cloud applications. That model worked when data movement followed known paths through corporate gateways. The assumption was straightforward — data moves through known systems, and controls can sit at those known exits.
That assumption no longer holds.
Browser-based AI tools, AI features embedded inside Microsoft 365 and Google Workspace, and third-party LLM integrations represent data pathways that sit outside traditional endpoint or network DLP. When users work directly through web browsers, data flows through copy-paste actions, API calls, and third-party integrations. Many of these interactions don't involve file transfers at all. There is no attachment to scan, no USB to flag, no email to inspect.
Traditional DLP classifiers target structured data patterns: Social Security numbers, credit card numbers, known file formats. The content employees feed into AI tools is overwhelmingly unstructured: meeting notes, draft documents, code snippets, financial models.
A user who pastes a paragraph describing customer account details, contract terms, or strategic plans creates exposure without triggering a single pattern match. Regex-based detection captures explicit formats. It misses meaning entirely.
The result is a detection gap that scales with AI adoption. Seventy percent of data leaks now happen directly in the browser, and fifty-three percent involve copying data into chat applications or AI prompts. Pattern matching was not designed for this. It was not designed for a world where sensitive information travels as plain prose.
Network DLP inspects traffic at egress points. Encrypted HTTPS sessions obscure payload contents unless SSL inspection is deployed. Direct-to-cloud connections from remote employees bypass corporate gateways entirely.
When a developer copies proprietary code into a desktop AI coding assistant, no inspectable network event occurs until data has already moved. The interception point that matters — the moment of input — is invisible to network-based tools.
ChatGPT Enterprise provides encryption, SSO, and data retention controls. These protections apply to the platform itself. They don't inspect or classify prompt content before submission.
The account is secured. The prompt is not.
That distinction matters. Enterprise agreements address how data is stored and whether it is used for training. They do not address what employees type into the chat window before hitting send. The exposure happens at input. Controls that operate after submission are already too late.
The problem with most DLP solutions is where they sit. They watch the network, scan file transfers, and flag known patterns. None of that helps when an employee pastes proprietary code into a browser tab. Wald AI DLP is built differently. It operates at the point where data actually leaves the organization.
Wald AI DLP runs through a browser extension that intercepts prompts at the moment of submission. Copy-paste actions, typed text, and file uploads are all captured before data reaches ChatGPT, Claude, Gemini, or any other generative AI tool. The interception happens locally on the endpoint.
Network-based DLP and gateway solutions cannot see this activity. Wald can. Coverage extends to the ChatGPT desktop application as well, not just browser sessions.
Traditional pattern-matching DLP identifies data by structure: Social Security numbers, credit card formats, known file types. That approach fails the moment sensitive information leaves as free text.
Wald's on-device small language model evaluates prompt meaning and intent. It recognizes a paragraph describing customer account details. It catches contract terms paraphrased in prose. It identifies strategic planning notes that contain no flagged strings at all. The classification engine processes semantic meaning locally, without sending prompt content to external servers for analysis.
This distinction matters. Most of what employees paste into AI tools is unstructured. Regex catches formats. Wald catches meaning.
The policy engine returns an enforcement decision before transmission completes. When a user submits a prompt containing classified data, the on-device SLM evaluates the content against policy rules and issues a verdict in real time.
Three outcomes are possible:
No sensitive data reaches OpenAI's servers until after the classification verdict executes. That sequence is what makes the control real.
Silent blocking creates workarounds. Employees retype content, switch devices, or use personal accounts to avoid friction. Wald's Warn action takes a different approach.
When a prompt triggers a policy match, users receive a context-specific notification explaining what was detected and why. They can provide a business justification and override the warning when legitimate work requires sharing the flagged content. Override capability is configurable by organization. Not every flagged prompt requires escalation.
The result is an audit trail for compliance documentation without removing the user's ability to make a judgment call when the situation warrants it.
The enforcement layer only matters if classification is accurate. Here is what Wald's on-device SLM detects in practice.
Wald intercepts code snippets copied from IDEs, terminal windows, and Git repositories before they reach ChatGPT or GitHub Copilot. The on-device SLM identifies proprietary algorithms, API keys embedded in configuration files, and internal library references. Pattern-based DLP misses these entirely. They contain no flagged strings.
Healthcare workers enter patient symptoms, treatment notes, and diagnostic details into ChatGPT to draft clinical summaries. Support agents paste ticket transcripts containing account numbers and contact information. HR teams upload performance reviews and salary data.
Wald's semantic classifier detects protected health information and personally identifiable information regardless of format. A prose description with no Social Security number or credit card pattern still triggers classification if the meaning warrants it.
Finance teams upload budget models, variance reports, and client account summaries. Deal structures, margin analysis, and competitive pricing intelligence embedded in Excel formulas and pivot tables are flagged before file upload completes. The risk is not always in what employees type. Sometimes it is in what they attach.
The browser extension monitors all generative AI endpoints, not just ChatGPT Enterprise. Activity flowing to personal ChatGPT accounts, Claude, Gemini, and unsanctioned tools receives the same classification and policy enforcement. Employees switching tools do not switch off controls.
Coverage extends to the ChatGPT desktop application on Windows and macOS. Prompts submitted outside browser tabs are captured. The endpoint is the control point, not the network.
Traditional DLP rollouts stall on hardware installation, database configuration, and policy-writing cycles that stretch across months. Wald takes a different approach. The browser extension distributes through standard MDM in minutes. The on-device SLM downloads automatically during initial setup. Pre-configured policies for banking, healthcare, and legal sectors are ready to enforce from day one.
There is no six-month policy project. No infrastructure dependency. Governance is operational before the end of the first day.
Most mid-market organizations don't staff a dedicated DLP analyst. Wald accounts for that. Semantic classification reduces the false positives that consume analyst time in pattern-matching systems. Policy management runs through a single console. Flagged incidents surface with context, not raw log files that require manual correlation.
One security administrator can operate the system. No specialized DLP expertise required.
Pre-built policy templates map directly to HIPAA protected health information elements, GDPR personal data categories, and PCI-DSS cardholder data requirements. Every flagged prompt generates a log entry: what data was detected, which employee submitted it, what action was taken, and when.
That audit trail addresses what compliance examiners ask for. No custom report development needed.
Coverage is not limited to ChatGPT Enterprise. The browser extension monitors all generative AI endpoints accessed through Chrome, Edge, and Firefox, including ChatGPT personal accounts, Claude, Gemini, Microsoft Copilot, and Perplexity. The same classification engine and policy rules apply regardless of which tool an employee opens.
Employees don't need to think about which AI tool is approved. The controls follow the data.
ChatGPT usage is already happening across your organization. Most of it is invisible to your current security stack.
Traditional DLP was not built for this. Pattern matching catches formats, not intent. Network inspection misses what happens at the endpoint. And ChatGPT Enterprise secures the platform, not what employees type into it.
The exposure is not theoretical. It is happening in engineering, support, finance, and HR, across both enterprise accounts and personal ones that your controls cannot touch.
For security leaders in regulated industries, this comes down to one question: are controls enforced at the point where data actually moves, or only documented in policy?
Wald enforces at the endpoint. Before the prompt reaches OpenAI. Without requiring a dedicated DLP team or a months-long deployment.
The leaks will not announce themselves. They will surface during an audit, a breach notification, or a regulator inquiry.
The controls need to be in place before that happens.
ChatGPT Enterprise provides encryption, SSO, and data retention controls for the platform itself, but it doesn't inspect or classify the actual content of prompts before they're submitted. While it secures the account and ensures data isn't used for training, it doesn't prevent users from typing or pasting sensitive information into the chat interface.
Traditional DLP systems were designed to monitor predictable channels like email attachments, file transfers, and USB ports. They struggle with browser-based AI tools because data often moves through copy-paste actions and direct web interactions rather than file transfers. Additionally, pattern-matching approaches fail to detect unstructured content like meeting notes or code snippets that don't contain obvious identifiers like credit card numbers.
Organizations can provide employees with compliant AI tools through enterprise agreements that include data protection guarantees, such as ChatGPT Team accounts with Data Processing Agreements or Microsoft Copilot. Combining approved AI tools with clear usage policies, employee training, and endpoint monitoring creates a framework that enables productivity while maintaining audit trails and compliance controls.
Common data leaks include source code and proprietary algorithms from development environments, customer personally identifiable information (PII) and protected health information (PHI) from support tickets and healthcare records, financial models and deal structures from spreadsheets, and strategic planning documents. These leaks often occur through copy-paste actions or file uploads that bypass traditional security controls.
Blocking public AI tools addresses part of the problem but doesn't eliminate it entirely. Employees may use personal devices, find workarounds, or simply retype information they can view on screen. A more effective approach combines technical controls with approved alternatives, clear policies, user education, and monitoring solutions that can detect and prevent sensitive data from being submitted regardless of the access method.
Purview provides DLP across Microsoft services, but organizations using ChatGPT and other third-party AI tools often need additional controls. Wald inspects prompts before they’re sent to AI models, extending protection across ChatGPT, Claude, Gemini, and more.
Wald inspects both typed and pasted prompts. Any content sent to ChatGPT is analyzed for sensitive data before it reaches the AI model.
Traditional DLP protects files, emails, and uploads. AI DLP protects conversations by inspecting prompts for sensitive information before they’re sent to AI assistants.
Wald supports browser-based ChatGPT. For desktop app coverage, check with the Wald team, as support depends on the deployment architecture and platform.
Yes. Wald explains why the prompt was blocked and, if your policy allows, lets users edit the prompt, provide a justification, or request an override.