the library
white paper
THE LIBRARY · WHITE PAPER · A THIRD ARK RECORD

The Shadow AI Problem

sealed 2 September 2026 AD·
the paper
11 min read

Why your people are leaking your firm to ChatGPT, and what to do about it. A Third ARK white paper, for the CISO, the General Counsel, and the CTO.

This paper makes claims. Every one of them is followed to its source at the foot of the document. That is not a courtesy. It is the entire thesis in miniature: a claim without its source is a rumour, and an age that has forgotten the difference is the age this paper is about. Read the claims. Then follow them down.

1. The pattern

It does not begin with a hacker. It begins with a copy and a paste.

An analyst is three hours into a due-diligence memo. The deadline is tonight. She has a folder of confidential financials and a model that, for ten dollars a month, will summarise them in nine seconds. She pastes. The summary is good. She pastes again tomorrow. By the end of the quarter she has moved more of the firm's confidential material through a third party's servers than any single breach in the company's history, and not one alert has fired, because no rule was broken that any system was watching for.

This is Shadow AI: the use of public, consumer-grade artificial intelligence on company work, outside the knowledge or control of the people accountable for the company's data. It is not malice. It is diligence, misdirected. The most conscientious employees are the heaviest users, because the tool genuinely makes their work better, and they have no reason to believe the paste leaves a mark.

It does leave a mark. It simply leaves it somewhere you cannot see.

The scale is no longer speculative. When researchers first measured it in 2023, roughly one in ten employees had pasted company data into ChatGPT, and around eleven per cent of what they pasted was confidential: source code, client records, internal strategy [1]. That was the floor. By 2026, the Verizon Data Breach Investigations Report found that forty-five per cent of employees were regular AI users on corporate devices, up from fifteen per cent the year before [2]. A January 2026 survey found that nearly half of employees admitted to uploading sensitive company or customer information into AI chats [3]. The curve is not slowing. It is the steepest adoption curve in the history of enterprise software, and it is happening below the line of sight of every governance function you have.

You already know this is happening inside your firm. You cannot prove where, you cannot prove how much, and you cannot prove it has stopped. That is the pattern.

2. The quantified risk

The risk is not one risk. It is four, and they compound.

Privacy and personal data. The moment an employee pastes a document containing a client's name, a patient's record, or an employee's grievance into a public model, the firm has, on a plain reading, transferred personal data to a third-party processor it has not assessed, under terms it has not negotiated, in a jurisdiction it may not know. Under the UK GDPR and its European parent, that is a processing event requiring a lawful basis, a controller-processor agreement, and a record. None of those exist for a paste. The exposure is not theoretical: the regime carries penalties to the higher of twenty million euro or four per cent of global annual turnover, and supervisory authorities have shown no reluctance to pursue novel processing.

Intellectual property. This is the Samsung lesson, and it is worth stating plainly because it is the most concrete case on the record. In April 2023, within a span of twenty days, Samsung engineers pasted proprietary source code and internal meeting notes into ChatGPT on three separate occasions. The company's response was not a memo. It was a ban on generative AI across the workforce [4]. The damage in such a case is rarely the single document. It is that confidential material, once submitted, cannot be recalled, cannot be proven destroyed, and may persist in systems the firm has no standing to inspect. For a business whose value is its know-how, a research boutique, a law firm, a med-tech, this is the asset walking out of the door in pieces, each piece too small to trigger an alarm.

Regulatory exposure, and the new layer above it. The EU AI Act entered into force in August 2024. Prohibited practices became enforceable in February 2025; transparency obligations for general-purpose AI followed in August 2025; and the broad obligations for high-risk systems land on 2 August 2026, carrying penalties to the higher of thirty-five million euro or seven per cent of global turnover [5]. A firm that cannot say which AI systems its people use, on what data, to what end, is a firm that cannot evidence compliance with a regime that is now months from full force. In regulated sectors the floor is higher still: pharmaceutical and medical-device work answers to the FDA and the MHRA, where an undocumented data flow is not a fine, it is a finding.

Training-data appropriation. The quieter risk, and the one that does not appear on a balance sheet until it is too late. Material submitted to a public model may, depending on the tier and the terms, become part of the substrate from which the next public model is built. Your proprietary method, your hard-won dataset, your distinctive way of framing a problem, the very things that make your firm worth more than its competitors, can be absorbed, generalised, and handed back to the whole market as a free feature. You will have paid to train your replacement.

Four risks. Each compounds the others, and all share one root: there is no record. You cannot govern what you cannot see, and the public tools are built, deliberately, so that you cannot see.

3. Why the obvious fixes fail

Three responses are reached for first. Each fails, and it is worth being precise about why, because the failure points directly at the answer.

The ban fails because the tool works. Samsung banned generative AI, and so have many others. A ban does not remove the incentive that created the behaviour; it only removes the behaviour from view. The analyst still has her deadline and the tool still saves her three hours. She will use it on her phone, on her home machine, through a personal account, and now the firm has converted a visible problem into an invisible one. A prohibition you cannot enforce is not a control. It is a liability with a paper trail that says you knew.

The enterprise seat fails because it solves the wrong half. Buying a managed, enterprise tier of a public model is a real improvement on the consumer free tier: the contract is better, the data-handling commitments are stronger, the promise not to train on your inputs is written down. But it does not change the architecture. The data still leaves your perimeter. The provenance still vanishes at the moment of submission. You are still trusting a third party's word, audited by a third party's process, that your most sensitive material is being handled as promised, and when General Counsel asks you to prove, from your own records, exactly what was sent and what came back, you cannot. You have rented a better promise. You have not built a record.

Conventional data-loss prevention fails because it was built for a different shape of problem. SaaS DLP watches for known patterns leaving known channels: a credit-card number in an email, a file moving to an unsanctioned cloud drive. Shadow AI does not move a file. It moves an idea, paraphrased, in a chat box, over an encrypted connection your own staff initiated for a legitimate reason. The leak does not match a signature because the leak is the work itself. DLP can tell you a document was attached. It cannot tell you that an analyst described its contents, in her own words, to a model that remembered.

Each fix fails at the same seam: it tries to police the data on its way out, rather than removing the reason it had to leave at all. The answer is not a better wall around the same broken arrangement. It is to change the arrangement, so the data never has to leave to be useful.

4. The architectural answer

The principle is simple to state and hard to fake: if the intelligence comes to the data, the data never has to go to the intelligence. Four properties make that real. Together they are what we call the Trust Wall, not a wall around the cloud, but the removal of the cloud from the path.

Local-first inference. The model runs on hardware the firm controls, on the Operator's own machine, or a node inside the firm's own walls. The confidential document is reasoned over where it already sits. It is not uploaded, because there is nowhere to upload it to. The paste that begins the Shadow AI pattern has no destination outside the perimeter, because the intelligence is inside it. This is not a private cloud, which is someone else's computer with a better contract. It is your computer, doing the work, answering to you.

Bring your own keys. Where external models are used at all, and there are tasks for which they are the right instrument, they are reached through keys the firm holds, under terms the firm sets, with the boundary drawn by the firm and not by the vendor's default. The firm decides what may cross the line and what may never. The decision is configuration, not trust.

Cryptographic provenance. Every claim the system produces is deep-linked to the source it was drawn from. Not a footnote that gestures at a document, the document itself, held, sealed, and carried with the claim wherever it travels. This is the Atomic Passport, and it is the subject of the next section, because it is the property that turns a security posture into an evidentiary one.

Immutable audit. Every action leaves a record that cannot be quietly revised. When the regulator, the auditor, or General Counsel asks what was processed, by whom, on what data, and to what end, the answer is not a reconstruction from memory and goodwill. It is a log, sealed, that no one, including the firm itself, can rewrite after the fact. The record is collateral. Its value is precisely that it binds the firm as much as it protects it.

Notice what these four properties do together. They do not ask the employee to be more careful. They remove the unsafe path entirely and replace it with one that is faster, because the model is local and the source is already to hand. Security that depends on discipline fails the day the deadline is tight. Security that is built into the shape of the tool holds on exactly that day.

5. The Atomic Passport

Most of this paper could have been written by any vendor with a local model and a logging feature. This section could not, and it is the part that matters most to the office of the General Counsel.

The deepest risk of the public AI is not that it leaks. It is that it speaks with total confidence and leaves no way to check it. It tells you what to think in a fluent, authoritative voice, and the chain back to whatever it was actually drawn from has been severed. For a firm whose product is judgement, a legal opinion, a research note, a regulatory submission, an answer you cannot trace is not an asset. It is a liability dressed as an asset. The frame, the words, and the confidence are supplied; the ground is not.

The Atomic Passport is the opposite arrangement. The unit of work is not a document but an atom, a single, sealed fragment of meaning that carries, with it and inseparable from it, the source it came from. When the system asserts a fact, the assertion is deep-linked to the primary material: the page, the transcript, the captured record. Follow it down and you arrive at bedrock, the actual source, archived, so that it remains verifiable even if the website changes, the link rots, or the internet itself moves on.

Two consequences follow, and both are commercial.

First, the answer is scrutable. You do not have to trust the machine, because you can check it, cheaply, every time. The burden of proof is inverted: instead of taking a confident assertion on faith, you take the chain to its source and read it yourself. For a regulated firm, this is the difference between a tool you must keep away from anything that matters and a tool you can put at the centre of the work.

Second, the record is staked. Because every atom is sealed and every action is logged, the firm's account of its own work is not a story it can revise when the story becomes inconvenient. The provenance binds the firm. That is not a weakness; it is the point. A record that can be edited proves nothing. A record that cannot be edited is evidence, to a regulator, to a court, to a client, and to the firm's own future self on a day it cannot fully remember what it decided.

Scrutable, and staked. That is the standard a public model cannot meet without ceasing to be a public model, and it is the standard a firm that lives by its judgement should refuse to work below.

6. Implementation: the ninety-day pilot

This is not a platform migration. It is a contained proof, run on the work that matters most, and designed to produce evidence rather than enthusiasm.

Days 1 to 15. Scope. Choose one team and one class of confidential work: the due-diligence desk, the litigation-support group, the regulatory-affairs function. Name the data that must never leave the perimeter. Define, with General Counsel, the specific evidentiary question the pilot must answer: can we prove, from our own records, what was processed and what was produced.

Days 16 to 45. Stand up. Deploy local inference on the team's own hardware or a node within the firm's walls. Configure the boundary: what stays local, what may reach an external model under the firm's own keys, what may never cross. Turn on the Passport and the audit log from the first hour, so the record begins with the pilot.

Days 46 to 75. Run live. The team does its real work through the system, on real confidential material, against real deadlines, the only honest test, because the unsafe path is only taken when the pressure is on. Measure two things: the work product, against the team's existing standard; and the record, against the evidentiary question from day one.

Days 76 to 90. Prove and decide. Produce, from the firm's own logs, a complete account of what was processed and what was generated across the pilot. Put it in front of the regulator's question, the auditor's question, and General Counsel's question. The pilot succeeds not when the team is happy but when the firm can answer those three questions from its own records, without recourse to any third party's goodwill. That answer is the deliverable.

A pilot framed this way cannot fail quietly. It either produces the evidentiary record or it does not, and either outcome is worth more than another quarter of not knowing.

7. The decision

Shadow AI is not a passing risk to be waited out. It is the new default behaviour of your most capable people, and it is invisible by design. You can ban it and lose sight of it. You can rent a better promise and still have no record. Or you can change the architecture, so that the intelligence comes to the data, the data never leaves, and every claim and every action is scrutable and staked.

The firms that will hold their advantage through this decade are not the ones that used the most AI. They are the ones that can still prove what their AI did. In an age of fluent, confident, unaccountable machines, the only currency that holds its value is bedrock truth, the claim you can follow down, and the record that cannot run from it.

That is the ground Third ARK is built on. It is scrutable, if you would care to look.


Sources and further reading

  1. Cyberhaven Labs, ChatGPT at Work: confidential data pasted into ChatGPT (2023); ~11% of pasted content confidential; ~2.3% of employees had entered confidential company data. cyberhaven.com
  2. Verizon, 2026 Data Breach Investigations Report, as reported by Push Security: 45% of employees regular AI users on corporate devices, up from 15%. pushsecurity.com
  3. Employee Shadow AI survey, January 2026: ~48% report uploading sensitive company or customer information into AI chats. clickondetroit.com
  4. Bloomberg / Forbes: Samsung bans generative AI after three sensitive-code leaks within 20 days, April and May 2023. bloomberg.com · forbes.com
  5. European Commission, AI Act: in force 1 August 2024; prohibited practices from February 2025; GPAI transparency from August 2025; high-risk (Annex III) obligations from 2 August 2026; penalties to the higher of €35M or 7% of global turnover. digital-strategy.ec.europa.eu

GDPR / UK GDPR penalties: to the higher of €20M or 4% of global annual turnover. Figures should be re-verified at point of publication (chairman review required, pre-counsel).

The cloth-bound edition.

Free to read here, sealed and numbered. A limited cloth-bound edition is being prepared. Request a copy and we will write to you when it is ready.

Request a copy → Limited edition · numbered · price on request
Sealed by Third ARK · sha-········ · Verify this piece →