An ivory folded-paper enclosure with one open gap, where a terracotta paper strip slips outside.

explainer · industry claim

Will AI labs steal your client data and train on it?

The short answer is no, and the fear points away from where a firm's data actually leaves. The economics, the technology, the provider terms, and what careful adoption looks like instead.

By Jeff Victorino · 22 August 2026

It is 9:40 on a Tuesday night. A junior bookkeeper is behind on a payroll reconciliation, so she pastes a client’s staff schedule, names and tax file numbers included, into a free chatbot on her phone. No approval was sought because no approval process exists. No log was kept because there is nothing to log it in. The partner believes the firm’s AI policy, which is to say its refusal to have AI, is protecting client data. The leak happened anyway, unwitnessed and uninsurable. [Illustrative scenario.]

That scene is the real data risk inside most professional-services firms, and it runs in the opposite direction to the fear this essay is about. The fear says: if we use AI, the provider gets our information, learns it, and sells it. A firm that bans AI has not removed the risks around AI. It has removed the controls. The same exposure remains, with none of the oversight.

The belief deserves a fair hearing, because it bundles a duty with a theory. The duty is real: client information must be handled responsibly, it is written into professional codes, and for lawyers it approaches the absolute. This essay never argues with the duty. The theory is something else: that a large AI laboratory wants your files specifically, will drink them into a model, and will resell what it learns. That theory is wrong about the mechanism, and acting on it is expensive. Take the three parts in order.

A. Your data is not the prize

Start with scale. Frontier models train on tens of trillions of tokens of public and licensed text. A large practice’s entire written output across twenty years, every return, workpaper, letter and spreadsheet, might reach a billion words, which is a few hundred million tokens. Against a training corpus tens of thousands of times larger, one firm’s life work is a rounding error. The dilution argument proves absence of motive, not immunity from memorisation, and the next section deals with memorisation honestly. But as a statement of incentive: no laboratory that can scrape the public internet needs the contents of one practice in Glen Iris. [Illustrative arithmetic.]

The business model points the same way. Laboratories earn their revenue from businesses: API contracts, enterprise seats, platform agreements. Major providers’ business and API tiers publish terms excluding customer data from model training by default, and zero-retention configurations are available on some platforms. Those are vendor-published commitments, and one honest sentence belongs here: the author runs a consultancy that builds governed AI systems for firms, so weigh that as you read. But the incentive direction is checkable arithmetic, not salesmanship. A provider caught brokering customer data would not lose one account. It would lose the enterprise market that constitutes its revenue, and the litigation record shows laboratories get sued over public web data, the lowest-stakes material they touch.

What about takers who are not buyers? This is where the original version of this argument overstated, so state it plainly. Extortion does not need a buyer: the Medibank breach was leverage, not resale. Litigation discovery and regulatory compulsion can reach a vendor’s servers. Insiders exist, at vendors as anywhere. Aggregated across thousands of firms, verified private outcomes would be valuable, though there is no evidence anyone is assembling that pool. Every item on that list is real, and every item is the same risk class your firm already carries today through its cloud practice platform, its email provider and its document store. The choice was never between trusting a vendor and trusting nobody. It is between governed third parties with contractual terms and ungoverned ones without.

The deeper mistake is about what the files contain. Here the author can speak from the trade he knows. [Founder experience.] Two decades in software engineering carried the same belief: source code is the crown jewels, and anyone who sees it will steal it. AI made the belief obviously wrong. A capable model can rewrite most of a codebase with better architecture inside a week. When work can be regenerated on demand, nobody steals the artefact. The moat was never the artefact. It was judgement: which problem, for whom, and what done means. The same inversion holds in every profession this firm serves. Clients never paid for the template. They paid for judgement applied to their situation, the relationship, and someone accountable when it goes wrong. None of that fits in a stolen file.

B. What the technology actually does, provider by provider

The mechanical error is imagining a model as a filing cabinet. Training compresses statistical patterns across enormous volumes of text. Afterwards the model holds the shape of language and reasoning, not a retrievable index of any document that went in.

Can specific documents resurface? Security research says: rarely, and through two documented routes. Unusual or duplicated text can memorise into weights out of proportion, and extraction can be attempted either through unusual access to model internals or through sustained adversarial querying through the ordinary interface. Both routes are studied, real, and nothing like “the vendor will sell your information”. They are closer to lock-picking knowledge: specialised, dangerous only with sustained access to your particular door, and defended by exactly the controls your profession already understands.

The deployment choice matters more than the provider choice. Retrieval-based systems, the standard pattern for firm tools, read your documents inside your tenancy at the moment of a query: the model learns nothing from the exchange, while logs, search indexes and backups stay in your tenancy under your retention settings. Fine-tuning is the different thing: it embeds your documents into the model’s weights, memorises far more readily, and belongs under explicit review if it is used at all. Governed adoption is retrieval-first for this reason.

And the providers genuinely differ, so here is the published picture as of August 2026. [Vendor-published policies; verify current terms before signing anything.]

Tier Trains on your data by default Retention
ChatGPT free and personal plans Yes, opt-out available in settings Account history
OpenAI business, enterprise and API No About 30 days of abuse-monitoring logs; zero-retention on approved arrangements
Claude consumer plans Choice-based; sharing extends retention from 30 days to 5 years 30 days unless opted in
Anthropic commercial and API No, contractually excluded About 30 days; zero-retention available
Gemini personal accounts Yes while activity setting is on Off means not trained
Google Workspace business tiers No Admin-set retention
Self-hosted open weights No: the publisher never sees your prompts Yours, locally

Read the table and the pattern is hard to miss. The training exposure lives almost entirely in free consumer accounts, which is where staff already are when a firm refuses to give them anything better. Every commercial tier excludes training. Self-hosted open-weight models are the strongest containment of all: the weights run on your hardware, the publisher receives nothing, and their training corpus was frozen at release, so your files can never join it. The honest caveats: consumer defaults differ by provider and change over time, terms get updated, and any cloud vendor can suffer a breach or face legal process. All of that is vendor due diligence, a recurring discipline with a checklist. None of it is answered by abstention.

C. What waiting costs

The adoption gap is measured. CPA Australia’s Business Technology Report 2025 found 89 percent of accounting and finance teams using AI while only 16 percent report it widely integrated across operations (CPA Australia, 2025). Almost everybody has tasted the tool. Almost nobody has rebuilt work around it. That gap is capacity sitting unused while fees stay flat and seniors drown in review.

The do-nothing firm leaks more than the governed one. Staff already use AI quietly on personal accounts, without approval, training or logging, and reported surveys of the profession name shadow usage as a top compliance risk (Wolters Kluwer Future Ready Accountant 2025, as reported by Accountancy Age). Refusal does not stop client details leaving through a junior’s phone. It guarantees it happens unwitnessed. Meanwhile the capability is being built informally anyway: 66 percent of chartered accountants report never receiving any workplace AI training (CA ANZ and Ipsos, 2025), so the firm gets AI adoption without AI governance, which is the worst available combination.

The calendar makes it worse. The Tax Practitioners Board’s draft guidance on AI, issued in March 2026 with consultation still open, makes explicit that registered agents remain responsible for verifying AI-assisted output and need client permission before client information enters AI tools. Privacy Act amendments commencing 10 December 2026 add automated-decision transparency for APP entities above the $3 million turnover threshold, subject to carve-outs. From late August that is roughly fifteen weeks. Governance work is happening in your firm either way; the only question is whether it ends in a policy nobody follows or in workflows that are compliant and faster.

Price it, roughly, for a five-partner practice. If governed tooling returns even six review hours per senior per week, that is a full day of fee capacity per partner, call it thirty senior-days a month across the firm. [Illustrative estimate, not a measured result.] Whether that becomes turnaround cut, margin banked, or both is the partner’s choice. The rival firm across town is making the same calculation with a head start.

Set the columns side by side. Governed adoption: bounded downside, insurable, reversible, measured before it scales. Waiting: the same exposure with none of the control, an unpriced capacity drain, and a compliance obligation that arrives regardless.

What a careful firm actually does

Everything above argues one conclusion: the caution is right, but it is an implementation specification, not a reason to abstain.

A careful firm adopts with: business-tier agreements whose data terms are read and kept on file, including subprocessor and support-access clauses; firm-managed accounts rather than personal ones; retrieval-first deployments; permission boundaries so client files serve the workflow and not the internet; redaction where proportionate; human review designed in, with named reviewers and recorded sign-off; a vendor register with due-diligence dates; a consent workflow for client permission, including the honest answer for refusal, which is that the affected work stays manual and gets documented; and a baseline measured before anything changes.

Under the Board’s draft guidance, that list is close to a description of compliance. Firms that build it satisfy the duty and get the hours back from the same project.

Four versions of the same fear

Accountants. The Tax Act is public. Structuring knowledge is sold in every CPD course and printed in ATO rulings, and the platforms now ship their own assistants for exactly this work. Millions of structurally identical ledgers are already reflected in anything a model learned from public filings and documentation. Clients pay for judgement on their affairs and for someone answerable when the ATO asks. Neither transfers through a copied file.

Bookkeepers. Double-entry is five hundred years old and the reconciliation procedures live in the software’s own public guides. What distinguishes a bookkeeper is reliability, speed and trust with a specific client’s records. A thief cannot extract reliability from stolen files.

Financial advisers. Statements of Advice follow regulatory templates, and the strategies are taught in every diploma course. Advisers already send client data daily to platforms, paraplanners and licensee systems under contract. A model ingesting a thousand SOAs learns the genre. It never learns your clients.

Law firms. The strongest version of the fear, and it deserves its due: the precedent bank feels like the crown jewels. But the law is public, judgments are published, and precedent clauses circulate through databases, briefs and discovery. Your opponent has probably seen a version of your best clause. What clients buy is privilege management, judgement on their matter, and professional indemnity behind the work. For lawyers most of all, the conclusion is precision: matter data stays inside systems built for it, retention terms verified, verification steps recorded, privilege treated as the design constraint it is. Fear, correctly aimed, becomes a specification.

The actual choice

Nobody serious is asking a firm to choose between paranoia and recklessness. The choice is between governed adoption and ungoverned drift. One builds controls, measures results, satisfies the regulator and hands back capacity. The other loses the same data through doors nobody watches, and watches the faster firm across town compound the difference.

The firm that refuses AI is not protecting its clients. It is declining to control what already happens.

If you want to know where governed adoption would start in your practice, the five-minute assessment scores where your firm sits on workflow, operational memory and readiness for review-first AI.

Sources and notes: CPA Australia Business Technology Report 2025 (fieldwork July to September 2025, 1,117 accounting and finance professionals). CA ANZ and Ipsos member survey 2025. Wolters Kluwer Future Ready Accountant 2025 as reported by Accountancy Age (secondhand). Tax Practitioners Board draft guidance, March 2026, consultation open at time of writing. Privacy Act automated decision-making amendments commencing 10 December 2026 for APP entities above $3 million turnover, subject to carve-outs. Provider data-use policies from OpenAI, Anthropic, Google and Mistral published terms and help centres, accessed August 2026, all subject to change. The opening scenario is illustrative, the scale arithmetic is illustrative, and the capacity estimate is an illustrative estimate, not a measured result. The author operates a consultancy that builds governed AI systems for professional-services firms.

Next step · The assessment

Turn the general question into one workflow worth examining.

The assessment takes about five minutes. It helps identify where repeated work, context and review are getting in the way before deciding what should change.

Take the AI assessment