Your own data is the part AI never scraped
Turn the records you already own into a context pack a model can use, and check who keeps a copy before you paste any of it in.
on this page · 0 / 0 checked
You have four years of proposals on a disk somewhere, a folder of finished client work, and an inbox that quietly records every objection anyone has ever raised about your prices. None of it is doing anything. Meanwhile you open a new chat, and the first ten minutes go on explaining what your business does, who this client is, what you mean by “the usual format”, and why the last version was wrong. The model is fluent about everything on the public web and knows nothing about the one subject you are actually paid for.
The large labs have a version of the same shortage, and they are spending real money on it [1][3]. That spending is the clearest signal available about what is scarce now, and it is not intelligence. It is specific, situated, non-public records of how work actually goes. You already own a small pile of exactly that. This guide is about turning it into something a model can use, and checking who else gets to keep a copy before you hand it over. It is not for people training models, building robots, or running a compliance function that already dictates their tooling.
The shortage that made data expensive
Language models trained on text that already existed, in enormous quantity, sitting on the open web. Epoch AI puts the total effective stock of human-generated public text data on the order of 300 trillion tokens, with a 90% confidence interval of 100 trillion to 1,000 trillion, and gives an 80% confidence interval that the stock is fully utilised at some point between 2026 and 2032 [1]. The exact year matters less than the shape of the problem: the free pile is finite, and everyone is drawing from the same one.
So the money moved upstream, to manufacturing data rather than collecting it. NVIDIA sells Cosmos as world foundation models that “generate infinite plausible futures from text, image, video, ambient sound and action input”, pitched as a way to “train physical AI without being constrained by what’s been physically captured”, including amplifying existing data diversity with new weather, lighting and geolocation data for autonomous vehicle training [2]. In August 2026, Veeda AI, led by Sanja Fidler, who joined NVIDIA in 2018 to help establish its Toronto research unit, which later became the company’s Spatial Intelligence Lab, raised $90 million in seed funding co-led by Khosla Ventures and Radical Ventures, three months after the business was founded [3]. The pitch is that robots need to learn in a simulated reality rather than through real-world trial and error [3]. That is a large cheque for a company whose product is experience that never happened.
The durable lesson is not about robots. It is that generic data is now a commodity and specific data is not. Epoch’s assessment of the alternatives is blunt: synthetic data “has only been shown to reliably improve capabilities in relatively narrow domains like math and coding”, and non-public data “seems unlikely to be used at scale due to legal issues, and because it is fragmented over several platforms controlled by actors with competing interests” [1]. Read that second half again from where you sit. Fragmented, privately held, hard to acquire records are the constraint. You hold some.
Your files are the part that was never scraped
Every model you can buy is fluent in the general shape of your industry. None of them has read the email where a client explained why they killed the project, or your quote for the job you underpriced, or the three drafts a specific editor rejected before the fourth one landed. Genre conventions are in the pile. Your particulars are not. When a model’s output feels generic, that is usually the honest answer to a generic question, not a defect in the model.
The material worth collecting is duller than people expect. Finished deliverables you were paid for, in their final approved form. Proposals and quotes with the real numbers left in. The scoping questions you always end up asking. Client-specific facts, like which stakeholder signs off and what they hate. Decisions and their reasons, especially reversed ones. Your own edits, meaning the before and after of a draft you fixed, because the difference between them is the only written record of your taste. Support or sales replies you have sent more than twice.
None of this needs to be large. Keep it to something you could read in one sitting, and treat density rather than volume as the goal: real names, real numbers, real rejected versions. Half a page of your actual pricing logic beats 40 pages of your website copy, which was written to impress strangers rather than to be true.
Build a context pack, not a prompt library
Prompt libraries decay because they encode phrasing, and phrasing is the part models get better at without you. Context packs hold up because they encode facts, and facts stay yours. Make one folder of plain markdown files and keep it to about five. One on the business: what you sell, to whom, what you refuse. One on offers and prices, with the real rates and the discounts you actually give. One per major client or account, with names, constraints and history. One of decisions, dated, each with a sentence on why. One of voice samples, meaning two or three pieces you would be happy to be judged on, pasted in full rather than described.
Plain text and markdown, not PDFs, and certainly not screenshots. The rule that matters most is the same one that governs all model work: paste the actual material, not a description of it. “We write in a friendly but professional tone” is worth nothing. Two full emails you actually sent are worth the whole file.
Then attach the pack to whatever standing container your tool provides rather than re-pasting it. Claude’s Projects are included on both the Pro and Team plans, and exist to hold chats and documents together in one place [5]. Whatever the vendor calls it, the test is whether you can update one file and have the next twenty conversations inherit the change. If you find yourself pasting the same three paragraphs into new chats, the pack is not installed, it is just stored.
Keep the pack somewhere you can leave with
The pack is an asset only while you can carry it. Keep the canonical copy in ordinary files in a folder you control, synced or backed up like anything else you would hate to lose, and treat every AI tool as a renter of that folder rather than its landlord. Anything you can only get back by copying it out of a chat window one message at a time is not really yours.
That principle also decides how you handle the by-products. The outputs a model writes from your pack are derivative and cheap to regenerate; the inputs are not. So the discipline is to promote things upward. When a chat produces a decision, a number or a phrasing you will reuse, it goes into a file in the pack. When a chat produces a nice paragraph, it goes to the client and nowhere else. This is also why model changes stop being disruptive. The version numbers on the pricing page turn over every few months, and a pack made of your own facts survives the turnover, because nothing in it depends on which model reads it.
Read the training setting before the archive goes in
Consumer plans and business plans handle your material differently, and the difference is a policy setting, not a technical one. Anthropic trains new models using data from Free, Pro and Max accounts when the setting is on, with retention extended to five years for new or resumed chats and coding sessions; leave it off and the existing 30-day retention applies, and you can change your choice in Privacy Settings at any time [4]. That does not reach Claude for Work, Claude for Government, Claude for Education, or API use, including via third parties such as Amazon Bedrock and Google Cloud’s Vertex AI [4]. The pricing page states the same split plainly: model training is “opt-out” on Free, Pro and Max, and “none by default” on Team and Enterprise [5].
OpenAI draws the line in the same place. By default it does not use business data from ChatGPT Business, Enterprise, Edu and the API Platform to train models, and training happens only where you have explicitly opted in; API inputs and outputs may be retained for up to 30 days to provide the services and identify abuse, with zero data retention available for eligible endpoints on a qualifying use case [6]. For individual accounts, content may be used to train models, opting out applies to new conversations rather than retroactively, and Temporary Chat does not appear in history, create memories, or get used for training [7]. Google is the most explicit about human eyes: a subset of Gemini Apps chats are reviewed by human reviewers, including Google’s trained service providers, those reviewed chats are retained for up to three years even if you delete your activity, and Google tells you outright not to enter confidential information you would not want a reviewer to see [8]. Turn Keep Activity off and chats are retained with your account for 72 hours, rather than the default 18-month auto-delete [8].
Two practical conclusions. First, your context pack is exactly the kind of material these settings were written about, because it contains client names, prices and internal reasoning that are not yours alone to donate. Second, the fix is cheap. A Claude Team standard seat is $25 a month billed monthly, or $20 billed annually, against $20 a month for Pro billed monthly, and Team is sold for teams of 2 to 150 [5]. Roughly $5 a seat a month is the difference between “opt-out” and “none by default” on the same underlying models.
What still goes wrong
The pack goes stale, quietly. Prices change, a client leaves, a policy you wrote down gets reversed in a conversation nobody wrote down, and the model keeps confidently repeating the old version because you told it to. Stale context is worse than no context, because it is wrong with your authority behind it. Re-reading the whole pack every quarter and deleting aggressively is the only fix, and it is boring enough that most people skip it.
The permissions problem is real and does not go away by choosing a better vendor. Much of the good material is not solely yours. Client work, contracts, other people’s emails and anything covered by an NDA belong to somebody else, and your plan’s training setting does not settle whether you were allowed to share it in the first place. Google’s advice to avoid entering confidential information into Gemini Apps [8] is worth reading as a general principle, not a Google-specific one. If in doubt, strip names and numbers before pasting, or work on a plan whose default is no training [5][6].
Finally, do not over-read the gold rush. A $90 million seed round for simulated worlds [3] is a bet by professionals about a robotics bottleneck, not proof that your folder of invoices is an asset with a market value. Nobody is going to buy your context pack. Its worth is entirely internal: it removes the first ten minutes of every conversation, and it makes the output specific enough to use. That is a real gain, measured in hours, and it is the only one being claimed here.
- 01Epoch AI — Will we run out of data? Limits of LLM scaling based on human-generated dataepoch.ai
- 02NVIDIA — Cosmos world foundation modelsnvidia.com
- 03SiliconANGLE — Sanja Fidler's world model startup Veeda AI raises $90M in seed fundingsiliconangle.com
- 04Anthropic — Updates to Consumer Terms and Privacy Policyanthropic.com
- 05Anthropic — Claude plans and pricingclaude.com
- 06OpenAI — Enterprise privacyopenai.com
- 07OpenAI — How your data is used to improve model performancehelp.openai.com
- 08Google — Gemini Apps Privacy Hubsupport.google.com