Epoch AI estimates the stock of high quality public human text at roughly 300 trillion tokens, projected to be fully consumed by 2028.
Research answer

Create a landscape editorial hero image for this Studio Global article: What are the key details about OpenAI and Anthropic purchasing internal corporate data — including employee chats, emails, video recordings,. Article summary: I'll research this thoroughly across multiple angles. Topic tags: general, education, general web, user generated, academic. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clickbait thumbnails, icons, and tiny thumbnail layouts. Make it useful as an illustrative visual, not as factual evidence.
The public internet is running out of clean human-written text, and the world's leading AI labs are scrambling for a new source of fuel: your internal corporate communications. OpenAI and Anthropic are now aggressively purchasing employee Slack messages, emails, video recordings, and code histories from startups—paying up to $300,000 per deal in a quiet but rapidly growing market for proprietary human interaction data.
The push into corporate data is driven by a simple, unavoidable math problem. The most widely cited estimate, from research group Epoch AI, puts the total stock of high-quality human-generated public text at roughly 300 trillion tokens (90% confidence interval: 100 trillion to 1,000 trillion) . At current frontier training consumption rates, that supply is projected to be exhausted between 2026 and 2032, with a median estimate of 2028
. Some analyses suggest high-quality human-written text on the open web may be only 15–20 trillion tokens, pushing exhaustion as early as 2026
. PBS reported in late 2025 that the "gold rush" for chatbot training data could run out by 2026
.
As that ceiling approaches, OpenAI and Anthropic have turned their attention to a previously untapped trove: internal corporate collaboration data . Records of live human interactions—emails containing negotiations, meeting transcripts with real-time decision-making, and Slack threads showing authentic workplace dynamics—offer the kind of organic conversational data that has become scarce on the scraped public web.
A new ecosystem of intermediary firms has sprung up to broker these sensitive transactions. Two names appear most frequently: Mercor and SimpleClosure .
Broader market rates, reported by Reuters in 2024, offer a sense of the value per unit: approximately $0.001 per word for text, $1 to $2 per image, $2 to $4 per short video, and $100 to $300 per hour of longer film or video .
SimpleClosure, a startup wind-down firm, has processed nearly 100 such transactions over the past year, with payouts ranging from $10,000 to $100,000 per company . The firm's "Asset Hub" platform helps founders scrub personally identifiable information (PII) from the data before licensing it—though the effectiveness and consistency of that anonymization remain uncertain and a growing point of contention .
The lengths AI companies will go to secure training data were laid bare in the extraordinary case of Project Panama, Anthropic's secret internal program to build a massive book training dataset .
The project had two tracks. First, Anthropic purchased physical print books, cut off their bindings, and destructively scanned them page by page, converting them into searchable PDFs and discarding the paper originals. The goal: convert up to 2 million books in six months . An internal memo advised using a codename "because we don't want it to be known that we are working on this" .
Second, the company downloaded more than 7 million pirated ebook files from shadow libraries including LibGen and Pirate Library Mirror . That track led to a class action lawsuit filed in 2024 by authors Andrea Bartz, Kirk Wallace Johnson, and Charles Graeber, who accused Anthropic of using nearly half a million pirated books to train Claude .
On July 20, 2026, a federal judge in San Francisco approved a $1.5 billion settlement—the largest known copyright payout in U.S. history . More than 500,000 authors and publishers were part of the class action, and they will receive an estimated $3,000 per eligible work .
Crucially, the court drew a legal distinction that has major implications for AI training. Using pirated ebooks was infringing and triggered the settlement. But buying physical print books and destructively scanning them internally for training was ruled to potentially qualify as fair use, because each physical copy was replaced with a single digital copy in a process the court called "exceedingly transformative" . The $1.5 billion covered the pirated book liability; the physical destruction program was not penalized .
As companies rush to monetize internal data, significant questions remain. The effectiveness of data anonymization by brokers like SimpleClosure is uncertain and a growing point of contention . Employee Slack messages and emails contain deeply personal and sensitive information, and it is not clear whether current stripping techniques are sufficient to prevent re-identification.
Meanwhile, training costs are spiraling. A single GPT-4-scale training run may already cost $100 million or more. As high-quality public data runs out, the combined costs of licensing proprietary corporate data, paying landmark settlements like Anthropic's $1.5 billion, and building synthetic data pipelines are all escalating dramatically. The broader AI training dataset market is projected to grow rapidly, from an estimated $3.6 billion in 2025 to $23 billion by 2034, as human text licensing becomes a formalized asset class .
For individual users, the landscape has shifted. As of late 2025, both OpenAI and Anthropic changed their default policies to allow training on consumer-tier chats (ChatGPT Free and Plus, Claude Free, Pro, and Max) unless users actively opt out . Enterprise and API customers remain excluded from training by default under contractual Data Processing Agreements
.
The message is clear: in the race to feed AI's insatiable hunger for human-generated data, the boundaries between public and private, personal and commercial, are being redrawn—one Slack archive at a time.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Epoch AI estimates the stock of high quality public human text at roughly 300 trillion tokens, projected to be fully consumed by 2028.