🌐
Substack
clouddb.substack.com › cloud database report › report: openai is shopping for 5 exabytes of data storage
Report: OpenAI Is Shopping for 5 Exabytes of Data Storage
January 20, 2026 - Until now, much of the industry discussion has been about the fact that AI training consumes vast amounts of data from across the Web. But it’s becoming increasingly clear that AI systems and apps are generating data at unprecedented scale. We don’t know what kinds of data OpenAI needs to store and manage or how it’s being used.
Discussions

What is the size of the training set for GPT-3
I’m having difficulty finding the size of the data used to train GPT-3. Searches return wildly divergent answers, anywhere from 570GB to 45TB. Language Models are Few-shot Learners would seem to be the definitive source. The largest training set was CommonCrawl which “. . . was downloaded ... More on community.openai.com
🌐 community.openai.com
5
0
September 8, 2023
Does the size of the data and openai api usage related?
I’m trying to build a chatbot that interacts with my own data by translating natural language questions into SQL queries and then querying the database to get the final answer. I’m using langchain to get the work done. Does the size (volume) of the database and the OpenAI API costs have ... More on community.openai.com
🌐 community.openai.com
2
0
June 7, 2023
How much data were GPT models trained onb& how big are the final models?
Here's the original paper for GPT3 . The raw (before filtering) training set was 45TB of compressed plaintext. After filtering, it was ~570GB. It used text scraped from the internet, wikipedia and books. The final model is ~175B parameters. We don't really have details about the training set or parameter count for GPT4. The model is assumed to be around 1T parameters. They intentionally left these details out of the paper. Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar. More on reddit.com
🌐 r/ChatGPT
3
1
May 8, 2023
AI Search using big ammount Data without VECTOR
I’m building an AI-powered search engine using OpenAI, where I need to process around 4,000 titles. The goal is to send a search query and have GPT suggest relevant titles from this dataset. However, there are challenges: Including all 4,000 titles in a single prompt would exceed token limits ... More on community.openai.com
🌐 community.openai.com
3
0
November 26, 2024
🌐
OpenAI
openai.com › index › inside-our-in-house-data-agent
Inside OpenAI’s in-house data agent | OpenAI
In this post, we’ll break down ... than 3.5k internal users working across Engineering, Product, and Research, spanning over 600 petabytes of data across 70k datasets....
🌐
Interface
interface.media › home › “big data” isn’t big enough to train generative ai
“Big Data” isn’t big enough to train generative AI - Interface
March 25, 2024 - Training a large language model like the one that fuels OpenAI’s ChatGPT takes a lot of data. It took approximately 570 gigabytes of text data–about 300 billion words—to train ChatGPT.
🌐
AI Tools Club
aitoolsclub.com › how-openai-built-an-ai-data-agent-that-turned-days-of-analysis-into-minutes
How OpenAI Built an AI Data Agent That Turned Days of Analysis Into Minutes
February 23, 2026 - As for no one's surprise, OpenAI operates at enormous scale, and its internal data platform is staggering, spanning over 600 petabytes of data spread across roughly 70,000 datasets, serving more than 3,500 internal users across Engineering, ...
🌐
Compare Internet
compareinternet.com › home › blog › how much data does ai use? chatgpt, image gen, and video gen compared
How Much Data Does AI Use? Get the Full Rundown Here - Compare Internet
July 14, 2026 - Text-based AI like ChatGPT, Claude, and Gemini in chat mode actually don’t use as much data as you might expect. Usually, they use up 1 to 5 MB per hour of conversation, which is less than it takes to load a single photo-heavy webpage.
🌐
OpenAI Developer Community
community.openai.com › chatgpt
What is the size of the training set for GPT-3 - ChatGPT - OpenAI Developer Community
September 8, 2023 - I’m having difficulty finding the size of the data used to train GPT-3. Searches return wildly divergent answers, anywhere from 570GB to 45TB. Language Models are Few-shot Learners would seem to be the definitive source.
🌐
Milvus
milvus.io › ai-quick-reference › how-does-openai-handle-large-datasets
How does OpenAI handle large datasets?
Built for AI.Try Free Now → ... OpenAI handles large datasets through a combination of distributed computing, efficient data processing pipelines, and specialized infrastructure.
Find elsewhere
🌐
MIT Technology Review
technologyreview.com › artificial intelligence › openai’s hunger for data is coming back to bite it
OpenAI’s hunger for data is coming back to bite it | MIT Technology Review
April 20, 2023 - In AI development, the dominant paradigm is that the more training data, the better. OpenAI’s GPT-2 model had a data set consisting of 40 gigabytes of text. GPT-3, which ChatGPT is based on, was trained on 570 GB of data.
🌐
OpenAI Developer Community
community.openai.com › api
Does the size of the data and openai api usage related? - API - OpenAI Developer Community
June 7, 2023 - I’m trying to build a chatbot that interacts with my own data by translating natural language questions into SQL queries and then querying the database to get the final answer. I’m using langchain to get the work done. Does the size (volume) of the database and the OpenAI API costs have ...
🌐
Reddit
reddit.com › r/chatgpt › how much data were gpt models trained onb& how big are the final models?
r/ChatGPT on Reddit: How much data were GPT models trained onb& how big are the final models?
May 8, 2023 -

I tried to find numbers using google, GPT itself (3.5 and 4) and reddit.

  • For training data set size I find numbers in the range <100Gb.

  • For the model sizes (guess that means size of all x-million parameters) I find values in the range of >500 Gb.

This makes me wonder:

  1. Are the ballparks of these numbers correct? Should/can it be quantified in Gb?

  2. Is it true that those models have much "more parameters than training data"? I read somewhere generalization can get better for such "underconstrained" models...

  3. Somewhere I read they were trained on large parts of "the internet". Is that realy only <100 Gb?

  4. Presumbly, the paywalled internet is mostly not included in that training set. How much better would thr modes be when high quality content (good newspapers, entire recent book catalogues, scientific literature/papers more than just abstracts...) would be added to the corpus? Wouldn't that blow accuracy out of the water?

Appreciate any answers. Thankyou!

🌐
OpenMined
openmined.org › home › blog › no, ai hasn’t run out of data
No, AI hasn’t run out of data — OpenMined
December 3, 2025 - Around 80% of the tokens used to train OpenAI’s GPT-3 model came from Common Crawl, and between 2019 and 2023, more than 60% of all published LLMs relied on Common Crawl for their training data.
🌐
Substack
clouddb.substack.com › p › how-much-more-data-will-ai-generate
How Much More Data Will AI Generate? 10x, 100x, 1000x?
August 23, 2024 - They’re now the most popular type of databases tracked by DB Engines, according to Alex Woodie at Datanami. The DB Engines website lists 15 vector databases and that jumps to 26 if you include databases with a secondary model.
🌐
OpenAI
openai.com › index › approach-to-data-and-ai
Our approach to data and AI | OpenAI
May 7, 2024 - Unlike larger companies in the AI field, we do not have a large corpus of data collected over decades. We primarily rely on publicly available information to teach our models how to be helpful.
🌐
DataRobot
datarobot.com › home › blogs › how much data is enough for ai?
How Much Data Is Enough for AI? | DataRobot Blog
March 17, 2025 - It leads to lower complexity and lower maintenance costs and a simpler and more accessible software stack that allows us to iterate and get value faster. Our COVID-19 decision intelligence platform was built without using established big data approaches, despite having sufficiently large data sizes, and not using them allowed our data scientists to use familiar tools and get results quicker.
🌐
Stack Overflow
stackoverflow.blog › 2024 › 10 › 17 › training-data-scarcity-synthetic-quality-model-genai-ai
Brain Drain: David vs Goliath - Stack Overflow
October 17, 2024 - One of the most striking things about today's generative AI models is the absolutely enormous amount of data that they train on. Meta wrote, for example, that its Llama 3 model was trained on 15 trillion tokens, which is equal to roughly 44 ...
🌐
OpenAI Developer Community
community.openai.com › prompting
AI Search using big ammount Data without VECTOR - Prompting - OpenAI Developer Community
November 26, 2024 - I’m building an AI-powered search engine using OpenAI, where I need to process around 4,000 titles. The goal is to send a search query and have GPT suggest relevant titles from this dataset.
🌐
PubMed Central
pmc.ncbi.nlm.nih.gov › articles › PMC8164167
Big Data Requirements for Artificial Intelligence - PMC - NIH
Checking your browser before accessing pmc.ncbi.nlm.nih.gov · Click here if you are not automatically redirected after 5 seconds