What is the size of the training set for GPT-3
Does the size of the data and openai api usage related?
How much data were GPT models trained onb& how big are the final models?
AI Search using big ammount Data without VECTOR
I tried to find numbers using google, GPT itself (3.5 and 4) and reddit.
-
For training data set size I find numbers in the range <100Gb.
-
For the model sizes (guess that means size of all x-million parameters) I find values in the range of >500 Gb.
This makes me wonder:
-
Are the ballparks of these numbers correct? Should/can it be quantified in Gb?
-
Is it true that those models have much "more parameters than training data"? I read somewhere generalization can get better for such "underconstrained" models...
-
Somewhere I read they were trained on large parts of "the internet". Is that realy only <100 Gb?
-
Presumbly, the paywalled internet is mostly not included in that training set. How much better would thr modes be when high quality content (good newspapers, entire recent book catalogues, scientific literature/papers more than just abstracts...) would be added to the corpus? Wouldn't that blow accuracy out of the water?
Appreciate any answers. Thankyou!