DOI : 10.17577/Type a question into an AI chatbot and the answer arrives so smoothly that it’s easy to forget the model never looked anything up. Everything it says comes from patterns it picked up while reading an enormous pile of text during training.
That raises an obvious question, and it’s one most users never ask. Whose text was it, and how did it get there?
The honest answer is less mysterious than people expect. Most of it started life as ordinary web pages. Getting from a web page to something a model can learn from is a long, fairly unglamorous engineering job, and this article follows that job from start to finish.
The main sources
Training data for a language model usually comes from a handful of places, mixed together in different proportions.
- The open web. Articles, forums, documentation, product pages, question-and-answer sites. This is the biggest share by a wide margin, simply because nothing else exists in the same volume.
- Books and long-form writing. Smaller in volume, but carefully edited, which makes it valuable.
- Source code. Public repositories teach models to read and write programs.
- Licensed archives. Publishers and platforms sell or license access to their content.
- An organization’s own material. Support tickets, manuals and internal documents, used mostly to adapt an existing model to a specific job.
- Text written for the purpose. People are paid to write examples, rate answers and correct mistakes.
The last few sources are small but they punch above their weight. The web supplies the bulk, and that’s where the collection pipeline does most of its work.
From a list of URLs to raw pages
Collection starts with a list of addresses. Some teams build that list themselves from sitemaps and links. Many skip the step entirely and download a public web archive that somebody else has already crawled, which is cheaper and much quicker.
Either way, a program called a crawler does the fetching. It asks a server for a page, saves what comes back, pulls out the links, and adds the new ones to its queue. Then it does it again, millions of times.
A well-built crawler saves more than the page. It records the address, the time of the request, the status code the server returned and a fingerprint of the content. That record looks like bookkeeping, and it is. It’s also the only way to answer the question “where did this sentence come from?” six months later.
Most of what’s collected gets thrown away
Here’s the part that surprises people. The raw crawl is mostly junk, at least from a model’s point of view. A typical web page is wrapped in menus, cookie banners, sidebars, footers and ads. Plenty of pages are error messages, empty templates or the same press release copied across hundreds of sites.
So the bulk of the pipeline is cleanup, and it happens in a fairly standard order.
- Text extraction. The markup is stripped out and the main content is separated from the navigation around it.
- Language detection. Each document is labeled by language, so an English dataset doesn’t quietly fill up with other languages, and the other way round.
- Quality filtering. Simple rules catch a lot: pages that are mostly links, lines repeated dozens of times, text with almost no punctuation, strings of keywords written for search engines.
- Deduplication. Exact copies are easy to find. Near copies, where two pages differ by a date or a byline, take cleverer comparison. Both have to go, because a model that sees the same paragraph a thousand times learns to recite it.
- Personal data handling. Email addresses, phone numbers and similar details are removed or masked.
- Safety filtering. Sources known for malicious or adult content are dropped.
What comes out the other end is a small fraction of what went in. That’s normal. Teams that publish their methods tend to describe this stage in far more detail than the crawling, because it has the bigger effect on how good the finished model is.
When a ready-made dataset isn’t enough
If a public archive already contains what you need, use it. Running your own crawl costs money and engineering time, and it brings responsibilities with it.
Still, there are good reasons to collect your own data. Public archives are snapshots, so they go stale. They sample the web broadly, which means they may hold only a few pages from the one site you care about. Smaller languages and regional sites are often thinly covered. And pages that build their content with JavaScript can come back half empty unless a real browser renders them.
A university group building a corpus in a regional language runs into this. So does a company that wants a model fluent in the vocabulary of its own industry.
The fetch stage and the IP address problem
Once a team crawls for itself at any real volume, it meets a problem that has nothing to do with code quality.
Web servers keep count of how many requests arrive from each IP address. Cross a threshold and the server slows you down, shows a CAPTCHA or refuses outright. A crawler running on a few cloud machines hits that wall quickly. Cloud addresses are also easy for a site to recognize, and some sites answer them with a challenge page.
That last case is worse than a plain block, because the crawler doesn’t know anything went wrong. It saves the challenge page as though it were the article. Unless somebody checks, that page ends up in the dataset, and nobody finds out until much later.
Location causes a quieter version of the same trouble. Many sites change what they show depending on where the visitor appears to be: the language, the products, the local news. A crawler sitting in one country collects one country’s version of the web and never finds out what it missed.
This is the job proxies do in a data pipeline. Requests go out through a large pool of addresses, so each address carries only a light load. And the crawler can choose the country a request appears to come from, which is how a team collects text as local readers actually see it. ProxyEmpire has a longer guide on how proxies support LLM training that goes into which kind of proxy fits which kind of source.
It’s worth being clear about the limits. A proxy changes where a request comes from. It has no bearing on whether the request should be made in the first place. And a bigger pool is a reason to go easier on each address, with the same polite pace for the site overall.
Asking permission
Websites have a simple way to tell crawlers what they will and won’t accept: a small text file called robots.txt, sitting at the root of the site. It lists which crawlers may visit and which paths are off limits.
Nothing technical enforces it. A crawler can ignore the file, and the server will usually still answer. Respecting it is a matter of good conduct, and for AI training it has become the baseline that everyone expects. Many site owners now add rules aimed specifically at AI crawlers, saying in effect “index me for search, but don’t train on me.”
A responsible pipeline checks those rules before every fetch and keeps a record of what it skipped. The law around training data is still being worked out and differs between countries, so anyone collecting at scale should get proper legal advice. The engineering habit is the same everywhere, though: know what each site asked for, and be able to show that you followed it.
What this means if you’re building a dataset
For students and researchers putting together their own corpus, a few habits save a lot of pain.
Look for an existing dataset first, and write down exactly what it lacks before you fetch anything. Keep the raw pages separate from the cleaned text, so you can improve your filters later without crawling all over again. Open a random sample of what you collected and read it yourself, because block pages and login screens look like success to a script. Remove duplicates early, while it’s cheap. And document as you go: the sources, the filters, what you left out and why.
It’s easy to think of an AI model as a feat of mathematics and hardware. A good part of it is something plainer: a very large reading list that somebody had to assemble, clean and take responsibility for. Knowing how that list gets made is the first step to judging what a model can be trusted to know.
