Who provides data to artificial intelligence?
Who provides data to artificial intelligence?\n\nAI data comes from many different sources, but the main one is the web. During training, models learn from large amounts of text, code, images and other content collected from public sources, licensed sources or partner providers.\n\nIn real-time answers, instead, AI can use indexed content, accessible documents online, databases, APIs and sources the system considers reliable. So there is no single provider: data comes from publishers, websites, archives, platforms and users, according to specific rules and agreements. The quality of the answer depends a lot on the quality and structure of the available data.
Want to see AImpeto more often on Google?
Add us to preferred sources: we will show up more often in results and AI Overviews.

The main data sources for artificial intelligence
The data used to train models comes from a combination of public and private sources, commercial agreements and direct user contributions. None of these alone is enough: the quality of the result depends on the variety and reliability of the whole set.
The web is the most important source. Data is collected from the open web with automatic scraping tools (the so-called web scraping) from open archives such as Common Crawl, news sites, encyclopedias like Wikipedia, forums like Reddit and code platforms like GitHub. Alongside these come databases from public bodies and institutions, health, geography and statistics data, and specialized archives in science or finance.
Users and contributors also play a role. Daily interactions, prompts entered in chats, feedback and the work of trainers who correct the answers, often with techniques like reinforcement learning from human feedback (RLHF), help make the model more precise.
How AI uses the data it receives
The process splits into two distinct phases. During training, the model reads large amounts of text and images and learns to recognize recurring patterns: for example which words tend to appear one after another, or how concepts are structured in different contexts. This statistical analysis creates a base of knowledge.
When it answers, though, the AI does not draw directly from a finite archive. It calculates the most likely combination of words based on everything it learned during training. In some cases the AI can connect to the internet in real time to read fresh news and give an answer based on up-to-date data, as happens with the modern browsing features of ChatGPT and with Perplexity.
OpenAI and ChatGPT: the mix between base knowledge and live search
OpenAI explains that ChatGPT training data comes from large amounts of text freely collected from the web, first passing filters that remove unwanted content such as hate speech, spam, adult material or personal data. Paywalled content and dark web data are not used.
For GPT-3, the base architecture of ChatGPT, the company used different collections: a slice of Common Crawl from 2016-2019, WebText2 (pages linked from Reddit with at least three upvotes), English Wikipedia and large collections of digital books. These are joined by data from business partners, such as proprietary databases or high-quality curated datasets.
ChatGPT base knowledge is static and reflects the web up to what is called the cutoff, a date beyond which it does not update its memory on its own. This is where real-time search comes in: with the browsing function, ChatGPT can consult the web on request instead of relying only on the pre-trained model.
Google Gemini and its proprietary knowledge
Gemini, the family of Google models, is also trained using publicly available information and proprietary material owned by Google. Most answers rely on this pre-trained knowledge, built from public sources and the company's internal data.
Only in more advanced modes, such as the Deep Research feature, Gemini runs real-time searches through the Google engine to find updated or more detailed information. This shows the strong link between AI results and the content indexed by the search engine: in practice, any public content that is not explicitly protected can feed the Gemini dataset, from news pages to forum posts.
Perplexity: search always active
Perplexity stands out because it was built to search the web. Unlike ChatGPT, its answers are not based mainly on static base knowledge, but on content actually available online. Without the user needing to click anything, each query runs a real search, and the model builds the answer with reference to the original sources.
Perplexity has its own index that tends to favor well-known and popular web sources. This means that niche or lesser-known content gets less attention, while the most cited and trustworthy sites are more likely to appear. It is an important signal for people working on SEO: to be visible on Perplexity, domain authority and content quality matter a lot.
Need a GEO consultation?
The AImpeto team helps you get cited by ChatGPT, Gemini, Copilot and Perplexity.
Request a consultationMore GEO questions
Explore the other aspects of Generative Engine Optimization: what it is, how the strategy works, the differences from SEO and AEO, and how to choose an agency.
Written by

Marco Modonesi
Co-Founder of AImpeto
SEO and GEO expert with over 17 years of experience in digital marketing. Marco helps businesses get found and cited by artificial intelligence, combining strategy, content and automation.
Follow me on LinkedIn