
LLM training data is the text, code, structured records, and human feedback used to change a language model's parameters. Multimodal language models may also train on paired images, audio, or video. The term is used more broadly for retrieval corpora and evaluation sets, even though those datasets have different jobs and should not be managed as one undifferentiated pool.
The practical challenge is not simply collecting more tokens. A useful dataset needs a defined purpose, traceable sources, documented usage rights, reliable extraction, quality controls, deduplication, contamination checks, representative coverage, versioned splits, and evaluations that remain separate from training.
Quick Answer
LLM training data is the material used to pretrain, continue training, fine-tune, or align a language model. RAG documents and evaluation sets support the system without necessarily changing model weights. High-quality LLM data is relevant, traceable, legally reviewed, deduplicated, representative, and separated from held-out tests. Proxies matter only when an authorized web-collection workflow needs a different network route or location.
Key Takeaways
- “LLM data” is not one dataset. Pretraining corpora, fine-tuning examples, preference data, RAG indexes, and evaluation sets serve different purposes and require different formats and controls.
- There is no universal token requirement. Data volume depends on the model, compute budget, training objective, quality, expected inference demand, and whether you are training from scratch or adapting an existing model.
- Provenance must begin at collection. Store the source, retrieval date, content hash, collection method, applicable license or review status, transformation history, and final dataset split.
- Deduplication and decontamination solve different problems. Deduplication reduces repeated material; decontamination protects evaluations from training overlap.
- Publicly accessible does not mean unrestricted for AI use. Copyright, database rights, privacy, contracts, website terms, and sector-specific rules can all matter. Obtain qualified legal advice for the jurisdictions and data involved.
- A proxy is part of the acquisition layer. It can change the exit IP and provider-supported location, but it does not grant permission, parse pages, verify facts, remove personal data, deduplicate documents, or make a dataset suitable for training.
- Measure usable output, not downloaded volume. Cost per accepted document or accepted million tokens is more informative than requests, pages, or gigabytes transferred.
Methodology: This guide was reviewed on September 29, 2026 against primary research on scaling, retrieval, dataset documentation, deduplication, contamination, and data curation; official guidance from NIST, the U.S. Copyright Office, the EU, the UK Information Commissioner's Office, and Common Crawl; current Proxidize product pages; and the existing Proxidize AI-data collection cluster. We did not train models to benchmark the dataset recipes below. This article provides technical and operational guidance, not legal advice.
What Is LLM Training Data?
In the strict sense, LLM training data consists of examples processed during optimization so that the model's weights change. During pretraining, the objective is commonly next-token prediction over very large corpora. During supervised fine-tuning, examples may pair an instruction with a desired response. Preference or alignment stages can use comparisons, ratings, critiques, demonstrations, or other feedback signals.
In practice, teams often use “LLM training data” as shorthand for every dataset in an AI system. That broader usage is understandable, but it hides important distinctions:
| Dataset | Does it normally change model weights? | Primary job | Typical unit |
|---|---|---|---|
| Pretraining corpus | Yes | Teach broad language, code, knowledge, and representation patterns | Documents or token sequences |
| Continued-pretraining corpus | Yes | Adapt an existing base model to a domain, language, or newer distribution | Domain documents or token sequences |
| Supervised fine-tuning data | Yes | Teach desired task behavior and response formats | Input-output examples |
| Preference or alignment data | Yes, during post-training | Rank or shape model behavior | Comparisons, ratings, critiques, or rewards |
| RAG corpus | Usually no | Supply retrievable knowledge at inference time | Documents, chunks, metadata, and embeddings |
| Evaluation data | No | Measure capability, quality, safety, or regressions | Held-out prompts, references, rubrics, and outcomes |
This distinction matters operationally. Accidentally mixing evaluation examples into pretraining or fine-tuning data can inflate reported performance. Treating a RAG index as if it were a training corpus can lead to unnecessary retraining when the real requirement is fresher retrieval content. Calling every collected page “training data” also makes provenance and deletion requests harder to manage.
For a broader explanation of the model itself, see what an LLM is and how it works.
Pretraining vs. Fine-Tuning vs. RAG vs. Evaluation Data
Pretraining Data
Pretraining teaches a base model general statistical patterns across language, code, and other supported modalities. The corpus can combine web pages, books, papers, code, reference material, licensed collections, and purpose-built data mixtures.
At this scale, data design is a model-design decision. Domain proportions, language balance, code quality, repetition, freshness, and filtering thresholds can change what the model learns. The DataComp-LM benchmark found substantial differences between curation strategies under standardized training and evaluation conditions, illustrating why raw token count alone is not a quality measure.
Continued Pretraining or Domain Adaptation
Continued pretraining starts with a base model and exposes it to additional token sequences. It can adapt vocabulary, style, facts, languages, or domain patterns without requiring instruction-response pairs for every document.
This is useful when the model needs deeper familiarity with a specialized corpus, but it can also shift behavior or weaken capabilities outside that domain. Keep the original and adapted evaluation suites, document the mixture, and test for regressions rather than assuming that more domain text only adds knowledge.
Supervised Fine-Tuning Data
Supervised fine-tuning usually contains an input, an intended response, and sometimes system instructions, context, tool calls, reasoning annotations, or structured labels. It teaches the model how to perform a task or follow a response policy.
Fine-tuning examples need consistency more than sheer volume. Contradictory instructions, mislabeled outputs, answer leakage, and inconsistent formatting can teach undesirable behavior quickly. A smaller reviewed dataset can be more useful than a much larger automatically generated one.
Preference and Alignment Data
Preference data represents judgments between possible outputs. Depending on the post-training method, records may contain chosen and rejected responses, scalar scores, critiques, safety labels, or feedback from humans and other models.
Record who or what produced each judgment, the rubric used, annotator qualifications where relevant, disagreement rates, model versions, and quality-control results. A preference label without its collection context is difficult to audit later.
RAG Data
Retrieval-augmented generation combines a model with an external knowledge source. The original RAG paper described a system combining parametric memory with a retrievable non-parametric index. In a production system, that index can be updated independently of the model's weights.
RAG corpora need source-aware chunking, metadata, access controls, freshness rules, retrieval tests, and citation handling. They do not automatically require model training. If the business need is to answer questions from frequently changing product documentation, updating a RAG index may be simpler and more traceable than repeatedly fine-tuning the model.
Evaluation Data
Evaluation data should remain outside training and prompt-generation pipelines unless a particular experiment explicitly studies that overlap. It can include:
- capability benchmarks;
- domain-specific test cases;
- safety and red-team prompts;
- retrieval relevance judgments;
- factuality and citation checks;
- human preference evaluations;
- production samples used for regression testing.
An evaluation set needs a version, scope, sampling method, scoring rubric, expected uncertainty, and contamination status. A single public benchmark is not proof that a model works for your production distribution.
How Much Data Does an LLM Need?
There is no responsible one-number answer. The required amount depends on:
- whether you are pretraining, continuing pretraining, fine-tuning, building RAG, or evaluating;
- model size, architecture, tokenizer, and context length;
- training compute and optimization recipe;
- corpus quality and repetition;
- domain breadth and language coverage;
- the number of passes over the data;
- the target capability and acceptable error rate;
- expected inference volume and deployment cost.
The 2022 Chinchilla scaling study examined compute-optimal model and token allocation under a specific experimental setup. Its commonly repeated ratio of roughly 20 training tokens per parameter is a useful historical heuristic, not a universal minimum, maximum, or guarantee. Later work has also examined training smaller models for longer when inference demand changes the economic objective.
Use scaling experiments rather than folklore:
- Select representative dataset mixtures and held-out evaluations.
- Train smaller models or shorter runs across several data quantities.
- Measure capability, memorization, subgroup performance, and cost.
- Fit a scaling estimate within the range actually tested.
- Re-evaluate when the model, tokenizer, mixture, or objective changes.
For fine-tuning, learning curves are often more useful than token-to-parameter ratios. Train on progressively larger reviewed samples and stop when additional examples no longer improve the held-out task—or when they begin to harm other requirements.
Where Does LLM Data Come From?
Most production datasets combine several source classes.
First-Party Data
First-party data can include product documentation, support articles, approved transcripts, knowledge bases, code, internal policies, and structured business records. It can be highly relevant, but ownership does not remove privacy, confidentiality, retention, access-control, or purpose-limitation obligations.
Export data through supported first-party APIs or repositories when possible. Record the system of record, access policy, snapshot date, and business owner.
Licensed or Purchased Data
Commercial datasets can offer clearer contractual terms, specialist coverage, annotations, or update commitments. Review exactly what the license permits: research, internal use, model training, fine-tuning, embeddings, redistribution, generated outputs, and derivative datasets may be treated differently.
Do not rely on the vendor's dataset title alone. Retain the agreement version, covered assets, restrictions, term, renewal status, and deletion obligations in the source registry.
Open Datasets
Open datasets can accelerate research, but “open” is not a complete rights description. Review the dataset card, license, source lineage, included subsets, personal-data warnings, and known limitations.
For example, the current FineWeb dataset card documents its Common Crawl sources, processing, filtering, deduplication choices, intended use, license, and limitations. That documentation is more useful than a bare download link, but downstream users still need to assess their own intended use.
Public Web Data
Public websites can supply current, multilingual, and domain-specific information that packaged datasets lack. Collection may use first-party APIs, feeds, sitemaps, ordinary HTTP clients, crawling frameworks, or browser automation.
Accessibility is not the same as permission for every downstream use. Common Crawl's own terms of use warn that crawled content can remain subject to third-party rights and terms. Build source review into the collection plan rather than attempting to reconstruct it after the corpus has been mixed and tokenized.
The implementation details belong in the separate guide to web crawling for AI training data. The broader data-sourcing guide compares APIs, open data, purchased datasets, surveys, and web collection.
Human-Generated and Annotated Data
Humans can write demonstrations, rank responses, label documents, verify facts, construct edge cases, and review failures. Human involvement does not automatically make data correct. Annotation guidelines, training, qualification, double review, adjudication, and agreement measurements determine whether labels are consistent enough for the task.
Synthetic Data
Synthetic data is generated or transformed by models, simulators, templates, or rules. It can expand rare cases, create structured examples, translate formats, and support controlled experiments.
Keep synthetic records labeled. Store the generator model and version, prompts, sampling settings, source context, filters, reviewer status, and generation date. Validate against independently sourced real examples. Otherwise, the dataset can amplify the generator's mistakes, narrow linguistic diversity, reproduce protected or personal material from its context, or contaminate evaluations.
What Makes LLM Training Data High Quality?
Quality is fitness for a defined purpose—not a universal score. A polished encyclopedia paragraph may be useful for general pretraining and irrelevant for teaching a tool-calling schema. A real support transcript may be highly representative and unusable until personal and confidential information is handled appropriately.
Evaluate at least these dimensions:
| Dimension | Question to answer | Example check |
|---|---|---|
| Relevance | Does this material support the intended capability? | Domain classifier plus reviewed samples |
| Accuracy | Are important claims and labels correct? | Source validation, annotator review, reference checks |
| Freshness | Is the content current enough for the task? | Publication and retrieval dates, revisit policy |
| Coverage | Are required topics, languages, formats, and edge cases present? | Distribution report against a target inventory |
| Representation | Are important groups and usage contexts underrepresented or distorted? | Subgroup counts and held-out performance |
| Extractability | Did parsing retain the meaningful content and structure? | Text-to-page comparison on sampled documents |
| Provenance | Can each record be traced to a source and processing history? | Complete source and transformation fields |
| Rights and privacy | Has the intended use been reviewed? | License/terms status, personal-data checks, restrictions |
| Duplication | Is repeated material unintentionally overweighted? | Exact and near-duplicate clusters |
| Contamination | Did training material overlap held-out evaluations? | Benchmark matching before final split |
Filtering can introduce its own bias. Rules based on length, vocabulary, “educational quality,” toxicity, or resemblance to a preferred reference corpus can disproportionately remove certain languages, dialects, communities, or document styles. Test what each filter removes, not only what it retains.
How Should You Track LLM Data Provenance?
Provenance answers where a record came from, how it changed, and why it was admitted to a particular dataset version. NIST's Generative AI Profile includes tracking the provenance of training data and metadata among its recommended risk-management actions. The Data Provenance Initiative also found widespread gaps and errors in dataset licensing and attribution, demonstrating why a hosting-page label should not be the only record.
At minimum, retain:
- source URL, repository, or system identifier;
- source owner or publisher where known;
- retrieval timestamp and collector version;
- HTTP status, content type, and language signals;
- raw-content hash and normalized-content hash;
- applicable license, terms, contract, or review status;
- collection method and access policy;
- parsing, filtering, annotation, and redaction versions;
- duplicate-cluster identifier;
- dataset version and train/validation/test/RAG assignment;
- deletion, opt-out, retention, and restriction status where applicable.
A record can look like this:
rights_review_status should refer to evidence and an accountable review process. It should not be guessed from whether a page was reachable or whether a crawler received HTTP 200.
Publish a dataset card or datasheet describing purpose, composition, collection, processing, known limitations, permitted uses, and maintenance. The Datasheets for Datasets paper provides a useful documentation framework.
How Should LLM Data Be Deduplicated?
Deduplication reduces unintentional repetition and helps prevent one source or template from receiving excessive weight. Research on deduplicating language-model training data found that common datasets contained many near-duplicates and that deduplication reduced memorized output and train-test overlap in the studied setup.
Use several levels:
- URL and identifier deduplication: Collapse normalized URLs, repository IDs, or document IDs where appropriate.
- Exact content deduplication: Hash normalized content after defining which whitespace, casing, and boilerplate changes should be ignored.
- Near-duplicate detection: Use MinHash, locality-sensitive hashing, n-gram similarity, or another scalable method to cluster highly similar documents.
- Template and boilerplate detection: Separate repeated navigation, legal footers, product chrome, quoted threads, and mirrored page frames from the unique body.
- Semantic review: For smaller high-value datasets, identify paraphrases or multiple answers that convey the same example under different wording.
Do not blindly keep one item from every cluster. Repetition can be meaningful: a term may legitimately appear across independent sources, a code pattern may be common, and a long-lived page can occur in several dated snapshots. Define whether the goal is to remove accidental duplication, preserve time series, or deliberately upweight verified material.
Run deduplication before assigning final dataset splits. If one version of a document enters training and a near-identical version enters validation, the evaluation is no longer independent.
How Do You Prevent Evaluation Contamination?
Evaluation contamination occurs when test questions, answers, benchmark explanations, or close variants appear in material used to train or tune the model. The result can look like capability even when the model has memorized recognizable examples.
Practical controls include:
- register protected benchmarks and internal test sets before final curation;
- search for exact question-and-answer matches;
- search normalized n-grams and near-duplicate passages;
- check benchmark repositories, solution pages, discussions, and derived datasets;
- use source-, organization-, or time-based holdouts where appropriate;
- keep private production evaluations that are not used for prompt generation;
- record known or suspected contamination instead of silently discarding the result;
- refresh evaluations as production tasks change.
Decontamination is not a one-time filter. Synthetic-data generators, annotation tools, retrieval indexes, and later fine-tuning rounds can all reintroduce held-out material.
Licensing, Copyright, Privacy, and Data Governance
There is no universal rule saying that all public data may—or may not—be used for every AI purpose. The result can depend on the content, collection method, intended use, license, contract, website terms, personal data, jurisdiction, and applicable exceptions. This is an area for qualified legal review, not a checkbox inferred from technical accessibility.
The U.S. Copyright Office's AI initiative and training report analyzes copyright questions surrounding generative-AI training, while the EU AI Act includes transparency and copyright-policy obligations for providers of general-purpose AI models. The exact obligations depend on the role and jurisdiction; do not turn either source into a universal conclusion about one dataset.
For every source, ask:
- Who owns or controls the source and the individual records?
- What does the applicable license, agreement, or website policy permit?
- Does the planned use involve copying, model training, embeddings, redistribution, or generated outputs?
- Does the material include personal, confidential, sensitive, or regulated information?
- Are data-subject notices, rights, retention, deletion, or opt-out processes required?
- Can the organization identify and remove affected records from derived datasets and indexes?
- Which evidence and decision owner support the approved status?
The UK's Information Commissioner's Office explains that transformed training data can remain personal data and notes separately that collecting personal data is itself processing. Apply data minimization: collect only what the defined purpose requires, restrict access, set retention rules, and design a deletion path before training.
robots.txt communicates crawler preferences; it is not a copyright license, privacy assessment, or complete authorization framework. Respect it where applicable, along with rate limits, access decisions, terms, contracts, and relevant law.
A Practical LLM Data Pipeline
A defensible pipeline keeps raw evidence, transformations, dataset versions, and model evaluations connected:
The ETL pipeline guide explains the extraction and transformation layer. Use data parsing for document structure, data normalization for consistent fields, and the guide to bad data when defining validation and rejection rules.
1. Define the Decision the Data Must Support
Write the intended model or retrieval behavior, target languages, formats, domains, freshness, unacceptable material, and measurable evaluation criteria before sourcing data. Otherwise, the easiest available data will quietly define the system.
2. Build a Source Registry
Give every source an owner, collection method, update policy, rights-review state, sensitivity classification, and removal process. Keep rejected sources in the registry with the reason, so they are not reintroduced by another team.
3. Preserve Raw Evidence
Store an immutable or access-controlled raw snapshot where appropriate, plus hashes and retrieval metadata. This lets you reproduce parsing, investigate an error, or prove which version entered a dataset. Apply retention and access rules; “raw forever” is not a safe default.
4. Parse and Validate Before Counting Tokens
Remove navigation, repeated chrome, cookie banners, broken encoding, unrelated comments, and extraction artifacts. Preserve useful structure such as headings, code blocks, tables, speaker turns, and document hierarchy. A token count before extraction validation can mostly measure boilerplate.
5. Filter, Deduplicate, and Split
Apply documented rules in a reproducible order. Record the acceptance and rejection reason for each stage. Deduplicate before the final split, then run benchmark decontamination. Freeze manifests rather than referring to a mutable folder as “the training set.”
6. Evaluate the Dataset Through the Model
Document-level quality scores are proxies for usefulness. The final test is whether the dataset improves the intended held-out behavior without causing unacceptable regressions, memorization, bias, safety failures, or cost increases.
Where Do Proxies Fit in LLM Data Collection?
A proxy fits between a web collector and the source site:
The proxy's job is narrow: route the configured request through another exit IP and, when supported, another country, city, ISP, or network type. It can be useful when an approved collection workflow:
- needs to observe genuinely localized public content;
- encounters route- or IP-dependent availability;
- requires separate sticky sessions for multi-request page flows;
- needs provider-managed exit selection instead of maintaining individual servers;
- collects permitted public sources across several regions.
A proxy does not:
- determine whether the source may be used for training;
- override a login, paywall, explicit denial, or contractual restriction;
- make high concurrency acceptable to one host;
- extract main content or render JavaScript by itself;
- verify a document's truth, license, language, or relevance;
- remove personal or confidential information;
- deduplicate records or prevent benchmark leakage.
Many LLM data workflows do not need proxies. First-party exports, licensed bulk downloads, internal repositories, official APIs, and openly distributed datasets should normally use their supported delivery methods. Start with the simplest compliant route and add proxy infrastructure only when the network or location requirement is demonstrated.
Raw Proxies vs. Scraping APIs vs. Ready-Made Datasets
| Access model | Your team operates | Provider supplies | Best fit |
|---|---|---|---|
| First-party API or export | Transformation, governance, and downstream pipeline | Supported records or files | Available official access with sufficient fields |
| Ready-made dataset | Validation, rights review, mixing, and model pipeline | A packaged snapshot or stream | Faster experiments and known dataset boundaries |
| Scraping or extraction API | Source policy, schema, validation, and governance | Fetching and often rendering/extraction | Teams that want less retrieval infrastructure |
| Raw proxies | Crawler, browser, retries, rendering, extraction, validation, and storage | Network route and exit pool | Teams that need maximum collector control |
The raw proxies vs. scraping APIs comparison covers the operational and cost trade-offs. If raw proxies match the design, compare providers in the guide to the best proxies for AI training and LLM data collection.
Using Proxidize for Authorized AI Data Collection
Proxidize supplies the network layer for teams running their own crawler or browser. It is not a dataset vendor, extraction API, licensing service, or data-cleaning platform. The AI and LLM data-collection use case explains that product fit at a higher level.
Proxidize Residential Proxies currently provide access to residential IPs across 195+ countries, with country, city, and ISP targeting, rotating or sticky sessions, and HTTP, HTTPS, and SOCKS5 support. That is the usual fit for geographically distributed public-web collection.
Proxidize Mobile Proxies are relevant when an authorized workflow specifically needs US mobile-carrier routing or a mobile-network view. Mobile bandwidth should not be treated as a universal upgrade for general corpus collection; residential or direct access is normally more cost-effective when carrier identity is not part of the requirement.
A sensible deployment policy is:
- Use a first-party feed, API, or bulk export when available.
- Test direct access for permitted public pages.
- Add residential routing only for a demonstrated location or access requirement.
- Keep a sticky session for dependent requests that must share state.
- Rotate between independent observations according to the documented session policy.
- Apply per-host rate limits regardless of how many proxy exits are available.
- Store the observed exit location, session identifier, source URL, timestamp, and collection outcome in the crawl record.
The proxy should improve acquisition reliability or geographic coverage. It should not become a reason to retain low-quality sources or ignore a site's access policy.
How Should You Measure an LLM Data Pipeline?
Track the pipeline from source discovery through model results.
Collection Metrics
- fetch success by source, status class, and region;
- valid-content extraction rate;
- bytes transferred per accepted document;
- freshness lag between publication and collection;
- retry and rendering rates;
- collection failures by reason.
Curation Metrics
- acceptance rate by source and filter;
- exact and near-duplicate removal rate;
- language-identification accuracy on reviewed samples;
- provenance completeness;
- rights-review coverage and unresolved-source count;
- personal-data or sensitive-data findings;
- contamination matches against protected evaluations.
Dataset Metrics
- documents and tokens by source, domain, language, format, time, and quality tier;
- source concentration and long-tail coverage;
- stale-document share;
- label agreement and adjudication rate;
- synthetic-to-human data ratio;
- deletion and opt-out processing time.
Model and RAG Metrics
- held-out task quality;
- subgroup and language performance;
- factuality and citation support;
- retrieval recall, precision, and answer-grounding rate;
- memorization and privacy tests;
- safety and policy outcomes;
- regressions outside the adapted domain;
- serving cost and latency.
For the acquisition layer, calculate:
Also measure cost per valid document. A source with a high fetch rate can still be expensive if extraction fails, most pages are duplicates, or rights review rejects the material.
Common LLM Training Data Mistakes
Treating Every Retrieved Document as Training Data
Raw downloads are candidates, not an approved corpus. Preserve them in a controlled intake layer until provenance, extraction, quality, privacy, and use-policy checks are complete.
Using One “Quality Score” as the Decision
Quality is multidimensional and task-specific. Retain the features and policy version behind a score, review samples from every threshold band, and test downstream impact.
Deduplicating After Splitting
Near-identical documents can land in training and validation, creating leakage. Cluster first, then assign whole clusters to splits.
Letting Synthetic Data Evaluate Itself
If the same model family creates training examples, rubrics, and test answers, correlated errors can make the system look stronger than it is. Use independent sources, human review, and protected evaluations.
Losing Source Lineage During Transformation
Every extracted passage, chunk, label, and synthetic derivative should retain a path to its source record and transformation version. A text-only file without lineage is difficult to audit or remove from downstream products.
Assuming a Proxy Solves Data Quality
A successful request proves that bytes were returned through a route. It does not prove that the page is the right source, the location context is correct, the parser extracted the intended material, or the content may be used for the planned purpose.
Build the Dataset Around Its Intended Job
The best LLM dataset is not the largest collection you can download. It is the smallest well-governed dataset that provides enough coverage and quality to meet a measured objective—and that can scale when the evidence justifies it.
Separate pretraining, fine-tuning, RAG, and evaluation data. Track provenance before transformations erase it. Review licenses, terms, privacy, and other obligations before admitting a source. Deduplicate before splitting, protect evaluations from contamination, and measure cost per accepted result.
When public-web collection is part of that design, use a proxy only for a real network or geographic requirement. Proxidize can provide the residential or mobile route; your data pipeline remains responsible for collection policy, extraction, governance, quality, and model evaluation.
Frequently asked questions
LLM training data is the material processed during pretraining, continued pretraining, fine-tuning, or alignment so that a language model's parameters change. It can include text, code, images, audio, structured records, demonstrations, preferences, and feedback.
Pretraining data teaches broad patterns across a large corpus. Fine-tuning data adapts an existing model to particular tasks, formats, domains, or behavior, often using smaller and more structured examples.
Not usually. A RAG corpus is retrieved at inference time and can be updated without changing model weights. Teams often group it under “LLM data,” but it needs retrieval, chunking, freshness, and access controls rather than a conventional training recipe.
There is no universal amount. It depends on the training stage, model, compute, objective, corpus quality, domain breadth, and desired performance. Use smaller scaling experiments and held-out evaluations to estimate the useful amount.
Common sources include first-party documents, licensed collections, open datasets, public web pages, code repositories, human annotations, feedback, and synthetic data. Each source needs its own provenance, quality, privacy, and rights review.
Public accessibility alone does not answer that question. Copyright, database rights, privacy, contracts, website terms, the collection method, intended use, and jurisdiction can matter. Organizations should obtain legal advice for their specific sources and use.
Deduplication reduces accidental over-weighting, wasted compute, memorization risk, and train-test leakage. Use exact and near-duplicate methods, then assign duplicate clusters to dataset splits together.
Synthetic data can supplement real examples, especially for rare or structured cases, but it should remain labeled and independently validated. It can reproduce generator errors, reduce diversity, or contaminate evaluations if used without controls.
A high-quality dataset is relevant to the intended task, traceable to its sources, appropriately reviewed for use, accurately extracted, representative, current enough, deduplicated, privacy-aware, and evaluated through downstream model behavior.
Not always. APIs, exports, licensed downloads, internal repositories, and distributed open datasets often work without proxies. A proxy is useful only when an authorized web collector needs a different exit IP, supported geographic location, or managed session behavior.