Skip to main content
LLM training data

Published Jan 28, 2026 · Updated Oct 6, 2026

What Is Bad Data? Types, Examples, and How to Fix It

Learn what bad data is, see common examples, measure six data-quality dimensions, and decide whether to clean, reject, or recollect unreliable records.

What Is Bad Data? Types, Examples, and How to Fix It

Bad data is information that is not reliable enough for its intended purpose. It may be inaccurate, incomplete, inconsistent, stale, invalid, duplicated, mislabeled, unrepresentative, or impossible to trace to a trustworthy source. The dangerous examples are not always obviously broken: a record can be correctly formatted and still describe the wrong product, country, date, or person.

Finding bad data therefore requires more than deleting blank rows. You need to define what a valid observation looks like, test each record against those rules, retain its source and collection context, and decide whether a failure can be corrected safely or must be rejected and collected again.

Quick Answer

Bad data is data that cannot reliably support its intended use. Common problems include wrong values, missing required fields, duplicates, inconsistent formats or units, invalid categories, stale observations, biased samples, and missing provenance. Identify it with purpose-specific validation rules and source checks. Clean a record only when the correct value can be established; otherwise quarantine, reject, or recollect it instead of guessing.

Key Takeaways

  • Data quality means fitness for purpose. The same record can be adequate for one task and inadequate for another.
  • Six useful quality dimensions are accuracy, completeness, consistency, timeliness, validity, and uniqueness. Relevance, provenance, rights, and representativeness may also matter.
  • Unstructured data is not inherently bad. HTML, text, images, audio, and documents can be high-quality inputs when their origin, meaning, and processing requirements are understood.
  • A successful request is not necessarily a valid observation. An HTTP 200 response can contain a challenge page, empty JavaScript shell, wrong locale, or unrelated content.
  • Not every error should be cleaned. Correct formatting when the intended value is known; reject or recollect data when the underlying fact is uncertain.
  • Proxies affect the network route, not data quality by themselves. The collector still needs parsing, validation, provenance, monitoring, and lawful access controls.

Methodology: We reviewed the previous article, six months of Google Search Console demand, and current first-party guidance from the UK Government, the UK Information Commissioner's Office, NIST, and NASA/JPL. The code below uses synthetic records and demonstrates validation logic; it is not a benchmark of a live dataset or proxy network.

What Is Bad Data?

Bad data is data whose quality is insufficient for a specific purpose. The UK Government Data Quality Framework describes data quality in terms of whether data is fit for its intended use. That definition is more useful than treating every imperfect record as equally bad.

Consider a product price recorded as $20:

  • It may be adequate for a rough market summary.
  • It is inadequate for an invoice if the actual price is $19.99.
  • It is misleading in a comparison if the source price was €20.
  • It is stale for a live monitor if it was collected six months ago.
  • It is invalid for a U.S. observation if the collector accidentally loaded the Canadian page.

The value looks plausible in every case. Its fitness depends on the required precision, currency, market, time, and source.

This is why a data-quality rule should name the purpose it protects. “Price must be present” is weaker than “price must be a positive decimal, use the expected currency, come from the requested product variant and country, and be no more than 24 hours old.”

Good Data vs. Bad Data

Good data does not need to be perfect in every imaginable way. It needs to satisfy the documented requirements of the decision, model, report, or workflow using it.

QuestionGood dataBad data
Is it accurate?Agrees with an authoritative source or verified observationContains an incorrect value or describes the wrong entity
Is it complete enough?Includes every field required for the taskOmits a critical value, page, population, or time period
Is it consistent?Uses compatible definitions, units, formats, and identifiersMixes currencies, units, time zones, labels, or entity definitions
Is it timely?Is current enough for the decisionIs stale, delayed, or missing its observation time
Is it valid?Conforms to the expected type, format, range, and categoryFails the schema or contains a value outside the permitted domain
Is it unique where required?Represents each intended entity or event onceRepeats a record because of retries, joins, or ingestion errors
Can it be traced?Retains source, collection time, method, and transformationsHas unclear origin or an undocumented processing history
Is it representative?Covers the population relevant to the conclusionSystematically omits or overweights important groups

A dataset can pass one row-level rule and fail at the dataset level. Every record in a survey might be valid, for example, while the sample excludes a segment the final conclusion claims to represent.

The Six Core Data-Quality Dimensions

The Government Data Quality Framework uses six practical dimensions: completeness, uniqueness, consistency, timeliness, validity, and accuracy. These dimensions overlap, but each exposes a different failure mode.

DimensionQuestion to askExample failureUseful check
AccuracyDoes the value reflect the real object or event?A product page is assigned the wrong priceCompare with an authoritative source or reviewed sample
CompletenessAre all required values and observations present?Currency is missing from a price recordMeasure null rates and coverage against required fields
ConsistencyDo values agree across records, systems, and definitions?One feed uses kilograms and another uses pounds without conversionCompare units, identifiers, formats, and cross-system values
TimelinessIs the data current and available when needed?An inventory result is three days old in an hourly monitorRecord event and collection times; enforce freshness thresholds
ValidityDoes the value conform to the expected format, range, and domain?A country field contains `United Stats`Apply schema, type, range, and allowed-value checks
UniquenessIs each intended entity or event represented once?A retry inserts the same order twiceUse stable keys, idempotent writes, and duplicate checks

Accuracy is often the hardest dimension to automate. A value can pass its schema and range checks while still being wrong. A syntactically valid postal code may belong to another address; a plausible price may come from the wrong product variant. Automated validation should therefore be paired with source comparison and representative human review.

The six dimensions are a strong baseline, not the whole governance model. Depending on the task, also evaluate:

  • Relevance: Does the data answer the question being asked?
  • Provenance: Can you identify its origin and transformation history?
  • Representativeness: Does it cover the relevant population without hidden gaps?
  • Label quality: Are categories and annotations applied consistently?
  • Rights and permitted use: Is the data lawfully and contractually usable for the intended purpose?

Technical accuracy and lawful use are separate questions. A dataset can be factually correct but unsuitable to use because it contains personal data without an appropriate basis, lacks necessary rights, or was collected contrary to applicable obligations. The ICO's accuracy guidance also explains the duty to take reasonable steps to keep personal data accurate and, where necessary, up to date.

Is Unstructured Data Bad Data?

No. Unstructured data describes a format, not a quality level.

Free text, HTML, PDFs, images, audio, and video do not fit neatly into relational rows and columns, but they can still be accurate, complete, timely, traceable, and useful. A signed PDF contract may be a more authoritative source than a tidy spreadsheet copied from it.

Unstructured data becomes a quality problem when, for example:

  • the source or creation date is missing;
  • a document cannot be associated with the correct entity;
  • extraction removes important context;
  • optical character recognition introduces errors;
  • required pages or sections are absent;
  • the content cannot be processed reliably for the intended task;
  • different formats are mixed without documenting how they were normalized.

The fix is not to label all unstructured inputs as bad. Define the metadata, extraction, review, and acceptance rules needed to make each format usable.

Common Types and Examples of Bad Data

TypeExampleWhy it matters
InaccurateA customer address is assigned to the wrong personDecisions concern the wrong entity
IncompleteA product observation has a price but no currencyThe value cannot be compared safely
InconsistentOne system stores meters while another assumes feetCalculations combine incompatible values
InvalidA timestamp contains an impossible monthThe value violates the expected domain
DuplicateA timed-out submission is retried and stored twiceCounts, revenue, or model frequency become inflated
StaleLast year's stock status is presented as currentA formerly correct fact produces a wrong current decision
MislabeledA support ticket is tagged as a billing issue when it concerns account accessModels and routing rules learn the wrong category
IrrelevantA dataset includes unrelated pages because the crawler followed a broad URL patternStorage and analysis are diluted by noise
UnrepresentativeA market analysis uses only one region but claims a national resultThe conclusion is generalized beyond the sample
UntraceableA row has no source URL, collection time, or transformation recordThe team cannot verify or reproduce the observation

“Dirty data” is often used for repairable problems such as inconsistent capitalization, spacing, date formats, and duplicate rows. “Bad data” is broader. It includes dirty data, but also incorrect facts, biased samples, stale observations, mislabeled examples, corrupted joins, and data whose source cannot be trusted.

What Causes Bad Data?

Bad data can enter at any stage of the lifecycle. Cleaning the final table may hide symptoms without correcting the process that keeps producing them.

Lifecycle stageCommon causeExample control
PlanningThe intended use and required fields are undefinedWrite a data contract and acceptance criteria before collection
CollectionWrong source, location, account state, device context, or sampling methodRecord collection context and verify it inside the same client
IngestionTruncation, encoding errors, schema drift, or incorrect field mappingValidate response size, encoding, schema, and required fields
TransformationFaulty joins, unit conversion, timezone conversion, or business logicTest transformations against known fixtures and reconcile totals
StorageDuplicate writes, lost updates, or missing version historyUse stable keys, idempotent operations, constraints, and audit logs
LabelingAmbiguous rubrics or inconsistent reviewersDefine examples, measure agreement, and adjudicate disagreements
AnalysisSelection bias, leakage, inappropriate aggregation, or unsupported assumptionsDocument sampling and preserve held-out data
MaintenanceNo owner, freshness rule, monitoring, or correction processAssign ownership and alert on quality trends

Human error is one cause, but “be more careful” is not a reliable control. Good systems make the correct action easier, reject impossible values early, expose uncertainty, and retain enough evidence to investigate a failure.

Bad Data in Web Scraping and Public-Web Collection

Web collection creates quality problems that can look like successful requests. Saving every HTTP 200 response as valid data is one of the most common mistakes.

A challenge page is stored as the target page

A website can return a CAPTCHA, access notice, sign-in page, or generic error template with a 200 status. If the collector checks only the status code, its parser may store blank fields or extract text from the wrong page.

Validate page identity as well as transport success. Check stable content markers, canonical URL, page type, expected fields, and redirect destination.

The collector observes the wrong market

A retailer may select currency, inventory, language, or offers using the IP location, delivery address, cookie, account, hostname, locale, or a combination of signals. A U.S. exit IP alone does not prove that the result is a U.S. observation.

Record both the requested context and evidence from the returned page. Validate currency, country selector, delivery location, language, and product variant where they affect the result.

JavaScript has not produced the required content

The initial HTML may contain a shell while the data arrives through a later request. A collector that parses too early can save an empty but valid-looking page. Use the underlying permitted data endpoint when appropriate, or wait for the specific browser state that proves the required content loaded.

A template change silently breaks the parser

Selectors can continue returning text after a page redesign while pointing to the wrong element. Monitor field null rates, value distributions, page fingerprints, and sampled screenshots or HTML. Treat a sudden shift as a possible collection failure before calling it a market trend.

Retries create duplicate observations

If a client times out after the server has processed a request, a blind retry can create a second record. Use a stable observation key and idempotent writes. Keep attempt logs separate from accepted business records.

The record loses its provenance

A price without its source URL, collection time, country, currency, product variant, and method is difficult to audit. Capture provenance at collection time; reconstructing it later is unreliable.

For the broader collection architecture, see How to Use Proxies for Web Scraping and the guide to data sourcing.

How to Identify Bad Data

Use layered validation. No single test can establish that every record is correct.

1. Define an acceptance contract

State what an accepted record must contain:

  • entity or observation identifier;
  • required fields and data types;
  • allowed categories, currencies, countries, and units;
  • valid numeric ranges and cross-field relationships;
  • maximum age;
  • uniqueness rule;
  • required source and provenance fields;
  • expected page or response type;
  • requested location, language, device, or account context;
  • treatment of missing optional values.

Write these rules before inspecting the output when possible. Otherwise, it is easy to adjust the definition until questionable data passes.

2. Validate the retrieval

Check status, redirects, content type, response size, encoding, and whether the response is the expected resource. A successful TCP connection or HTTP status does not validate the business result.

3. Validate the schema and fields

Check required values, types, formats, allowed categories, and ranges. Then add cross-field rules. A currency can be individually valid but inconsistent with the requested market.

4. Check uniqueness and entity identity

Choose keys that represent the actual entity or observation. A job ID should usually be scoped to its employer; a product ID may need a merchant and market; a price observation may need product, variant, location, and time.

5. Check freshness and sequence

Store event time and collection time separately when both exist. Reject impossible future dates and flag observations older than the workload permits. Verify sequence rules for event streams.

6. Compare distributions and coverage

Watch null rates, counts, category proportions, price ranges, geographic coverage, and other domain-specific distributions. A pipeline can pass row-level validation while losing half its normal coverage.

7. Review a representative sample

Compare sampled records with their original evidence. Include failures, uncommon categories, boundary values, and each important source—not only easy successes.

8. Quarantine failures

Do not mix questionable records into the accepted dataset. Store the failed record, reason codes, relevant evidence, and pipeline version in a controlled quarantine area. This makes correction and root-cause analysis possible without silently treating failure as truth.

Python Example: Validate Collected Price Records

The following standard-library example validates synthetic price observations. It checks record identity, duplicates, response and page type, required values, location, currency, positive price, timestamp format, and freshness.

It deliberately includes an HTTP 200 challenge page to show why status alone is insufficient.

python

The output is:

bash

These checks do not prove the accepted price is factually correct. The next layer would compare sampled records with page evidence, verify product and variant identity, and monitor distributions over time. The code demonstrates a gate, not a universal quality system.

Should You Clean, Reject, or Recollect the Data?

Cleaning is appropriate only when the intended value can be established without inventing information.

ProblemUsually appropriate actionWhy
Extra whitespace or a known date-format differenceNormalize and retain the original valueThe transformation is deterministic
Known unit difference with documented unitsConvert and record the conversionThe original meaning is available
Exact duplicate with a stable keyDeduplicate while preserving audit historyThe records represent the same intended observation
Missing optional fieldAccept with a flag if the use permits itThe record may still be fit for purpose
Missing critical valueReject or recollectGuessing creates an unsupported fact
Challenge, sign-in, or error pageReject and investigate before a bounded recollectionThe response is not the intended observation
Wrong country, currency, product, or account contextReject and recollect with corrected contextA valid-looking value describes the wrong observation
Stale recordRecollect or label it historicalIts suitability depends on the time requirement
Unclear or unverifiable sourceQuarantine pending reviewReliability and permitted use cannot be established
Coverage or representation gapCollect additional data and limit claimsRow cleaning cannot repair a missing population
Incorrect label with authoritative evidenceCorrect it and retain the change historyThe replacement can be supported and audited

Do not replace unknown values with averages or generated text unless the method is explicitly appropriate for the analysis and the imputation is marked. An estimated value is not an observed value.

A Real Example: The Mars Climate Orbiter

NASA's Mars Climate Orbiter is a well-known example of why consistency and interface validation matter. NASA reports that the ground software used English units while the onboard software expected metric units. The mismatch caused trajectory errors and contributed to the loss of the spacecraft in 1999.

This was not a blank-cell problem. Values existed, but two systems interpreted them under incompatible measurement conventions. The lesson for modern data pipelines is straightforward: units and semantic definitions belong in the contract, and interface tests should validate them before downstream decisions depend on the result.

How Bad Data Affects Analytics and AI

Bad data can produce confident outputs because software processes supplied values consistently even when those values are wrong.

In analytics, it can cause:

  • incorrect counts, totals, and rates;
  • false trends caused by collection failures;
  • customer or product records joined to the wrong entity;
  • misleading geographic comparisons;
  • decisions based on stale or incomplete coverage.

In AI systems, quality problems can appear in several places:

  • Training data: duplicates, incorrect labels, unrepresentative sources, and contamination can distort learning or evaluation.
  • RAG corpora: stale documents, weak chunk metadata, and retrieval gaps can produce unsupported answers.
  • Evaluation sets: leakage into training data or ambiguous rubrics can inflate reported performance.
  • Production inputs: malformed or adversarial documents can cause extraction and grounding failures.

The NIST AI Risk Management Framework measurement guidance recommends documenting the origin, quality, representativeness, and suitability of data used in AI systems. The exact controls depend on the system and risk, but provenance and representative testing should not begin after deployment.

For a fuller pipeline, see the guides to LLM training data, ETL pipelines, data parsing, and data normalization.

How to Prevent Bad Data

Prevention is more effective than repeatedly repairing final outputs.

  1. Name the purpose and owner. Document who uses the data, which decisions it supports, and who resolves quality failures.
  2. Define a data contract. Specify fields, types, semantics, units, identifiers, freshness, context, and acceptance thresholds.
  3. Validate at ingestion. Reject malformed or clearly wrong records before they mix with accepted data.
  4. Preserve raw evidence. Keep an appropriately protected, immutable source record or content hash so transformations can be audited.
  5. Capture provenance. Store source, time, method, requested context, parser version, and transformation history.
  6. Make writes idempotent. Stable keys and database constraints prevent retries from creating duplicate business records.
  7. Separate failure from absence. A failed collection must not be interpreted as a product removal, zero value, or unchanged page.
  8. Monitor quality rates. Track rejection, null, duplicate, stale, wrong-context, and source-coverage rates over time.
  9. Review representative samples. Include boundary cases and failures, not only accepted records.
  10. Fix the source of recurring errors. Update the form, parser, mapping, contract, or workflow instead of applying the same cleanup indefinitely.

The UK Government's Data Quality Framework guidance similarly emphasizes defining rules around user needs, measuring the current state, monitoring quality, and addressing issues as close to their source as possible.

Where Do Proxies Fit?

A proxy provides a network route and exit IP for the traffic configured to use it. It can be useful when an authorized collection workflow needs location-specific observations, route separation between workers, or another provider-supported network context.

A proxy does not:

  • decide which pages are permitted to collect;
  • guarantee that a website will return the intended content;
  • execute JavaScript unless a browser does that work;
  • identify a challenge page;
  • repair a selector or parser;
  • verify a price, label, or source;
  • deduplicate records;
  • establish rights to use collected material.

Start with the simplest route that works. When a proxy is required, validate the exit IP and country from the same HTTP client or browser context that performs the collection. Then validate the returned content separately. A successful IP check proves the route used for that check; it does not prove that a later page contains the correct product, locale, currency, or data.

Proxidize Residential Proxies can provide country, city, and ISP targeting for permitted global collection workflows. The scraper remains responsible for request pacing, browser behavior, parsing, validation, storage, and compliance. Teams should measure cost per accepted record—not just requests sent or bytes downloaded—because invalid pages and retries consume resources without producing usable data.

Data-Quality Metrics Worth Tracking

Choose metrics that correspond to the failure modes of the workflow:

MetricExample calculationWhat it reveals
Valid-record rateAccepted records / attempted observationsOverall usable output from the pipeline
Required-field completenessRecords with every required field / accepted recordsMissing critical data
Duplicate rateDuplicate records / ingested recordsRetry or identity problems
Freshness pass rateRecords within the age limit / checked recordsStale sources or delayed pipelines
Correct-context rateRecords matching requested country, currency, or variant / checked recordsLocation and session mistakes
Source coverageSuccessfully observed sources / planned sourcesSilent collection gaps
Parser failure rateRecords rejected for structure or selector errors / responsesTemplate drift and extraction problems
Recollection recovery rateValid records recovered / recollection attemptsWhether retry or routing changes solve the failure

Report the denominator and rejection rules with every rate. “99% accurate” has little meaning if the sample, evidence, and definition of accuracy are missing.

Conclusion

Bad data is data that is unfit for its intended use. Some problems are easy to see, such as a missing field or malformed date. Others are plausible values attached to the wrong entity, location, unit, time, or source.

The practical response is to define acceptance rules before scaling, validate retrieval and business meaning separately, retain provenance, measure failures, and avoid guessing at facts you cannot verify. Clean deterministic formatting problems. Quarantine uncertainty. Recollect observations that describe the wrong context. Then fix the process that produced the failure so it does not return in the next batch.

Frequently asked questions

A product price of 19.99 without a currency is a simple example. The number is correctly formatted but cannot be compared safely. Another example is a scraper receiving an HTTP 200 challenge page and storing it as if it were the intended product page.

Dirty data usually refers to repairable issues such as inconsistent formatting, extra whitespace, duplicate rows, or standardized spelling differences. Bad data is broader. It also includes incorrect facts, stale observations, biased samples, wrong labels, missing provenance, and values collected under the wrong context.

No. Unstructured data is a format category, not a quality judgment. Text, HTML, PDFs, images, audio, and video can be high-quality sources. Their usefulness depends on provenance, completeness, extraction accuracy, metadata, and the intended task.

Define required fields, types, ranges, categories, units, freshness, uniqueness, source, and collection context. Apply automated checks, monitor distributions and coverage, and compare a representative sample with authoritative evidence. Separate failed collection from a valid observation with an empty or zero value.

No. Formatting and known unit differences can often be corrected deterministically. Missing critical facts, wrong-context observations, uncertain sources, and unrepresentative coverage usually require review, rejection, or recollection. Do not invent a value merely to complete a row.

Incorrect labels, duplicates, stale documents, provenance gaps, unrepresentative samples, and evaluation leakage can distort training, retrieval, and reported performance. AI data should be versioned, traceable, reviewed for its intended use, and tested against representative production cases.

No. A proxy changes the route and exit IP used by configured traffic. It can support location-specific or distributed collection, but it does not validate the returned page, parse fields, remove duplicates, verify facts, or grant permission to collect data. Those controls belong in the crawler and data pipeline.

Ready to launch?

Proxies built for real operations.

For teams that depend on stability, not luck.