
Bad data is information that is not reliable enough for its intended purpose. It may be inaccurate, incomplete, inconsistent, stale, invalid, duplicated, mislabeled, unrepresentative, or impossible to trace to a trustworthy source. The dangerous examples are not always obviously broken: a record can be correctly formatted and still describe the wrong product, country, date, or person.
Finding bad data therefore requires more than deleting blank rows. You need to define what a valid observation looks like, test each record against those rules, retain its source and collection context, and decide whether a failure can be corrected safely or must be rejected and collected again.
Quick Answer
Bad data is data that cannot reliably support its intended use. Common problems include wrong values, missing required fields, duplicates, inconsistent formats or units, invalid categories, stale observations, biased samples, and missing provenance. Identify it with purpose-specific validation rules and source checks. Clean a record only when the correct value can be established; otherwise quarantine, reject, or recollect it instead of guessing.
Key Takeaways
- Data quality means fitness for purpose. The same record can be adequate for one task and inadequate for another.
- Six useful quality dimensions are accuracy, completeness, consistency, timeliness, validity, and uniqueness. Relevance, provenance, rights, and representativeness may also matter.
- Unstructured data is not inherently bad. HTML, text, images, audio, and documents can be high-quality inputs when their origin, meaning, and processing requirements are understood.
- A successful request is not necessarily a valid observation. An HTTP 200 response can contain a challenge page, empty JavaScript shell, wrong locale, or unrelated content.
- Not every error should be cleaned. Correct formatting when the intended value is known; reject or recollect data when the underlying fact is uncertain.
- Proxies affect the network route, not data quality by themselves. The collector still needs parsing, validation, provenance, monitoring, and lawful access controls.
Methodology: We reviewed the previous article, six months of Google Search Console demand, and current first-party guidance from the UK Government, the UK Information Commissioner's Office, NIST, and NASA/JPL. The code below uses synthetic records and demonstrates validation logic; it is not a benchmark of a live dataset or proxy network.
What Is Bad Data?
Bad data is data whose quality is insufficient for a specific purpose. The UK Government Data Quality Framework describes data quality in terms of whether data is fit for its intended use. That definition is more useful than treating every imperfect record as equally bad.
Consider a product price recorded as $20:
- It may be adequate for a rough market summary.
- It is inadequate for an invoice if the actual price is $19.99.
- It is misleading in a comparison if the source price was €20.
- It is stale for a live monitor if it was collected six months ago.
- It is invalid for a U.S. observation if the collector accidentally loaded the Canadian page.
The value looks plausible in every case. Its fitness depends on the required precision, currency, market, time, and source.
This is why a data-quality rule should name the purpose it protects. “Price must be present” is weaker than “price must be a positive decimal, use the expected currency, come from the requested product variant and country, and be no more than 24 hours old.”
Good Data vs. Bad Data
Good data does not need to be perfect in every imaginable way. It needs to satisfy the documented requirements of the decision, model, report, or workflow using it.
| Question | Good data | Bad data |
|---|---|---|
| Is it accurate? | Agrees with an authoritative source or verified observation | Contains an incorrect value or describes the wrong entity |
| Is it complete enough? | Includes every field required for the task | Omits a critical value, page, population, or time period |
| Is it consistent? | Uses compatible definitions, units, formats, and identifiers | Mixes currencies, units, time zones, labels, or entity definitions |
| Is it timely? | Is current enough for the decision | Is stale, delayed, or missing its observation time |
| Is it valid? | Conforms to the expected type, format, range, and category | Fails the schema or contains a value outside the permitted domain |
| Is it unique where required? | Represents each intended entity or event once | Repeats a record because of retries, joins, or ingestion errors |
| Can it be traced? | Retains source, collection time, method, and transformations | Has unclear origin or an undocumented processing history |
| Is it representative? | Covers the population relevant to the conclusion | Systematically omits or overweights important groups |
A dataset can pass one row-level rule and fail at the dataset level. Every record in a survey might be valid, for example, while the sample excludes a segment the final conclusion claims to represent.
The Six Core Data-Quality Dimensions
The Government Data Quality Framework uses six practical dimensions: completeness, uniqueness, consistency, timeliness, validity, and accuracy. These dimensions overlap, but each exposes a different failure mode.
| Dimension | Question to ask | Example failure | Useful check |
|---|---|---|---|
| Accuracy | Does the value reflect the real object or event? | A product page is assigned the wrong price | Compare with an authoritative source or reviewed sample |
| Completeness | Are all required values and observations present? | Currency is missing from a price record | Measure null rates and coverage against required fields |
| Consistency | Do values agree across records, systems, and definitions? | One feed uses kilograms and another uses pounds without conversion | Compare units, identifiers, formats, and cross-system values |
| Timeliness | Is the data current and available when needed? | An inventory result is three days old in an hourly monitor | Record event and collection times; enforce freshness thresholds |
| Validity | Does the value conform to the expected format, range, and domain? | A country field contains `United Stats` | Apply schema, type, range, and allowed-value checks |
| Uniqueness | Is each intended entity or event represented once? | A retry inserts the same order twice | Use stable keys, idempotent writes, and duplicate checks |
Accuracy is often the hardest dimension to automate. A value can pass its schema and range checks while still being wrong. A syntactically valid postal code may belong to another address; a plausible price may come from the wrong product variant. Automated validation should therefore be paired with source comparison and representative human review.
The six dimensions are a strong baseline, not the whole governance model. Depending on the task, also evaluate:
- Relevance: Does the data answer the question being asked?
- Provenance: Can you identify its origin and transformation history?
- Representativeness: Does it cover the relevant population without hidden gaps?
- Label quality: Are categories and annotations applied consistently?
- Rights and permitted use: Is the data lawfully and contractually usable for the intended purpose?
Technical accuracy and lawful use are separate questions. A dataset can be factually correct but unsuitable to use because it contains personal data without an appropriate basis, lacks necessary rights, or was collected contrary to applicable obligations. The ICO's accuracy guidance also explains the duty to take reasonable steps to keep personal data accurate and, where necessary, up to date.
Is Unstructured Data Bad Data?
No. Unstructured data describes a format, not a quality level.
Free text, HTML, PDFs, images, audio, and video do not fit neatly into relational rows and columns, but they can still be accurate, complete, timely, traceable, and useful. A signed PDF contract may be a more authoritative source than a tidy spreadsheet copied from it.
Unstructured data becomes a quality problem when, for example:
- the source or creation date is missing;
- a document cannot be associated with the correct entity;
- extraction removes important context;
- optical character recognition introduces errors;
- required pages or sections are absent;
- the content cannot be processed reliably for the intended task;
- different formats are mixed without documenting how they were normalized.
The fix is not to label all unstructured inputs as bad. Define the metadata, extraction, review, and acceptance rules needed to make each format usable.
Common Types and Examples of Bad Data
| Type | Example | Why it matters |
|---|---|---|
| Inaccurate | A customer address is assigned to the wrong person | Decisions concern the wrong entity |
| Incomplete | A product observation has a price but no currency | The value cannot be compared safely |
| Inconsistent | One system stores meters while another assumes feet | Calculations combine incompatible values |
| Invalid | A timestamp contains an impossible month | The value violates the expected domain |
| Duplicate | A timed-out submission is retried and stored twice | Counts, revenue, or model frequency become inflated |
| Stale | Last year's stock status is presented as current | A formerly correct fact produces a wrong current decision |
| Mislabeled | A support ticket is tagged as a billing issue when it concerns account access | Models and routing rules learn the wrong category |
| Irrelevant | A dataset includes unrelated pages because the crawler followed a broad URL pattern | Storage and analysis are diluted by noise |
| Unrepresentative | A market analysis uses only one region but claims a national result | The conclusion is generalized beyond the sample |
| Untraceable | A row has no source URL, collection time, or transformation record | The team cannot verify or reproduce the observation |
“Dirty data” is often used for repairable problems such as inconsistent capitalization, spacing, date formats, and duplicate rows. “Bad data” is broader. It includes dirty data, but also incorrect facts, biased samples, stale observations, mislabeled examples, corrupted joins, and data whose source cannot be trusted.
What Causes Bad Data?
Bad data can enter at any stage of the lifecycle. Cleaning the final table may hide symptoms without correcting the process that keeps producing them.
| Lifecycle stage | Common cause | Example control |
|---|---|---|
| Planning | The intended use and required fields are undefined | Write a data contract and acceptance criteria before collection |
| Collection | Wrong source, location, account state, device context, or sampling method | Record collection context and verify it inside the same client |
| Ingestion | Truncation, encoding errors, schema drift, or incorrect field mapping | Validate response size, encoding, schema, and required fields |
| Transformation | Faulty joins, unit conversion, timezone conversion, or business logic | Test transformations against known fixtures and reconcile totals |
| Storage | Duplicate writes, lost updates, or missing version history | Use stable keys, idempotent operations, constraints, and audit logs |
| Labeling | Ambiguous rubrics or inconsistent reviewers | Define examples, measure agreement, and adjudicate disagreements |
| Analysis | Selection bias, leakage, inappropriate aggregation, or unsupported assumptions | Document sampling and preserve held-out data |
| Maintenance | No owner, freshness rule, monitoring, or correction process | Assign ownership and alert on quality trends |
Human error is one cause, but “be more careful” is not a reliable control. Good systems make the correct action easier, reject impossible values early, expose uncertainty, and retain enough evidence to investigate a failure.
Bad Data in Web Scraping and Public-Web Collection
Web collection creates quality problems that can look like successful requests. Saving every HTTP 200 response as valid data is one of the most common mistakes.
A challenge page is stored as the target page
A website can return a CAPTCHA, access notice, sign-in page, or generic error template with a 200 status. If the collector checks only the status code, its parser may store blank fields or extract text from the wrong page.
Validate page identity as well as transport success. Check stable content markers, canonical URL, page type, expected fields, and redirect destination.
The collector observes the wrong market
A retailer may select currency, inventory, language, or offers using the IP location, delivery address, cookie, account, hostname, locale, or a combination of signals. A U.S. exit IP alone does not prove that the result is a U.S. observation.
Record both the requested context and evidence from the returned page. Validate currency, country selector, delivery location, language, and product variant where they affect the result.
JavaScript has not produced the required content
The initial HTML may contain a shell while the data arrives through a later request. A collector that parses too early can save an empty but valid-looking page. Use the underlying permitted data endpoint when appropriate, or wait for the specific browser state that proves the required content loaded.
A template change silently breaks the parser
Selectors can continue returning text after a page redesign while pointing to the wrong element. Monitor field null rates, value distributions, page fingerprints, and sampled screenshots or HTML. Treat a sudden shift as a possible collection failure before calling it a market trend.
Retries create duplicate observations
If a client times out after the server has processed a request, a blind retry can create a second record. Use a stable observation key and idempotent writes. Keep attempt logs separate from accepted business records.
The record loses its provenance
A price without its source URL, collection time, country, currency, product variant, and method is difficult to audit. Capture provenance at collection time; reconstructing it later is unreliable.
For the broader collection architecture, see How to Use Proxies for Web Scraping and the guide to data sourcing.
How to Identify Bad Data
Use layered validation. No single test can establish that every record is correct.
1. Define an acceptance contract
State what an accepted record must contain:
- entity or observation identifier;
- required fields and data types;
- allowed categories, currencies, countries, and units;
- valid numeric ranges and cross-field relationships;
- maximum age;
- uniqueness rule;
- required source and provenance fields;
- expected page or response type;
- requested location, language, device, or account context;
- treatment of missing optional values.
Write these rules before inspecting the output when possible. Otherwise, it is easy to adjust the definition until questionable data passes.
2. Validate the retrieval
Check status, redirects, content type, response size, encoding, and whether the response is the expected resource. A successful TCP connection or HTTP status does not validate the business result.
3. Validate the schema and fields
Check required values, types, formats, allowed categories, and ranges. Then add cross-field rules. A currency can be individually valid but inconsistent with the requested market.
4. Check uniqueness and entity identity
Choose keys that represent the actual entity or observation. A job ID should usually be scoped to its employer; a product ID may need a merchant and market; a price observation may need product, variant, location, and time.
5. Check freshness and sequence
Store event time and collection time separately when both exist. Reject impossible future dates and flag observations older than the workload permits. Verify sequence rules for event streams.
6. Compare distributions and coverage
Watch null rates, counts, category proportions, price ranges, geographic coverage, and other domain-specific distributions. A pipeline can pass row-level validation while losing half its normal coverage.
7. Review a representative sample
Compare sampled records with their original evidence. Include failures, uncommon categories, boundary values, and each important source—not only easy successes.
8. Quarantine failures
Do not mix questionable records into the accepted dataset. Store the failed record, reason codes, relevant evidence, and pipeline version in a controlled quarantine area. This makes correction and root-cause analysis possible without silently treating failure as truth.
Python Example: Validate Collected Price Records
The following standard-library example validates synthetic price observations. It checks record identity, duplicates, response and page type, required values, location, currency, positive price, timestamp format, and freshness.
It deliberately includes an HTTP 200 challenge page to show why status alone is insufficient.
The output is:
These checks do not prove the accepted price is factually correct. The next layer would compare sampled records with page evidence, verify product and variant identity, and monitor distributions over time. The code demonstrates a gate, not a universal quality system.
Should You Clean, Reject, or Recollect the Data?
Cleaning is appropriate only when the intended value can be established without inventing information.
| Problem | Usually appropriate action | Why |
|---|---|---|
| Extra whitespace or a known date-format difference | Normalize and retain the original value | The transformation is deterministic |
| Known unit difference with documented units | Convert and record the conversion | The original meaning is available |
| Exact duplicate with a stable key | Deduplicate while preserving audit history | The records represent the same intended observation |
| Missing optional field | Accept with a flag if the use permits it | The record may still be fit for purpose |
| Missing critical value | Reject or recollect | Guessing creates an unsupported fact |
| Challenge, sign-in, or error page | Reject and investigate before a bounded recollection | The response is not the intended observation |
| Wrong country, currency, product, or account context | Reject and recollect with corrected context | A valid-looking value describes the wrong observation |
| Stale record | Recollect or label it historical | Its suitability depends on the time requirement |
| Unclear or unverifiable source | Quarantine pending review | Reliability and permitted use cannot be established |
| Coverage or representation gap | Collect additional data and limit claims | Row cleaning cannot repair a missing population |
| Incorrect label with authoritative evidence | Correct it and retain the change history | The replacement can be supported and audited |
Do not replace unknown values with averages or generated text unless the method is explicitly appropriate for the analysis and the imputation is marked. An estimated value is not an observed value.
A Real Example: The Mars Climate Orbiter
NASA's Mars Climate Orbiter is a well-known example of why consistency and interface validation matter. NASA reports that the ground software used English units while the onboard software expected metric units. The mismatch caused trajectory errors and contributed to the loss of the spacecraft in 1999.
This was not a blank-cell problem. Values existed, but two systems interpreted them under incompatible measurement conventions. The lesson for modern data pipelines is straightforward: units and semantic definitions belong in the contract, and interface tests should validate them before downstream decisions depend on the result.
How Bad Data Affects Analytics and AI
Bad data can produce confident outputs because software processes supplied values consistently even when those values are wrong.
In analytics, it can cause:
- incorrect counts, totals, and rates;
- false trends caused by collection failures;
- customer or product records joined to the wrong entity;
- misleading geographic comparisons;
- decisions based on stale or incomplete coverage.
In AI systems, quality problems can appear in several places:
- Training data: duplicates, incorrect labels, unrepresentative sources, and contamination can distort learning or evaluation.
- RAG corpora: stale documents, weak chunk metadata, and retrieval gaps can produce unsupported answers.
- Evaluation sets: leakage into training data or ambiguous rubrics can inflate reported performance.
- Production inputs: malformed or adversarial documents can cause extraction and grounding failures.
The NIST AI Risk Management Framework measurement guidance recommends documenting the origin, quality, representativeness, and suitability of data used in AI systems. The exact controls depend on the system and risk, but provenance and representative testing should not begin after deployment.
For a fuller pipeline, see the guides to LLM training data, ETL pipelines, data parsing, and data normalization.
How to Prevent Bad Data
Prevention is more effective than repeatedly repairing final outputs.
- Name the purpose and owner. Document who uses the data, which decisions it supports, and who resolves quality failures.
- Define a data contract. Specify fields, types, semantics, units, identifiers, freshness, context, and acceptance thresholds.
- Validate at ingestion. Reject malformed or clearly wrong records before they mix with accepted data.
- Preserve raw evidence. Keep an appropriately protected, immutable source record or content hash so transformations can be audited.
- Capture provenance. Store source, time, method, requested context, parser version, and transformation history.
- Make writes idempotent. Stable keys and database constraints prevent retries from creating duplicate business records.
- Separate failure from absence. A failed collection must not be interpreted as a product removal, zero value, or unchanged page.
- Monitor quality rates. Track rejection, null, duplicate, stale, wrong-context, and source-coverage rates over time.
- Review representative samples. Include boundary cases and failures, not only accepted records.
- Fix the source of recurring errors. Update the form, parser, mapping, contract, or workflow instead of applying the same cleanup indefinitely.
The UK Government's Data Quality Framework guidance similarly emphasizes defining rules around user needs, measuring the current state, monitoring quality, and addressing issues as close to their source as possible.
Where Do Proxies Fit?
A proxy provides a network route and exit IP for the traffic configured to use it. It can be useful when an authorized collection workflow needs location-specific observations, route separation between workers, or another provider-supported network context.
A proxy does not:
- decide which pages are permitted to collect;
- guarantee that a website will return the intended content;
- execute JavaScript unless a browser does that work;
- identify a challenge page;
- repair a selector or parser;
- verify a price, label, or source;
- deduplicate records;
- establish rights to use collected material.
Start with the simplest route that works. When a proxy is required, validate the exit IP and country from the same HTTP client or browser context that performs the collection. Then validate the returned content separately. A successful IP check proves the route used for that check; it does not prove that a later page contains the correct product, locale, currency, or data.
Proxidize Residential Proxies can provide country, city, and ISP targeting for permitted global collection workflows. The scraper remains responsible for request pacing, browser behavior, parsing, validation, storage, and compliance. Teams should measure cost per accepted record—not just requests sent or bytes downloaded—because invalid pages and retries consume resources without producing usable data.
Data-Quality Metrics Worth Tracking
Choose metrics that correspond to the failure modes of the workflow:
| Metric | Example calculation | What it reveals |
|---|---|---|
| Valid-record rate | Accepted records / attempted observations | Overall usable output from the pipeline |
| Required-field completeness | Records with every required field / accepted records | Missing critical data |
| Duplicate rate | Duplicate records / ingested records | Retry or identity problems |
| Freshness pass rate | Records within the age limit / checked records | Stale sources or delayed pipelines |
| Correct-context rate | Records matching requested country, currency, or variant / checked records | Location and session mistakes |
| Source coverage | Successfully observed sources / planned sources | Silent collection gaps |
| Parser failure rate | Records rejected for structure or selector errors / responses | Template drift and extraction problems |
| Recollection recovery rate | Valid records recovered / recollection attempts | Whether retry or routing changes solve the failure |
Report the denominator and rejection rules with every rate. “99% accurate” has little meaning if the sample, evidence, and definition of accuracy are missing.
Conclusion
Bad data is data that is unfit for its intended use. Some problems are easy to see, such as a missing field or malformed date. Others are plausible values attached to the wrong entity, location, unit, time, or source.
The practical response is to define acceptance rules before scaling, validate retrieval and business meaning separately, retain provenance, measure failures, and avoid guessing at facts you cannot verify. Clean deterministic formatting problems. Quarantine uncertainty. Recollect observations that describe the wrong context. Then fix the process that produced the failure so it does not return in the next batch.
Frequently asked questions
A product price of 19.99 without a currency is a simple example. The number is correctly formatted but cannot be compared safely. Another example is a scraper receiving an HTTP 200 challenge page and storing it as if it were the intended product page.
Dirty data usually refers to repairable issues such as inconsistent formatting, extra whitespace, duplicate rows, or standardized spelling differences. Bad data is broader. It also includes incorrect facts, stale observations, biased samples, wrong labels, missing provenance, and values collected under the wrong context.
No. Unstructured data is a format category, not a quality judgment. Text, HTML, PDFs, images, audio, and video can be high-quality sources. Their usefulness depends on provenance, completeness, extraction accuracy, metadata, and the intended task.
Define required fields, types, ranges, categories, units, freshness, uniqueness, source, and collection context. Apply automated checks, monitor distributions and coverage, and compare a representative sample with authoritative evidence. Separate failed collection from a valid observation with an empty or zero value.
No. Formatting and known unit differences can often be corrected deterministically. Missing critical facts, wrong-context observations, uncertain sources, and unrepresentative coverage usually require review, rejection, or recollection. Do not invent a value merely to complete a row.
Incorrect labels, duplicates, stale documents, provenance gaps, unrepresentative samples, and evaluation leakage can distort training, retrieval, and reported performance. AI data should be versioned, traceable, reviewed for its intended use, and tested against representative production cases.
No. A proxy changes the route and exit IP used by configured traffic. It can support location-specific or distributed collection, but it does not validate the returned page, parse fields, remove duplicates, verify facts, or grant permission to collect data. Those controls belong in the crawler and data pipeline.