0% found this document useful (0 votes)

8 views

Data_Preprocessing-1-19

Chapter 3 discusses data preprocessing, emphasizing the importance of data quality and the major tasks involved, including data cleaning, integration, reduction, and transformation. It outlines the challenges of handling missing, noisy, and inconsistent data, as well as techniques for data cleaning and integration. The chapter also covers methods for evaluating data quality and resolving conflicts during data integration.

Uploaded by

Mahim Jain Anwa

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

8 views

Data_Preprocessing-1-19

Uploaded by

Mahim Jain Anwa

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

You are on page 1/ 19

Chapter 3: Data Preprocessing

◼ Data Preprocessing: An Overview

◼ Data Quality

◼ Major Tasks in Data Preprocessing

◼ Data Cleaning

◼ Data Integration

◼ Data Reduction

◼ Data Transformation and Data Discretization

◼ Summary
1
Data Quality: Why Preprocess the Data?

◼ Measures for data quality: A multidimensional view

◼ Accuracy: correct or wrong, accurate or not
◼ Completeness: not recorded, unavailable, …
◼ Consistency: some modified but some not, dangling, …
◼ Timeliness: timely update?
◼ Believability: how trustable the data are correct?
◼ Interpretability: how easily the data can be
understood?

2
Major Tasks in Data Preprocessing
◼ Data cleaning
◼ Fill in missing values, smooth noisy data, identify or remove
outliers, and resolve inconsistencies
◼ Data integration
◼ Integration of multiple databases, data cubes, or files
◼ Data reduction
◼ Dimensionality reduction
◼ Numerosity reduction
◼ Data compression
◼ Data transformation and data discretization
◼ Normalization
◼ Concept hierarchy generation

3
Chapter 3: Data Preprocessing

◼ Data Preprocessing: An Overview

◼ Data Quality

◼ Major Tasks in Data Preprocessing

◼ Data Cleaning

◼ Data Integration

◼ Data Reduction

◼ Data Transformation and Data Discretization

◼ Summary
4
Data Cleaning
◼ Data in the Real World Is Dirty: Lots of potentially incorrect data,
e.g., instrument faulty, human or computer error, transmission error
◼ incomplete: lacking attribute values, lacking certain attributes of
interest, or containing only aggregate data
◼ e.g., Occupation=“ ” (missing data)
◼ noisy: containing noise, errors, or outliers
◼ e.g., Salary=“−10” (an error)
◼ inconsistent: containing discrepancies in codes or names, e.g.,
◼ Age=“42”, Birthday=“03/07/2010”
◼ Was rating “1, 2, 3”, now rating “A, B, C”
◼ discrepancy between duplicate records
◼ Intentional (e.g., disguised missing data)
◼ Jan. 1 as everyone’s birthday?
5
Incomplete (Missing) Data

◼ Data is not always available

◼ E.g., many tuples have no recorded value for several
attributes, such as customer income in sales data
◼ Missing data may be due to
◼ equipment malfunction
◼ inconsistent with other recorded data and thus deleted
◼ data not entered due to misunderstanding
◼ certain data may not be considered important at the
time of entry
◼ not register history or changes of the data
◼ Missing data may need to be inferred
6
How to Handle Missing Data?
◼ Ignore the tuple: usually done when class label is missing
(when doing classification)—not effective when the % of
missing values per attribute varies considerably
◼ Fill in the missing value manually: tedious + infeasible?
◼ Fill in it automatically with
◼ a global constant : e.g., “unknown”, a new class?!
◼ the attribute mean
◼ the attribute mean for all samples belonging to the
same class: smarter
◼ the most probable value: inference-based such as
Bayesian formula or decision tree
7
Noisy Data
◼ Noise: random error or variance in a measured variable
◼ Incorrect attribute values may be due to
◼ faulty data collection instruments

◼ data entry problems

◼ data transmission problems

◼ technology limitation

◼ inconsistency in naming convention

◼ Other data problems which require data cleaning

◼ duplicate records

◼ incomplete data

◼ inconsistent data

8
How to Handle Noisy Data?

◼ Binning
◼ first sort data and partition into (equal-frequency) bins

◼ then one can smooth by bin means, smooth by bin

median, smooth by bin boundaries, etc.

◼ Regression
◼ smooth by fitting the data into regression functions

◼ Clustering
◼ detect and remove outliers

◼ Combined computer and human inspection

◼ detect suspicious values and check by human (e.g.,

deal with possible outliers)

9
Data Cleaning as a Process
◼ Data discrepancy detection
◼ Use metadata (e.g., domain, range, dependency, distribution)

◼ Check field overloading

◼ Check uniqueness rule, consecutive rule and null rule

◼ Use commercial tools

◼ Data scrubbing: use simple domain knowledge (e.g., postal

code, spell-check) to detect errors and make corrections

◼ Data auditing: by analyzing data to discover rules and

relationship to detect violators (e.g., correlation and clustering

to find outliers)
◼ Data migration and integration
◼ Data migration tools: allow transformations to be specified

◼ ETL (Extraction/Transformation/Loading) tools: allow users to

specify transformations through a graphical user interface
◼ Integration of the two processes
◼ Iterative and interactive (e.g., Potter’s Wheels)

10
Chapter 3: Data Preprocessing

◼ Data Preprocessing: An Overview

◼ Data Quality

◼ Major Tasks in Data Preprocessing

◼ Data Cleaning

◼ Data Integration

◼ Data Reduction

◼ Data Transformation and Data Discretization

◼ Summary
11
Data Integration
◼ Data integration:
◼ Combines data from multiple sources into a coherent store
◼ Schema integration: e.g., A.cust-id  B.cust-#
◼ Integrate metadata from different sources
◼ Entity identification problem:
◼ Identify real world entities from multiple data sources, e.g., Bill
Clinton = William Clinton
◼ Detecting and resolving data value conflicts
◼ For the same real world entity, attribute values from different
sources are different
◼ Possible reasons: different representations, different scales, e.g.,
metric vs. British units
12
Handling Redundancy in Data Integration

◼ Redundant data occur often when integration of multiple

databases
◼ Object identification: The same attribute or object
may have different names in different databases
◼ Derivable data: One attribute may be a “derived”
attribute in another table, e.g., annual revenue
◼ Redundant attributes may be able to be detected by
correlation analysis and covariance analysis
◼ Careful integration of the data from multiple sources may
help reduce/avoid redundancies and inconsistencies and
improve mining speed and quality
13
Correlation Analysis (Numeric Data)

◼ Correlation coefficient (also called Pearson’s product

moment coefficient)

i=1 (ai − A)(bi − B) 

n n
(ai bi ) − n AB
rA, B = = i =1
(n − 1) A B (n − 1) A B

where n is the number of tuples, A and B are the respective

means of A and B, σA and σB are the respective standard deviation
of A and B, and Σ(aibi) is the sum of the AB cross-product.
◼ If rA,B > 0, A and B are positively correlated (A’s values
increase as B’s).
◼ rA,B = 0: independent; rAB < 0: negatively correlated

14
2/12/2025 Data Mining: Concepts and Techniques 15
Visually Evaluating Correlation

Scatter plots
showing the
similarity from
–1 to 1.

16
Correlation (viewed as linear relationship)
◼ Correlation measures the linear relationship
between objects
◼ To compute correlation, we standardize data
objects, A and B, and then take their dot product

a 'k = (ak − mean( A)) / std ( A)

b'k = (bk − mean( B )) / std ( B)

correlation( A, B) = A'• B '

17
Covariance (Numeric Data)
◼ Covariance is similar to correlation

Correlation coefficient:

where n is the number of tuples, A and B are the respective mean or

expected values of A and B, σA and σB are the respective standard
deviation of A and B.
◼ Positive covariance: If CovA,B > 0, then A and B both tend to be larger
than their expected values.
◼ Negative covariance: If CovA,B < 0 then if A is larger than its expected
value, B is likely to be smaller than its expected value.
◼ Independence: CovA,B = 0 but the converse is not true:
◼ Some pairs of random variables may have a covariance of 0 but are not
independent. Only under some additional assumptions (e.g., the data follow
multivariate normal distributions) does a covariance of 0 imply independence18
Co-Variance: An Example

◼ It can be simplified in computation as

◼ Suppose two stocks A and B have the following values in one week:
(2, 5), (3, 8), (5, 10), (4, 11), (6, 14).

◼ Question: If the stocks are affected by the same industry trends, will
their prices rise or fall together?

◼ E(A) = (2 + 3 + 5 + 4 + 6)/ 5 = 20/5 = 4

◼ E(B) = (5 + 8 + 10 + 11 + 14) /5 = 48/5 = 9.6

◼ Cov(A,B) = (2×5+3×8+5×10+4×11+6×14)/5 − 4 × 9.6 = 4

◼ Thus, A and B rise together since Cov(A, B) > 0.

Multivariate Statistical Inference and Applications
50% (10)
Multivariate Statistical Inference and Applications
634 pages
Chapter 3 - Tagged
No ratings yet
Chapter 3 - Tagged
63 pages
Wk6 Preprocessing
No ratings yet
Wk6 Preprocessing
64 pages
DM_merged
No ratings yet
DM_merged
169 pages
Lec7
No ratings yet
Lec7
45 pages
Chapter 3: Data Preprocessing
No ratings yet
Chapter 3: Data Preprocessing
63 pages
03Preprocessing_20160222
No ratings yet
03Preprocessing_20160222
65 pages
Chapter 3: Data Preprocessing
No ratings yet
Chapter 3: Data Preprocessing
56 pages
03 Preprocessing
No ratings yet
03 Preprocessing
63 pages
Chapter 3
No ratings yet
Chapter 3
63 pages
Concepts and Techniques: - Chapter 3
No ratings yet
Concepts and Techniques: - Chapter 3
63 pages
03 Preprocessing
No ratings yet
03 Preprocessing
63 pages
Data Preprocessing
No ratings yet
Data Preprocessing
63 pages
IT446 Wk03.2 HanKamberPei 03preprocessing PDF
No ratings yet
IT446 Wk03.2 HanKamberPei 03preprocessing PDF
64 pages
Concepts and Techniques: - Chapter 3
No ratings yet
Concepts and Techniques: - Chapter 3
63 pages
Mining
No ratings yet
Mining
63 pages
Concepts and Techniques: - Chapter 3
No ratings yet
Concepts and Techniques: - Chapter 3
64 pages
03 Preprocessing
No ratings yet
03 Preprocessing
54 pages
Data Preprocessing
No ratings yet
Data Preprocessing
77 pages
03Preprocessing
No ratings yet
03Preprocessing
65 pages
Module 2
No ratings yet
Module 2
62 pages
Concepts and Techniques: Data Mining
No ratings yet
Concepts and Techniques: Data Mining
66 pages
Data Mining: Dosen: Dr. Vitri Tundjungsari
No ratings yet
Data Mining: Dosen: Dr. Vitri Tundjungsari
64 pages
Data Pre Processing
No ratings yet
Data Pre Processing
63 pages
Module 5 03preprocessing
No ratings yet
Module 5 03preprocessing
63 pages
03 Preprocessing
No ratings yet
03 Preprocessing
64 pages
Lecture 2.3.1-2.3.3
No ratings yet
Lecture 2.3.1-2.3.3
67 pages
Concepts and Techniques: Data Mining
No ratings yet
Concepts and Techniques: Data Mining
61 pages
_03Preprocessing
No ratings yet
_03Preprocessing
60 pages
Chapter 3
No ratings yet
Chapter 3
56 pages
03 Pre Processing
No ratings yet
03 Pre Processing
63 pages
Data Mining and Knowledge Discovery
No ratings yet
Data Mining and Knowledge Discovery
65 pages
Preprocessing Techniques
No ratings yet
Preprocessing Techniques
63 pages
03 Pre Processing
No ratings yet
03 Pre Processing
89 pages
Lecture 3
No ratings yet
Lecture 3
47 pages
Chapter 3: Data Preprocessing
No ratings yet
Chapter 3: Data Preprocessing
62 pages
Concepts and Techniques: Data Mining
No ratings yet
Concepts and Techniques: Data Mining
52 pages
Unit 2 Data Preprocessing
No ratings yet
Unit 2 Data Preprocessing
40 pages
data mining 3
No ratings yet
data mining 3
57 pages
03Preprocessing
No ratings yet
03Preprocessing
38 pages
Slide 05 Chapter3 Data Preprocessing
No ratings yet
Slide 05 Chapter3 Data Preprocessing
58 pages
Chapter 3: Data Preprocessing
100% (1)
Chapter 3: Data Preprocessing
41 pages
Concepts and Techniques: Data Mining
No ratings yet
Concepts and Techniques: Data Mining
54 pages
Data Preprocessing (Sagar)
No ratings yet
Data Preprocessing (Sagar)
31 pages
Unit2 Part2
No ratings yet
Unit2 Part2
67 pages
Chapter 3: Data Preprocessing
No ratings yet
Chapter 3: Data Preprocessing
30 pages
PPT1
No ratings yet
PPT1
93 pages
2020 Preprocessing
No ratings yet
2020 Preprocessing
63 pages
Data Preprocessing (DWDM MOD 2)
No ratings yet
Data Preprocessing (DWDM MOD 2)
62 pages
Data Mining Requires Collecting Great Amount of Data (Available in Data Warehouses or Databases) To Achieve The Intended Objective
No ratings yet
Data Mining Requires Collecting Great Amount of Data (Available in Data Warehouses or Databases) To Achieve The Intended Objective
37 pages
Concepts and Techniques: Data Mining
No ratings yet
Concepts and Techniques: Data Mining
50 pages
Unit 3
No ratings yet
Unit 3
164 pages
03 Preprocessing
No ratings yet
03 Preprocessing
59 pages
Concepts and Techniques: - Chapter 3
No ratings yet
Concepts and Techniques: - Chapter 3
55 pages
03preprocessing DMDW
No ratings yet
03preprocessing DMDW
81 pages
TTDS Lecture 2
No ratings yet
TTDS Lecture 2
40 pages
TTDS Lecture 2
No ratings yet
TTDS Lecture 2
40 pages
HIT391-week 3-New
No ratings yet
HIT391-week 3-New
43 pages
DP
No ratings yet
DP
44 pages
IT Specialist: Data Analytics Certification Prep - 500 Exam Questions and Explanations
From Everand
IT Specialist: Data Analytics Certification Prep - 500 Exam Questions and Explanations
Steve Brown
No ratings yet
Illuminating Data: A hands on guide to data visualization in R
From Everand
Illuminating Data: A hands on guide to data visualization in R
Eman Ahmad
No ratings yet
DataScience_1
No ratings yet
DataScience_1
22 pages
Assignment 1
No ratings yet
Assignment 1
2 pages
Assignment (1)
No ratings yet
Assignment (1)
3 pages
adverse_reactions.csv
No ratings yet
adverse_reactions.csv
1 page
Quantum Theory
No ratings yet
Quantum Theory
65 pages
assignment_8_report
No ratings yet
assignment_8_report
13 pages
4 EL EMWaves Vaccum BC
No ratings yet
4 EL EMWaves Vaccum BC
46 pages
1 EL Vectors
No ratings yet
1 EL Vectors
59 pages
2 EL Div Curl EB
No ratings yet
2 EL Div Curl EB
50 pages
3 EL Maxwells Eqns
No ratings yet
3 EL Maxwells Eqns
23 pages
Ch_6_7_Homework_assignment
No ratings yet
Ch_6_7_Homework_assignment
5 pages
(Bethea, Robert M) - Statistical Methods For Engineers and Scientists, Third Edition-Routledge (2018)
75% (4)
(Bethea, Robert M) - Statistical Methods For Engineers and Scientists, Third Edition-Routledge (2018)
681 pages
CP / CPK Calculation Sheet: Specification
No ratings yet
CP / CPK Calculation Sheet: Specification
6 pages
465 - C Applied Mathematics
No ratings yet
465 - C Applied Mathematics
23 pages
Spearman Rho
0% (1)
Spearman Rho
26 pages
Session 22,23 - Interval Estimates
No ratings yet
Session 22,23 - Interval Estimates
68 pages
Problem Set 6. Statistics and Probability
No ratings yet
Problem Set 6. Statistics and Probability
3 pages
Robust Detection of Multiple Outliers in A Multivariate Data Set
No ratings yet
Robust Detection of Multiple Outliers in A Multivariate Data Set
30 pages
Practice Questions - Statistical Process Control
No ratings yet
Practice Questions - Statistical Process Control
3 pages
Evaluating Statistical Claims (Level 1) Answer Key
No ratings yet
Evaluating Statistical Claims (Level 1) Answer Key
2 pages
Exploring Business FINAL EXAM
No ratings yet
Exploring Business FINAL EXAM
2 pages
Statistics for Criminology and Criminal Justice 4th Edition Bachman Test Bank - Full Book Is Now Available For Download
100% (2)
Statistics for Criminology and Criminal Justice 4th Edition Bachman Test Bank - Full Book Is Now Available For Download
43 pages
Curve Fitting Linear 1
No ratings yet
Curve Fitting Linear 1
43 pages
Statistics Exercise
No ratings yet
Statistics Exercise
14 pages
Random Motors Project
No ratings yet
Random Motors Project
10 pages
Probability and Statistical Course.: Instructor: DR - Ing. (C) Sergio A. Abreo C
No ratings yet
Probability and Statistical Course.: Instructor: DR - Ing. (C) Sergio A. Abreo C
25 pages
Causal Research Design Summary Chapter 7
No ratings yet
Causal Research Design Summary Chapter 7
7 pages
Binomial and Geometric Distributions
No ratings yet
Binomial and Geometric Distributions
12 pages
Mathematics P2
No ratings yet
Mathematics P2
24 pages
Meta AnalysisofPrevalenceStudiesusingR August2021
No ratings yet
Meta AnalysisofPrevalenceStudiesusingR August2021
6 pages
IITM B.Sc. Qualifier Exam Revision
No ratings yet
IITM B.Sc. Qualifier Exam Revision
3 pages
Notes 12
No ratings yet
Notes 12
41 pages
المعرفة السوقية
No ratings yet
المعرفة السوقية
40 pages
Stats
No ratings yet
Stats
19 pages
Statistics and Probability
No ratings yet
Statistics and Probability
196 pages
Jurnal Mengenai Regresi Linear
No ratings yet
Jurnal Mengenai Regresi Linear
19 pages
11 - ch10 - p364-424.qxd 9/7/11 12:47 PM Page 373: MINITAB Output For
No ratings yet
11 - ch10 - p364-424.qxd 9/7/11 12:47 PM Page 373: MINITAB Output For
4 pages
SQC
100% (2)
SQC
44 pages
INFERENTIAL STATISTICS (Project)
No ratings yet
INFERENTIAL STATISTICS (Project)
17 pages

Data_Preprocessing-1-19

Uploaded by

Data_Preprocessing-1-19

Uploaded by

Chapter 3: Data Preprocessing

◼ Data Preprocessing: An Overview

◼ Major Tasks in Data Preprocessing

◼ Data Transformation and Data Discretization

◼ Measures for data quality: A multidimensional view

◼ Data Preprocessing: An Overview

◼ Major Tasks in Data Preprocessing

◼ Data Transformation and Data Discretization

◼ Data is not always available

◼ data entry problems

◼ data transmission problems

◼ inconsistency in naming convention

◼ Other data problems which require data cleaning

◼ then one can smooth by bin means, smooth by bin

median, smooth by bin boundaries, etc.

◼ Combined computer and human inspection

deal with possible outliers)

◼ Check field overloading

◼ Check uniqueness rule, consecutive rule and null rule

◼ Use commercial tools

◼ Data scrubbing: use simple domain knowledge (e.g., postal

code, spell-check) to detect errors and make corrections

relationship to detect violators (e.g., correlation and clustering

◼ ETL (Extraction/Transformation/Loading) tools: allow users to

◼ Data Preprocessing: An Overview

◼ Major Tasks in Data Preprocessing

◼ Data Transformation and Data Discretization

◼ Redundant data occur often when integration of multiple

◼ Correlation coefficient (also called Pearson’s product

i=1 (ai − A)(bi − B) 

where n is the number of tuples, A and B are the respective

a 'k = (ak − mean( A)) / std ( A)

b'k = (bk − mean( B )) / std ( B)

correlation( A, B) = A'• B '

where n is the number of tuples, A and B are the respective mean or

◼ It can be simplified in computation as

◼ E(A) = (2 + 3 + 5 + 4 + 6)/ 5 = 20/5 = 4

◼ E(B) = (5 + 8 + 10 + 11 + 14) /5 = 48/5 = 9.6

◼ Cov(A,B) = (2×5+3×8+5×10+4×11+6×14)/5 − 4 × 9.6 = 4

◼ Thus, A and B rise together since Cov(A, B) > 0.

You might also like