Intro Data Mining
Intro Data Mining
Data Mining:
Concepts and Techniques
— Chapter 1 —
2
August 19, 2020 Data Mining: Concepts and Techniques 3
August 19, 2020 Data Mining: Concepts and Techniques 4
Chapter 1. Introduction
Why Data Mining?
What Is Data Mining?
A Multi-Dimensional View of Data Mining
What Kind of Data Can Be Mined?
What Kinds of Patterns Can Be Mined?
What Technology Are Used?
What Kind of Applications Are Targeted?
5
Why Data Mining?
6
Evolution of Database Technology
1960s:
Data collection, database creation, IMS and network DBMS
1970s:
Relational data model, relational DBMS implementation
1980s:
RDBMS, advanced data models (extended-relational, OO, deductive, etc.)
Application-oriented DBMS (spatial, scientific, engineering, etc.)
1990s:
Data mining, data warehousing, multimedia databases, and Web
databases
2000s
Stream data management and mining
Data mining and its applications
Web technology (XML, data integration) and global information systems
8
Chapter 1. Introduction
Why Data Mining?
What Is Data Mining?
A Multi-Dimensional View of Data Mining
What Kind of Data Can Be Mined?
What Kinds of Patterns Can Be Mined?
What Technology Are Used?
What Kind of Applications Are Targeted?
9
What Is Data Mining?
10
August 19, 2020 Data Mining: Concepts and Techniques 11
Knowledge Discovery (KDD) Process
This is a view from typical
database systems and data
warehousing communities Pattern Evaluation
Identify interesting patterns
Data mining plays an essential role
in the knowledge discovery process
Data Mining
Task-relevant Data
Data Transformation
Selection
To retrieve data relevant to analysis
Data
Warehouse
Data Cleaning
Remove noise & inconsistent data
Data Integration
Combine multiple data Sources
Databases
12
Example: A Web Mining Framework
13
Focuses search towards interesting patterns
19
Multi-Dimensional View of Data Mining
Data to be mined
Database data (extended-relational, object-oriented, heterogeneous,
Techniques utilized
Data-intensive, data warehouse (OLAP), machine learning, statistics,
21
Data Mining: On What Kinds of Data?
Database-oriented data sets and applications
Relational database, data warehouse, transactional database
Advanced data sets and advanced applications
Data streams and sensor data
Time-series data, temporal data, sequence data (incl. bio-sequences)
Structure data, graphs, social networks and multi-linked data
Object-relational databases
Heterogeneous databases and legacy databases
Spatial data and spatiotemporal data
Multimedia database
Text databases
The World-Wide Web
Ref Section 1.3– Self Study
22
Chapter 1. Introduction
Why Data Mining?
What Is Data Mining?
A Multi-Dimensional View of Data Mining
What Kind of Data Can Be Mined?
What Kinds of Patterns Can Be Mined?
What Technology Are Used?
What Kind of Applications Are Targeted?
23
Data Mining Function: Association and
Correlation Analysis
Frequent patterns (or frequent item sets)
What items are frequently purchased together in your
Chaldal/Lavender?
Association, correlation vs. causality
A typical association rule
Diaper Milk [0.5%, 75%] (support, confidence)
How to mine such patterns and rules efficiently in large
datasets?
25
Correlation Vs Causality
27
Classification
29
Data Mining Function: Cluster
Analysis
31
Data Mining Function: Outlier
Analysis
36
Data Mining: Confluence of Multiple Disciplines
37
Chapter 1. Introduction
Why Data Mining?
What Is Data Mining?
A Multi-Dimensional View of Data Mining
What Kind of Data Can Be Mined?
What Kinds of Patterns Can Be Mined?
What Technology Are Used?
What Kind of Applications Are Targeted?
39
Applications of Data Mining
Web page analysis: from web page classification, clustering to
PageRank & HITS algorithms
Collaborative analysis & recommender systems
Basket data analysis to targeted marketing
Biological and medical data analysis: classification, cluster analysis
(microarray data analysis), biological sequence analysis, biological
network analysis
Major dedicated data mining systems/tools (e.g., SAS, MS SQL-Server
Analysis Manager, Oracle Data Mining Tools).
40