Tree Based Classifiers: Dinesh R

Tree-based classifiers are machine learning models that use a decision tree as a predictive model. Decision trees classify instances by starting at the root node and moving through the tree recursively according to a test at each node until a leaf node is reached, which provides the classification or predicted value. Tree-based classifiers are powerful tools for classification and prediction that represent rules in an interpretable way. Building decision trees involves splitting the training data into nodes based on attribute values to create branches until the data is partitioned into distinct target classes.

Uploaded by

dhruva

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PPTX, PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

75 views54 pages

Tree Based Classifiers: Dinesh R

Uploaded by

dhruva

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PPTX, PDF, TXT or read online on Scribd

You are on page 1/ 54

Tree Based Classifiers

Dinesh R
Principal Engineer,
Samsung, Bangalore
Email: [email protected]
Contents
• Background
• Tree classifiers
• Applications
• Building decision trees
• Entropy and GINI index for tree building
• Tree Pruning
• Challenges
Classifiers
• Bayesian Classifier
• K-Nearest Neighborhood classifier
• Decision Trees
• Boosting Classifiers
• SVM
• Neural Networks
Classification: Definition
• Given a collection of records (training set )
– Each record contains a set of attributes, one of the attributes is the
class.
• Find a model for class attribute as a function
of the values of other attributes.
• Goal: previously unseen records should be
assigned a class as accurately as possible.
– A test set is used to determine the accuracy of the model. Usually,
the given data set is divided into training and test sets, with training
set used to build the model and test set used to validate it.
Illustrating Classification Task
Tid Attrib1 Attrib2 Attrib3 Class Learning
No
1 Yes Large 125K
algorithm
2 No Medium 100K No
3 No Small 70K No

4 Yes Medium 120K No

Induction
5 No Large 95K Yes
6 No Medium 60K No
7 Yes Large 220K No Learn
8 No Small 85K Yes Model
9 No Medium 75K No
10 No Small 90K Yes
Model
10

Training Set
Apply
Tid Attrib1 Attrib2 Attrib3 Class Model
11 No Small 55K ?
12 Yes Medium 80K ?
13 Yes Large 110K ? Deduction
14 No Small 95K ?

15 No Large 67K ?
10

Test Set
Classification Using Distance
• Place items in class to which they are
“closest”.
• Must determine distance between an item
and a class.
• Classes represented by
– Centroid: Central value.
– Medoid: Representative point.
– Individual points
• Algorithm: KNN
K Nearest Neighbor (KNN):
• Training set includes classes.
• Examine K items near item to be classified.
• New item placed in class with the most
number of close items.
• O(q) for each tuple to be classified. (Here q
is the size of the training set.)
KNN
Limitation of KNN
No Model learning
Extremely slow
Definition

 Decision tree is a classifier in the form of a tree structure

– Decision node: specifies a test on a single attribute
– Leaf node: indicates the value of the target attribute
– Arc/edge: split of one attribute
– Path: a disjunction of test to make the final decision

 Decision trees classify instances or examples by starting at the

root of the tree and moving through it until a leaf node.
Why decision tree?

• Decision trees are powerful and popular tools for

classification and prediction.
• Decision trees represent rules, which can be understood
by humans and used in knowledge system such as
database.
key requirements
• Attribute-value description: object or case must be expressible
in terms of a fixed collection of properties or attributes (e.g., hot,
mild, cold).
• Predefined classes (target values): the target function has
discrete output values (bollean or multiclass)
• Sufficient data: enough training cases should be provided to learn
the model.
Example of a Decision Tree
l l us
ir ca ir ca o
go go tinu ss
te te n la
ca ca co c
Tid Refund Marital Taxable
Splitting Attributes
Status Income Cheat

1 Yes Single 125K No

2 No Married 100K No Refund
No
Yes No
3 No Single 70K
4 Yes Married 120K No NO MarSt
5 No Divorced 95K Yes Married
Single, Divorced
6 No Married 60K No
7 Yes Divorced 220K No TaxInc NO
8 No Single 85K Yes < 80K > 80K
9 No Married 75K No
NO YES
10 No Single 90K Yes
10

Training Data Model: Decision Tree

Another Example of Decision Tree
l l us
rica rica o
go go tinu ss
te te n l a Single,
ca ca co c MarSt
Married Divorced
Tid Refund Marital Taxable
Status Income Cheat
NO Refund
1 Yes Single 125K No
Yes No
2 No Married 100K No
3 No Single 70K No NO TaxInc
4 Yes Married 120K No < 80K > 80K
5 No Divorced 95K Yes
NO YES
6 No Married 60K No
7 Yes Divorced 220K No
8 No Single 85K Yes
9 No Married 75K No There could be more than one tree that fits
10 No Single 90K Yes the same data!
10
Decision Tree Classification Task
Tid Attrib1 Attrib2 Attrib3 Class
Tree
1 Yes Large 125K No Induction
2 No Medium 100K No algorithm
3 No Small 70K No

4 Yes Medium 120K No

Induction
5 No Large 95K Yes
6 No Medium 60K No
7 Yes Large 220K No Learn
8 No Small 85K Yes Model
9 No Medium 75K No
10 No Small 90K Yes
Model
10

Training Set
Apply Decision Tree
Tid Attrib1 Attrib2 Attrib3 Class
Model
11 No Small 55K ?
12 Yes Medium 80K ?
13 Yes Large 110K ?
Deduction
14 No Small 95K ?
15 No Large 67K ?
10

Test Set
Apply Model to Test Data
Test Data
Start from the root of tree. Refund Marital Taxable
Status Income Cheat

No Married 80K ?
Refund 10

Yes No

NO MarSt
Single, Divorced Married

TaxInc NO
< 80K > 80K