01
Decision trees
A decision tree recursively partitions the feature space using questions about feature values. A leaf stores a prediction or distribution over classes.
Tree construction chooses splits according to an impurity or loss criterion. Deep trees can fit highly complex patterns but can also overfit, which motivates pruning and ensemble methods.
02
Random forests
A random forest averages many decision trees trained with randomness in both examples and feature selection. Averaging reduces variance and makes the method substantially more robust than a single unconstrained tree.
03
Boosting
Boosting constructs a sequence of weak or moderately strong learners, with later learners focused on correcting errors made by earlier ones. Gradient boosting interprets this process as optimization in function space.
Modern gradient-boosted decision-tree systems are especially effective on structured and tabular datasets.
04
Support vector machines
A support vector machine seeks a decision boundary with a large margin between classes. Only selected training examples, the support vectors, directly determine the optimal boundary.
minimize 1/2 ||w||² subject to yᵢ(wᵀxᵢ+b) ≥ 1
Kernel methods extend the idea by computing inner products in implicit feature spaces, allowing nonlinear decision boundaries without explicitly constructing the transformed coordinates.