01
A metric defines what counts as success
There is no universally best machine-learning metric. Accuracy, precision, recall, squared error, cross-entropy, ranking metrics and task-specific utility measure different properties.
The metric should reflect the actual cost of mistakes. A system for detecting a rare dangerous event may care more about recall and false negatives than overall accuracy.
02
Classification metrics
Precision = TP / (TP + FP) Recall = TP / (TP + FN) F1 = 2PR / (P + R)
Precision measures how many predicted positives are correct. Recall measures how many actual positives are detected. F1 combines them through their harmonic mean.
03
Regression metrics
MSE = (1/n)Σ(yᵢ - ŷᵢ)² RMSE = √MSE
Mean squared error strongly penalizes large errors. Mean absolute error is less sensitive to outliers. The appropriate metric depends on how prediction error translates into downstream consequences.
04
Calibration
A probabilistic classifier can have good discrimination while being poorly calibrated. A calibrated model producing predictions around 0.8 should be correct roughly 80 percent of the time within an appropriate population and grouping.
Calibration matters when predictions are consumed as probabilities rather than only as rankings or hard classes.
05
Comparing models
Model comparisons are meaningful only when the evaluation protocol is held constant and uncertainty is considered. Small differences in a metric may not represent meaningful improvements, particularly when the test sample is small.