siddhant

Knowledge / Machine Learning

Practical Machine Learning

How real machine-learning projects are designed, evaluated, deployed and maintained.

By Siddhant Krishna · Published 2026-10-06 · Updated 2026-10-06

01

Start with the problem

A machine-learning project should begin with a precise task and success criterion, not with a model. The team should define what input is available at prediction time, what the target represents, what decisions depend on the output and what errors matter.

02

Baselines

A simple baseline provides a reference point. Depending on the task this may be a constant predictor, linear model, heuristic, retrieval system or existing business rule.

Without a baseline, it is easy to mistake model complexity for progress.

03

Data quality

  • Missing values
  • Incorrect labels
  • Duplicate records
  • Sampling bias
  • Temporal leakage
  • Train-test contamination
  • Unrepresentative deployment data

In many projects, improving the data-generating and labelling process produces more reliable gains than swapping one sophisticated model for another.

04

Validation strategy

The validation scheme should match deployment. Random splits can be inappropriate for temporal prediction, grouped observations, repeated measurements or data with strong correlations.

A test set should remain sufficiently isolated from iterative model development to preserve its role as an unbiased estimate of final performance.

05

Error analysis

Aggregate metrics hide structure. A useful error-analysis process identifies representative failures, groups them by cause, estimates their frequency and determines which intervention is likely to reduce them.

  • Improve data
  • Change labels
  • Change features
  • Change model
  • Change threshold
  • Add a specialized component
  • Redefine the task

06

Deployment and monitoring

A model is part of a system. Production reliability depends on input validation, latency, resource usage, model versioning, observability, rollback procedures and monitoring for data and performance changes.

References

  1. Stanford University, CS229 Machine Learning. Course materials covering supervised and unsupervised learning, learning theory, regularization, SVMs and reinforcement learning.
    https://cs229.stanford.edu/
  2. Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani and Jonathan Taylor. An Introduction to Statistical Learning.
    https://www.statlearning.com/
  3. Ian Goodfellow, Yoshua Bengio and Aaron Courville. Deep Learning. MIT Press.
    https://www.deeplearningbook.org/

Related

Contact

Get in Touch

Want to chat? Just shoot me a dm with a direct question on twitter and I'll respond whenever I can. I will ignore all soliciting.