A collection of Machine Learning algorithms implemented from scratch using mathematics and Python, with their results compared against equivalent models from scikit-learn.
The main goal is to understand how ML algorithms work internally instead of using only ready-made libraries.
Each model focuses on:
- Mathematical foundation
- From-scratch Python implementation
- Training and prediction
- Model evaluation
- Comparison with
scikit-learn
Used for predicting continuous values.
Linear Model:
Mean Squared Error:
Normal Equation:
scikit-learn: LinearRegression
Used for binary classification.
Sigmoid Function:
where:
Binary Cross-Entropy Loss:
scikit-learn: LogisticRegression
Classifies a sample based on its nearest neighbors.
Euclidean Distance:
The predicted class is generally the majority class among the (k) nearest samples.
scikit-learn: KNeighborsClassifier
Builds a tree by recursively splitting data based on feature values.
Entropy:
Information Gain:
scikit-learn: DecisionTreeClassifier
A probabilistic classification algorithm based on Bayes' theorem.
Bayes' Theorem:
With the conditional independence assumption:
scikit-learn: GaussianNB
An unsupervised learning algorithm that divides data into (k) clusters.
Euclidean Distance:
Centroid Update:
Objective Function:
scikit-learn: KMeans
An ensemble learning algorithm that combines multiple Decision Trees.
For classification, the final prediction is generally based on majority voting:
where
scikit-learn: RandomForestClassifier
| Model | From Scratch | scikit-learn |
|---|---|---|
| Linear Regression | Mathematical implementation | LinearRegression |
| Logistic Regression | Sigmoid + Gradient Descent | LogisticRegression |
| KNN | Distance-based implementation | KNeighborsClassifier |
| Decision Tree | Entropy + Information Gain | DecisionTreeClassifier |
| Naive Bayes | Bayes Theorem | GaussianNB |
| K-Means | Centroid-based clustering | KMeans |
| Random Forest | Multiple Decision Trees | RandomForestClassifier |
- Python
- NumPy
- Pandas
- Matplotlib
- scikit-learn
- Jupyter Notebook
This project is created for understanding Machine Learning algorithms from their mathematical foundations.
Instead of directly using ML libraries, the algorithms are first implemented from scratch and then compared with optimized scikit-learn implementations.
Possible algorithms to add:
- Support Vector Machine
- PCA
- Gradient Boosting
- Neural Networks
- DBSCAN
- AdaBoost
ARC
Machine Learning / Data Analytics Learning Project