From a Line to a Neural Network
An interactive journey through machine learning: least squares, gradient descent, logistic regression, k-nearest neighbors, decision trees, feature engineering, and a neural network that trains in your browser.
How does a machine go from fitting one line to learning its own representation?
Machine learning can look like a bag of unrelated algorithms. It makes more sense as a sequence of ideas: define a function, measure its mistakes, optimize its parameters, introduce nonlinearity, and finally let the model learn the features themselves.
Linear regression: the smallest useful learning machine.
Suppose we observe pairs of numbers — study hours and exam score, house size and price, temperature and energy use. We want a rule that predicts one from the other. The simplest useful assumption is that the relationship is approximately linear.
A linear model has only two learnable parameters: an intercept and a slope. Training means choosing those two numbers so that the line passes as close as possible to the observed data.
For every observed point, the vertical gap between the prediction and the observation is a residual. Least-squares regression chooses the line that minimizes the sum of the squared residuals. Squaring matters: positive and negative errors cannot cancel, and large mistakes receive a larger penalty.
What the model is actually learning
Nothing inside linear regression “knows” what a house, temperature, or exam is. The model only sees numbers and searches for parameters that make a chosen loss function small. That separation — model + loss + optimization — becomes the template for almost everything that follows.
Gradient descent: learning by walking downhill.
For ordinary least squares, we can solve for the best line directly. But many machine-learning models have millions or billions of parameters and no practical closed-form solution. We need an iterative way to improve them.
Imagine every possible parameter setting as a point on a landscape. The height of the landscape is the loss. A gradient tells us the direction of steepest increase, so we move in the opposite direction.
The learning rate η controls the step size. Too small and learning crawls. Too large and the optimizer can overshoot or diverge. This same basic idea later trains neural networks.
Use all examples.
Compute one gradient from the full dataset. The update is stable but can be expensive when the dataset is huge.
Use a small sample.
Estimate the gradient from a subset. Updates become noisy, but training becomes scalable — the standard recipe for deep learning.
Logistic regression: turn a score into a probability.
Regression predicts a continuous value. Classification asks a different question: which class does this example belong to? Logistic regression keeps the linear score, then squeezes it through a sigmoid function so the output lies between 0 and 1.
A probability is not yet a decision. We still choose a threshold. Moving that threshold changes the trade-off between false positives and false negatives. That is why “accuracy” alone rarely tells the whole story.
The first important boundary
Despite the sigmoid, logistic regression still creates a linear decision boundary in the original feature space. It can separate two groups with a line, plane, or hyperplane — but it cannot naturally carve out complicated regions.
Choose a function. Define a loss. Optimize parameters.
Different algorithms now disagree mainly about what function family to use and how much structure to assume.
k-nearest neighbors: what if we barely train at all?
k-NN takes a radically different approach. It stores the training examples. When a new point arrives, it asks which stored examples are closest and lets those neighbors vote.
There is almost no conventional training step. The inductive bias is local: nearby points should have similar labels. The hyperparameter k controls how local the decision is.
Flexible, but noisy.
A tiny neighborhood can follow intricate boundaries but is sensitive to outliers and mislabeled examples.
Smooth, but biased.
A large neighborhood averages away noise, but can erase small local structures. This is the bias–variance trade-off in visible form.
Decision trees: learn rules by slicing feature space.
A tree repeatedly asks simple questions such as “is feature 1 smaller than 0.42?” Each question divides the data, and the recursive sequence of splits creates a piecewise decision surface.
A classification tree searches for splits that make the child groups purer. A common impurity measure is Gini impurity. Deep trees can model surprisingly complicated patterns, but they can also memorize noise.
From one tree to a forest
Single trees are unstable: a small change in data can produce a different structure. Random forests reduce that variance by averaging many decorrelated trees. Gradient-boosted trees take another route: add weak trees sequentially, with each new tree focusing on the current residual errors.
The hidden bottleneck: the model can only use the representation you give it.
Classical machine learning often depends heavily on feature engineering. A linear classifier may fail not because the optimizer is weak, but because the useful structure is invisible in the original coordinates.
A simple pattern that defeats a straight boundary.
In raw coordinates, opposite corners share a class. No single straight line can separate them. Add the interaction feature x₁×x₂, however, and the classes become linearly separable in the transformed representation.
This is the conceptual bridge to neural networks. Instead of asking a human to invent every useful interaction, can the model learn a sequence of useful transformations directly from data?
Stop hand-designing every feature. Learn the representation.
A neural network stacks parameterized transformations. Training adjusts not just the final decision rule, but the intermediate representation itself.
Neural networks: linear layers become powerful when we insert nonlinearity.
A single neuron looks familiar: weighted inputs plus a bias, followed by an activation. The power comes from composing many such units into layers.
If we stacked only linear layers, the entire stack would collapse into one linear transformation. Nonlinear activations are what allow the network to bend, fold, and reorganize feature space.
Backpropagation is gradient descent with efficient bookkeeping
We compute the prediction, evaluate the loss, then use the chain rule to propagate how each parameter contributed to the error. The gradient flows backward through the computational graph; the parameter update still looks like the gradient-descent rule we used for the simple line.
Press TRAIN. At initialization the colored probability field is essentially random. As backpropagation updates the weights, the hidden layer discovers a representation that makes the XOR classes separable. The final output unit can then solve a problem that logistic regression could not solve in the raw coordinates.
The surprising continuity
A modern neural network is vastly more expressive than linear regression, but the conceptual skeleton is remarkably similar:
- There is still a parameterized function.
- There is still a prediction.
- There is still a loss that measures error.
- There is still gradient-based optimization.
- The major upgrade is that intermediate features are learned jointly with the predictor.
# forward pass h = tanh(W1 @ x + b1) ŷ = sigmoid(W2 @ h + b2) # measure error loss = binary_cross_entropy(ŷ, y) # backward pass: chain rule gives gradients gradients = backpropagate(loss) # same optimization idea as before parameters = parameters - learning_rate * gradients
One field, many inductive biases.
These models are not a simple ladder where newer always means better. Each algorithm encodes assumptions about the shape of the problem, the amount of data, interpretability, compute, and the representation available.
| MODEL | CORE IDEA | STRENGTH | LIMITATION |
|---|---|---|---|
| Linear regression | Fit a weighted sum | Simple, interpretable, data-efficient | Only linear relationships unless features are engineered |
| Logistic regression | Linear score → probability | Strong baseline for classification | Linear decision boundary in input space |
| k-NN | Let nearby examples vote | Very flexible local behavior | Prediction cost and distance sensitivity |
| Decision tree | Recursive feature-space splits | Readable nonlinear rules | Single trees can overfit and be unstable |
| Neural network | Learn stacked nonlinear representations | Expressive feature learning | More data, compute, tuning, and analysis |
The mental model I keep
When I encounter a new machine-learning architecture, I ask four questions: What family of functions can it represent? What loss gives it a learning signal? How are the parameters optimized? What representation makes the task easy? Those questions connect classical machine learning to deep learning far better than memorizing an isolated list of algorithms.
Machine learning is the art of choosing — or learning — a useful function space.
Linear models choose a tiny function family. Trees build piecewise rules. Nearest-neighbor methods borrow structure from local examples. Neural networks go further: they learn the representation and the predictor together.
what can it express?
what counts as wrong?
how do parameters improve?
what makes the task easy?