← All writing
September 27, 2026 · INTERACTIVE ESSAY

From a Line to a Neural Network

An interactive journey through machine learning: least squares, gradient descent, logistic regression, k-nearest neighbors, decision trees, feature engineering, and a neural network that trains in your browser.

A visual journey through model capacity

How does a machine go from fitting one line to learning its own representation?

Machine learning can look like a bag of unrelated algorithms. It makes more sense as a sequence of ideas: define a function, measure its mistakes, optimize its parameters, introduce nonlinearity, and finally let the model learn the features themselves.

01 · LINEAR REGRESSION fit 02 · GRADIENT DESCENT optimize 03 · LOGISTIC REGRESSION classify 04 · k-NN remember 05 · DECISION TREES partition 06 · FEATURES represent 07 · NEURAL NETWORKS learn features
01 / FIT A LINE

Linear regression: the smallest useful learning machine.

Suppose we observe pairs of numbers — study hours and exam score, house size and price, temperature and energy use. We want a rule that predicts one from the other. The simplest useful assumption is that the relationship is approximately linear.

A linear model has only two learnable parameters: an intercept and a slope. Training means choosing those two numbers so that the line passes as close as possible to the observed data.

ŷ = b₀ + b₁xPREDICTION = INTERCEPT + SLOPE × INPUT

For every observed point, the vertical gap between the prediction and the observation is a residual. Least-squares regression chooses the line that minimizes the sum of the squared residuals. Squaring matters: positive and negative errors cannot cancel, and large mistakes receive a larger penalty.

LEAST-SQUARES LAB
DRAG ANY WHITE POINT · THE FIT UPDATES IN REAL TIME
Slope b₁—
Intercept b₀—
Mean squared error—
R²—
ŷ = …

What the model is actually learning

Nothing inside linear regression “knows” what a house, temperature, or exam is. The model only sees numbers and searches for parameters that make a chosen loss function small. That separation — model + loss + optimization — becomes the template for almost everything that follows.

Important limitationA straight line can only express a straight-line relationship. If the true structure bends, branches, loops, or depends on interactions between features, the model is structurally incapable of representing it.
02 / OPTIMIZE

Gradient descent: learning by walking downhill.

For ordinary least squares, we can solve for the best line directly. But many machine-learning models have millions or billions of parameters and no practical closed-form solution. We need an iterative way to improve them.

Imagine every possible parameter setting as a point on a landscape. The height of the landscape is the loss. A gradient tells us the direction of steepest increase, so we move in the opposite direction.

θ ← θ − η ∇L(θ)PARAMETERS ← PARAMETERS − LEARNING RATE × LOSS GRADIENT

The learning rate η controls the step size. Too small and learning crawls. Too large and the optimizer can overshoot or diverge. This same basic idea later trains neural networks.

GRADIENT-DESCENT TRAINER
START FROM A BAD LINE · WATCH OPTIMIZATION REPAIR IT
Slope—
Intercept—
Loss—
Iteration0
0.080
Batch gradient descent

Use all examples.

Compute one gradient from the full dataset. The update is stable but can be expensive when the dataset is huge.

Stochastic / mini-batch

Use a small sample.

Estimate the gradient from a subset. Updates become noisy, but training becomes scalable — the standard recipe for deep learning.

03 / CLASSIFY

Logistic regression: turn a score into a probability.

Regression predicts a continuous value. Classification asks a different question: which class does this example belong to? Logistic regression keeps the linear score, then squeezes it through a sigmoid function so the output lies between 0 and 1.

p(y=1|x) = 1 / (1 + e−(wᵀx+b))A LINEAR SCORE BECOMES A PROBABILITY

A probability is not yet a decision. We still choose a threshold. Moving that threshold changes the trade-off between false positives and false negatives. That is why “accuracy” alone rarely tells the whole story.

CLASSIFICATION THRESHOLD LAB
MOVE THE THRESHOLD · WATCH THE DECISION BOUNDARY AND ERRORS CHANGE
0.50
TRUE +—
FALSE +—
FALSE −—
TRUE −—
Precision—
Recall—

The first important boundary

Despite the sigmoid, logistic regression still creates a linear decision boundary in the original feature space. It can separate two groups with a line, plane, or hyperplane — but it cannot naturally carve out complicated regions.

At this point we have a pattern

Choose a function. Define a loss. Optimize parameters.

Different algorithms now disagree mainly about what function family to use and how much structure to assume.

04 / REMEMBER

k-nearest neighbors: what if we barely train at all?

k-NN takes a radically different approach. It stores the training examples. When a new point arrives, it asks which stored examples are closest and lets those neighbors vote.

There is almost no conventional training step. The inductive bias is local: nearby points should have similar labels. The hyperparameter k controls how local the decision is.

k-NN PROBE
MOVE YOUR CURSOR THROUGH THE FIELD · LINES SHOW THE k NEAREST TRAINING POINTS
5
PREDICTION: — —
Small k

Flexible, but noisy.

A tiny neighborhood can follow intricate boundaries but is sensitive to outliers and mislabeled examples.

Large k

Smooth, but biased.

A large neighborhood averages away noise, but can erase small local structures. This is the bias–variance trade-off in visible form.

05 / PARTITION

Decision trees: learn rules by slicing feature space.

A tree repeatedly asks simple questions such as “is feature 1 smaller than 0.42?” Each question divides the data, and the recursive sequence of splits creates a piecewise decision surface.

A classification tree searches for splits that make the child groups purer. A common impurity measure is Gini impurity. Deep trees can model surprisingly complicated patterns, but they can also memorize noise.

Gini = 1 − Σ pk²ZERO MEANS A NODE CONTAINS ONLY ONE CLASS
LIVE CART PARTITIONER
THIS MINI TREE SEARCHES REAL SPLITS IN JAVASCRIPT
Leaf regions—
Training accuracy—
3

From one tree to a forest

Single trees are unstable: a small change in data can produce a different structure. Random forests reduce that variance by averaging many decorrelated trees. Gradient-boosted trees take another route: add weak trees sequentially, with each new tree focusing on the current residual errors.

06 / REPRESENT

The hidden bottleneck: the model can only use the representation you give it.

Classical machine learning often depends heavily on feature engineering. A linear classifier may fail not because the optimizer is weak, but because the useful structure is invisible in the original coordinates.

The XOR problem

A simple pattern that defeats a straight boundary.

In raw coordinates, opposite corners share a class. No single straight line can separate them. Add the interaction feature x₁×x₂, however, and the classes become linearly separable in the transformed representation.

RAW FEATURES

This is the conceptual bridge to neural networks. Instead of asking a human to invent every useful interaction, can the model learn a sequence of useful transformations directly from data?

The deep-learning move

Stop hand-designing every feature. Learn the representation.

A neural network stacks parameterized transformations. Training adjusts not just the final decision rule, but the intermediate representation itself.

07 / LEARN FEATURES

Neural networks: linear layers become powerful when we insert nonlinearity.

A single neuron looks familiar: weighted inputs plus a bias, followed by an activation. The power comes from composing many such units into layers.

h = φ(Wx + b)   →   ŷ = σ(Vh + c)LINEAR TRANSFORM → NONLINEAR ACTIVATION → ANOTHER TRANSFORM

If we stacked only linear layers, the entire stack would collapse into one linear transformation. Nonlinear activations are what allow the network to bend, fold, and reorganize feature space.

Backpropagation is gradient descent with efficient bookkeeping

We compute the prediction, evaluate the loss, then use the chain rule to propagate how each parameter contributed to the error. The gradient flows backward through the computational graph; the parameter update still looks like the gradient-descent rule we used for the simple line.

2 → 6 → 1 NEURAL NETWORK · XOR
THE NETWORK BELOW REALLY TRAINS IN YOUR BROWSER — NO PRECOMPUTED ANIMATION
Epoch0
Cross-entropy loss—
Training accuracy—
Architecture2·6·1
0.35

Press TRAIN. At initialization the colored probability field is essentially random. As backpropagation updates the weights, the hidden layer discovers a representation that makes the XOR classes separable. The final output unit can then solve a problem that logistic regression could not solve in the raw coordinates.

The surprising continuity

A modern neural network is vastly more expressive than linear regression, but the conceptual skeleton is remarkably similar:

  • There is still a parameterized function.
  • There is still a prediction.
  • There is still a loss that measures error.
  • There is still gradient-based optimization.
  • The major upgrade is that intermediate features are learned jointly with the predictor.
THE CORE IDEA · PSEUDOCODE
# forward pass
h = tanh(W1 @ x + b1)
ŷ = sigmoid(W2 @ h + b2)

# measure error
loss = binary_cross_entropy(ŷ, y)

# backward pass: chain rule gives gradients
gradients = backpropagate(loss)

# same optimization idea as before
parameters = parameters - learning_rate * gradients
08 / CONNECT IT

One field, many inductive biases.

These models are not a simple ladder where newer always means better. Each algorithm encodes assumptions about the shape of the problem, the amount of data, interpretability, compute, and the representation available.

MODELCORE IDEASTRENGTHLIMITATION
Linear regressionFit a weighted sumSimple, interpretable, data-efficientOnly linear relationships unless features are engineered
Logistic regressionLinear score → probabilityStrong baseline for classificationLinear decision boundary in input space
k-NNLet nearby examples voteVery flexible local behaviorPrediction cost and distance sensitivity
Decision treeRecursive feature-space splitsReadable nonlinear rulesSingle trees can overfit and be unstable
Neural networkLearn stacked nonlinear representationsExpressive feature learningMore data, compute, tuning, and analysis

The mental model I keep

When I encounter a new machine-learning architecture, I ask four questions: What family of functions can it represent? What loss gives it a learning signal? How are the parameters optimized? What representation makes the task easy? Those questions connect classical machine learning to deep learning far better than memorizing an isolated list of algorithms.

The whole journey in one sentence

Machine learning is the art of choosing — or learning — a useful function space.

Linear models choose a tiny function family. Trees build piecewise rules. Nearest-neighbor methods borrow structure from local examples. Neural networks go further: they learn the representation and the predictor together.

01MODEL
what can it express?
02LOSS
what counts as wrong?
03OPTIMIZER
how do parameters improve?
04REPRESENTATION
what makes the task easy?