raw-neural-net
A small deep learning framework built on NumPy, with every backward pass written by hand.
- Sole engineer
- 2023
- archived
- Python, NumPy, OpenCV
- Source
- 4
- 4
This is a learning project. It follows the standard from-scratch curriculum, it is not an original architecture, and the repository is untidy in the way a repository is when its purpose was the writing rather than the result.
The goal was to stop treating the training loop as a black box. Calling
model.fit() teaches you the API, not what the gradient flowing back through a
softmax looks like.
What is in it
Dense layers with L1 and L2 regularisation, dropout, ReLU, softmax, sigmoid and linear activations, four optimizers in SGD with momentum, Adagrad, RMSprop and Adam, four losses in categorical cross-entropy, binary cross-entropy, mean squared error and mean absolute error, and accuracy implementations for both classification and regression.
On top sits a model class that links layers to their neighbours, runs forward and backward passes, batches, validates, and saves and loads trained parameters.
NumPy does the array arithmetic. “No framework” here means no autograd, no computation graph, and no library computing a derivative for me. Every backward pass is written by hand. The matrix multiplication is NumPy’s.
The two parts worth the time
Most of the code is mechanical once the derivation is done. Two parts are not.
The fused softmax and cross-entropy backward pass. Taken separately, the derivative of softmax is a Jacobian per sample, and backpropagating through it is a matrix multiply. Composed with categorical cross entropy, almost all of it cancels and the combined gradient collapses to:
self.dinputs = dvalues.copy()
self.dinputs[range(samples), y_true] -= 1
self.dinputs = self.dinputs / samples
Subtract one from the predicted probability at the correct class, divide by the batch size. In a framework that looks like a trick. Deriving it makes it obvious, and it is the best argument for doing this exercise.
Adam’s bias correction. The momentum and cache terms both start at zero, so early in training they are biased toward zero and the first steps come out too small. The correction divides each by one minus the decay rate raised to the step number, which is large early and decays to nothing:
weight_momentums_corrected = layer.weight_momentums / \
(1 - self.beta_1 ** (self.iterations + 1))
The + 1 is there because the iteration counter starts at zero, and at zero the
denominator would be exactly zero.
Then I photographed my own clothes
Fashion-MNIST is a solved dataset, so a good score on it proves little. The real test was a shirt, a pair of trousers and a sneaker, photographed on my own phone.
That surfaced something the dataset hides. Fashion-MNIST is a light garment on a black background. A photograph of clothing is usually the opposite, a darker object against a lighter surface, so a model trained on one and shown the other gets handed a negative of what it learned. The preprocessing has to read the image as greyscale, resize to 28 by 28, and invert it before scaling into the range the network was trained on.
That is not difficult and it is not in the dataset documentation, because it only shows up once real input arrives. The training distribution is a choice somebody made, and real inputs do not have to match it.
It classified the photographs correctly. I did not record an accuracy figure at the time and I am not going to invent one, so there is no number to quote here. Three garments is not an evaluation. It answered the question I was asking, which was whether the thing worked outside the sandbox it was built in.
What this is worth
Not a system anyone runs. It is the answer to one interview question: does this person understand what the library is doing for them, or only how to call it.
I can say why softmax and cross entropy get fused, what Adam’s correction is for and why it disappears as training goes on, why dropout scales its output during training rather than at inference, and where regularisation enters the gradient, because the training loop did not converge until I got each of those right.