Training a 2-Layer Network in NumPy: From Scalar to Vectorized Backprop

Training a 2-Layer Network in NumPy: From Scalar to Vectorized Backprop The previous note did the backward pass for a small 2-layer network by hand, one scalar at a time. That is the best way to understand what backprop actually does. This note takes the next step: turn that scalar walk into compact, vectorized NumPy code. The math is unchanged; the only thing that changes is notation. I will introduce every matrix slowly — what its rows and columns mean, where the division by the batch size comes from, and why it is exactly the same algorithm you already did by hand. ...

September 14, 2026 · 20 min