Embeddings: From One-Hot Vectors to Learned Representations

Embeddings: From One-Hot Vectors to Learned Representations Everything in the previous notes assumed the network’s input was already a list of numbers — x1 = 1, x2 = 2, pixel intensities, whatever. But most interesting data is not numeric. “cat”, “dog”, “bank”, user IDs, product IDs, words of a sentence. Neural networks cannot multiply the string "cat" by a weight matrix. Somewhere between the raw symbol and the first linear layer, a translation to numbers must happen, and the way we do it — the embedding layer — turns out to be one of the most consequential ideas in modern deep learning. ...

September 14, 2026 · 31 min