ESSAY · PROJECT 0002
A computer has never seen a cat — only numbers
The first session of the program covered convolutional networks and ResNet, the part of the curriculum where deep learning becomes something you can watch. This page rebuilds it in plain words: what a computer actually receives when it looks at a picture, the small walking window that starts everything, why stacking layers turns edges into parts into objects, and the shortcut trick that finally made deep networks trainable.
Where this comes from
- The curriculum: ARENA 3.0, part 0.2: CNNs & ResNets, where you can build every piece of this page yourself, window by window.
- The shortcut result: Deep Residual Learning (He et al., 2015), the paper behind "ResNet" and the express-lane fix.
- Every number and picture on the page is hand-set and illustrative; nothing is measured output.
- Next in the series: what a language model does with words, once the pictures are done.