How a network learns to see

convolutions and ResNets, in plain words · no code, no math · about 8 minutes
BLOCK 1 · THE RAW MATERIAL

To a computer, a picture is a grid of numbers

You see a cat. The computer sees a grid of brightness values — that is the entire input. Every image task starts here: numbers, arranged in a grid, where each cell is called a pixel.

Flip the switch, then hover the cells.

Hover a cell to read its value. The whiskers, the ears, the eyes — all of it is just brighter and darker cells.
BLOCK 2 · THE FIRST MOVE

A small window asks one question, everywhere

Nobody looks at a whole picture at once. The first thing a vision network does is send a tiny window walking across the grid. The window carries one pattern — say, a vertical edge, dark on the left and light on the right — and at every position it asks: how much is my pattern here?

Step the window across the cat. The score column on the right fills in as it goes.

This walking window is the famous convolution. One window, one pattern, every position — and a real network trains hundreds of different windows, each hunting its own pattern.
BLOCK 3 · THE STACK

Stack the layers and features build up

Here is the payoff of all those little windows. Layer one finds edges. Layer two finds patterns made of edges — corners, curves, textures. Layer three finds patterns made of those: ears, eyes, whiskers. Layer four doesn't see parts anymore. It sees cat. Nobody programmed any of the levels; they emerge from stacking.

Slide through the depth and watch the description change.

BLOCK 4 · THE PROBLEM

Deep should be better — so why did deep networks refuse to train?

By the stack logic, more layers should mean smarter features. For years it didn't work: past a certain depth, networks trained worse, not better. The plain-words version of why: learning is a message passed backwards from the last layer to the first — “adjust this window a little.” Through enough layers, the message fades before it arrives. A whisper through forty people.

The 2015 fix is almost cheeky: give every layer a shortcut. Its input flows past it, unchanged, in addition to whatever the layer adds. The message no longer has to survive every layer — it has an express lane. Networks got 100+ layers deep overnight, and the shortcut design (ResNet) is still the default skeleton today.

LEARNING SIGNAL REACHING THE FIRST LAYER · NO SHORTCUTS
LEARNING SIGNAL REACHING THE FIRST LAYER · WITH SHORTCUTS
Illustrative decay curves. Drag the depth and watch the no-shortcut signal die while the shortcut signal survives by construction.
BLOCK 5 · CHECK YOURSELF

Why images first? Because you can watch it happen

Vision is where deep learning became visible: the features a network builds — edges, then parts, then objects — can be drawn, and that is why the program's first session starts here. The same stacking logic, run on words instead of pixels, is the transformer of the next page.

What does the walking window actually do?
What problem did the shortcut (skip connection) solve?
“Edges, then parts, then cat” describes…
TARA COMPANION · WEEK 1 · CNNS AND RESNETS · HAND-SET DATA, ILLUSTRATIVE · ARENA 3.0 PART 0.2