To a computer, a picture is a grid of numbers
You see a cat. The computer sees a grid of brightness values — that is the entire input. Every image task starts here: numbers, arranged in a grid, where each cell is called a pixel.
Flip the switch, then hover the cells.
A small window asks one question, everywhere
Nobody looks at a whole picture at once. The first thing a vision network does is send a tiny window walking across the grid. The window carries one pattern — say, a vertical edge, dark on the left and light on the right — and at every position it asks: how much is my pattern here?
Step the window across the cat. The score column on the right fills in as it goes.
Stack the layers and features build up
Here is the payoff of all those little windows. Layer one finds edges. Layer two finds patterns made of edges — corners, curves, textures. Layer three finds patterns made of those: ears, eyes, whiskers. Layer four doesn't see parts anymore. It sees cat. Nobody programmed any of the levels; they emerge from stacking.
Slide through the depth and watch the description change.
Deep should be better — so why did deep networks refuse to train?
By the stack logic, more layers should mean smarter features. For years it didn't work: past a certain depth, networks trained worse, not better. The plain-words version of why: learning is a message passed backwards from the last layer to the first — “adjust this window a little.” Through enough layers, the message fades before it arrives. A whisper through forty people.
The 2015 fix is almost cheeky: give every layer a shortcut. Its input flows past it, unchanged, in addition to whatever the layer adds. The message no longer has to survive every layer — it has an express lane. Networks got 100+ layers deep overnight, and the shortcut design (ResNet) is still the default skeleton today.
Why images first? Because you can watch it happen
Vision is where deep learning became visible: the features a network builds — edges, then parts, then objects — can be drawn, and that is why the program's first session starts here. The same stacking logic, run on words instead of pixels, is the transformer of the next page.