“Artificial intelligence learned to see because human beings first showed it the world.”
The Pioneers Who Gave Artificial Intelligence Its Eyes
For most of human history, sight appeared effortless.
A child looks at a photograph and sees a dog. A face is recognised across a crowded room. A tree remains a tree whether photographed in sunshine, rain or snow. We distinguish a bicycle from a motorcycle, a smile from a frown and a street from a field almost without conscious thought.
For a computer, none of this was obvious.
An image entering a computer is not initially a face, a dog, a mountain or Freddie Mercury standing before a stadium crowd. At its most fundamental level, it is numerical information: pixels, intensities, colours and patterns.
Teaching machines to extract meaning from that information has therefore been one of the extraordinary scientific journeys of the computer age.
And that journey did not begin with ChatGPT or today’s generative artificial intelligence. Its roots stretch back through decades of mathematics, engineering, psychology and computer science, and through the work of people whose names deserve to be remembered.
Nasir Ahmed and the Mathematics Behind the Picture
One of them was the engineer Nasir Ahmed.
In the early 1970s, Ahmed conceived what became known as the Discrete Cosine Transform, or DCT. Working with T. Natarajan and K. R. Rao, he published the landmark work in 1974.
The mathematics did something immensely important. Rather than treating an image merely as a huge collection of individual pixels, the DCT could represent blocks of image information in terms of different spatial frequencies.
Natural pictures contain enormous amounts of redundancy. Neighbouring pixels frequently resemble one another. Large areas may change gradually rather than chaotically. The DCT helped concentrate much of the significant visual information into relatively few coefficients, allowing less perceptually important information to be represented less precisely.
That may sound like an obscure mathematical achievement. Its consequences were anything but obscure.
The DCT became fundamental to JPEG compression and enormously influential in subsequent digital image and video technology. Every generation that has casually photographed a birthday, emailed a picture, browsed photographs on the Web or watched compressed digital video has inhabited a technological landscape that Ahmed and his collaborators helped create.
Before computers could learn to understand the world’s pictures, those pictures first had to become practical things for computers to store, manipulate and transmit.
Ahmed helped make that possible.
Fei-Fei Li and a Radical Idea
Decades later another problem confronted computer scientists.
Computers could store millions of photographs. But how could they learn what was actually in them?
Professor Fei-Fei Li came to believe that part of the answer was scale.
Human beings do not learn what a dog is by seeing one carefully selected photograph of a Labrador sitting obediently against a white background. We encounter dogs large and small, running and sleeping, partially obscured, photographed from strange angles and seen under different lighting.
The visual world is gloriously untidy.
If machines were going to recognise it, Li reasoned, they needed experience of that diversity.
The result was ImageNet.
Li and her collaborators set out to construct an immense organised collection of labelled photographs. It became a colossal human undertaking. Stanford has described how nearly 50,000 workers from 167 countries participated over two and a half years in cleaning, sorting and labelling imagery drawn from an enormous pool of candidate pictures. ImageNet eventually encompassed millions of images organised across thousands of categories.
Behind every label was an apparently simple statement.
This is a cat.
This is a chair.
This is a strawberry.
This is a car.
Yet collectively those labels helped change artificial intelligence.
The Moment the Machine Began Learning
ImageNet provided something researchers desperately needed: an enormous visual classroom.
Then, in 2012, came one of the pivotal moments in modern artificial intelligence.
Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton used a deep convolutional neural network, subsequently famous as AlexNet, in the ImageNet competition. Its performance represented a dramatic improvement over competing approaches and demonstrated the extraordinary potential of combining neural networks, large datasets and increasingly powerful computing hardware.
Computer vision accelerated.
Instead of engineers attempting to specify every visual rule themselves, increasingly sophisticated neural networks could learn useful representations from examples.
Layers within these networks could progressively respond to visual structures, from comparatively elementary patterns towards increasingly complex features useful for identifying objects.
The machine was no longer simply being handed a gigantic book of rigid visual instructions.
It was learning from experience.
From Recognition to Understanding
What followed has been extraordinary.
Computer vision systems can now identify objects, locate them within scenes, separate one object from another, recognise text, analyse medical imagery, help robots navigate their surroundings and assist people with visual impairments.
Modern multimodal artificial intelligence has pushed the idea further still.
Show a sufficiently capable system a photograph today and it may do considerably more than answer “dog”.
It can potentially describe the dog’s surroundings, distinguish the animal from nearby objects, read a sign behind it, discuss the composition of the photograph and reason about relationships among things within the scene.
That distinction is important.
Recognition asks: What is this?
Visual understanding asks: What is happening here, how do these things relate to one another, and what can reasonably be inferred from what is visible?
We should nevertheless resist seductive statistics suggesting that artificial vision has become universally “99.99 per cent accurate”.
There is no single meaningful accuracy figure for machine vision.
On narrowly defined and carefully controlled recognition tasks, modern systems can achieve extraordinarily high accuracy. The untidy real world is considerably harder. Unusual lighting, obscured objects, unfamiliar situations, manipulated photographs, cultural context and ambiguous scenes can still deceive machines.
Even extremely sophisticated artificial intelligence can be confidently wrong.
That makes human judgement no less important today than it was when the first datasets were being labelled.
The Humans Behind the Machine
Perhaps that is the most beautiful part of this history.
Artificial intelligence learned to see because human beings first showed it the world.
Mathematicians found ways of representing images. Engineers found ways of compressing them. Computer scientists designed algorithms capable of learning from them. Semiconductor engineers built processors powerful enough to perform the calculations. And tens of thousands of ordinary people sat at computers attaching humble words to pictures.

Cat.
House.
Tree.
Car.
Person.
Those little acts of classification became part of something enormous.
Nasir Ahmed helped establish mathematics that allowed digital imagery to be represented and compressed efficiently. Fei-Fei Li and her collaborators helped provide machines with a visual encyclopedia from which a new generation of algorithms could learn. Krizhevsky, Sutskever and Hinton demonstrated how deep neural networks could exploit such enormous datasets with remarkable effectiveness.
None of them alone “invented computer vision”.
Science almost never works that way.
Instead, each generation laid another piece of track, frequently without knowing quite where the railway would eventually lead.
And Now the Machine Looks Back
Today we can place a century-old photograph before an artificial intelligence system and ask what it sees.
We can show it a crowded street, a painting, a diagram, a planet photographed from space or a singer standing at a piano before tens of thousands of people.
Increasingly, the machine can tell us not merely that pixels are present, but that those pixels form people, buildings, expressions, words and relationships.
It remains imperfect. Sometimes spectacularly so.
But consider the distance travelled.
From Nasir Ahmed’s cosine mathematics, through JPEG and the digital-image revolution, to Fei-Fei Li’s millions of labelled photographs, ImageNet, AlexNet and today’s multimodal artificial intelligence, we have spent half a century teaching silicon something that a young child begins doing almost instinctively.
We have been teaching the machine to see.
And behind those artificial eyes remain countless human ones: the scientists, engineers, mathematicians, programmers and anonymous image labellers who looked first.
Artificial intelligence may now examine the picture.
But humanity showed it where to look.
