Image recognition: the promise of a new approach to energy-efficient AI
- According to neuroscientist Simon Thorpe, the key to understanding how an image is constructed in the brain lies in the order in which the neurons in the retina send electrical signals to the brain.
- As early as the 1980s, he developed an AI model for image recognition based on this principle.
- This model was commercially deployed as early as 1999 to identify advertising logos at sporting events.
- Although subsequently superseded by current models based on statistical science, this model is designed to be much more resource-efficient, which could now give it a competitive advantage.
- Microsoft and the French company Mistral AI have released their own frugal AI models: Phi 3 and Mistral 7B.
To communicate with other cells, such as those in muscles and various glands, neurons emit electrical signals. The amplitude of each of these electrical signals is approximately 100 millivolts; a reading that remains constant regardless of the neuron or the information. However, the neuron will not emit any electrical signal if the stimulus it is exposed to is too weak. In other words, it is the presence or absence of an electrical signal that constitutes information in itself, rather than the variation in its amplitude.
“But electrical signals are extremely rare,” says Simon Thorpe. “For me, what matters in transmitting an image and its composition from the retina to the brain is, above all, the order in which the electrical signals travel from the 100 million neurons in each retina to brain cell.” The first electrical signals will determine the brightest or highest-contrast points, then the image will gradually take shape, moving from the brightest to the dimmest, until the stimuli are too weak to trigger an electrical signal from the neurons.
Image recognition
It was on this theoretical foundation, developed as early as the 1980s by Simon Thorpe, that he built the basis for his own artificial intelligence model for recognising objects within landscapes. Indeed, the edges of an object within an image can be compared to abrupt changes in brightness. This makes it easily identifiable by his model. “And, if there is a face in this image, we can detect it, for example, by identifying an upper edge – the top of the skull – a lower edge – the chin – and a middle edge – the mouth,” adds Simon Thorpe.
In 1999, Simon Thorpe co-founded Spikenet Technology, an image-processing company based on this technology. One of its main aims was to identify logos along the route of sporting events such as cycling or Formula 1, to help marketing departments optimise their ad placements. “It was a market that was growing rapidly at the time, both at physical events and in digital spaces. Professionals wanted to use this type of technology to optimise the placement of their adverts, so that they would reach the audience with s them feeling overwhelmed,” explains Jean-Gabriel Ganascia, a professor at the Sorbonne and an expert on these issues.
Back to the future
From 2010 to 2017, a competition well-known within the field, known as the ImageNet Large Scale Visual Recognition Challenge (ILSVRC), brought together teams of researchers from around the world, who submitted their models to classify image databases. In 2012, a model based on statistical science won this competition. This event helped to establish this type of model – which forms the basis of LLMs (Large Language Models) – as the dominant approach on the global stage. “Indeed, models based on statistical science were then considered to have a lower margin of error than those based on neuroscience,” explains Jean-Gabriel Ganascia.
“For models like mine, the momentum had passed,” says Simon Thorpe. “But – as recent events clearly show us – we are realising that these models are very energy-intensive.” This is mainly due to the astronomical quantities of images they must process during the training phase to learn to distinguish a saucepan from a drum. “But that’s not how human beings work. Otherwise, 10 years of education wouldn’t be enough to pass your A‑levels!” exclaims Simon Thorpe.
Simon Thorpe’s model operates in two successive phases: first, the identification of an image based on the order of the discharges rather than their amplitude, in order to construct the meaning of an image. Next, this image is stored in a system known as ‘silent neurons’. These are the digital equivalents of inactive neural connections in the brain that learn and then go into standby mode – to prevent the brain from becoming overloaded. They can then be reactivated when the image reappears before the eyes.
In fact, it follows in the tradition of so-called ‘frugal’ AI models, which seek to achieve their objectives with the minimum of resources, rather than with maximum processing power. For example, the multinational Microsoft and the French company Mistral AI have also released their own frugal AI models: Phi 3 and Mistral 7B. These models train on limited datasets and are based on statistical science.

