Home / Chroniques / Image recognition: the promise of a new approach to energy-efficient AI
Généré par l'IA / Generated using AI
π Digital π Science and technology

Image recognition: the promise of a new approach to energy-efficient AI

Jean-Gabriel Ganascia
Jean-Gabriel Ganascia
Philosopher and Professor Emeritus of Computer Science at the Faculty of Sciences at Sorbonne University
Avatar
Simon Thorpe
CNRS Research Director
Key takeaways
  • According to neuroscientist Simon Thorpe, the key to understanding how an image is constructed in the brain lies in the order in which the neurons in the retina send electrical signals to the brain.
  • As early as the 1980s, he developed an AI model for image recognition based on this principle.
  • This model was commercially deployed as early as 1999 to identify advertising logos at sporting events.
  • Although subsequently superseded by current models based on statistical science, this model is designed to be much more resource-efficient, which could now give it a competitive advantage.
  • Microsoft and the French company Mistral AI have released their own frugal AI models: Phi 3 and Mistral 7B.

To com­mu­nic­ate with oth­er cells, such as those in muscles and vari­ous glands, neur­ons emit elec­tric­al sig­nals. The amp­litude of each of these elec­tric­al sig­nals is approx­im­ately 100 mil­li­volts; a read­ing that remains con­stant regard­less of the neur­on or the inform­a­tion. How­ever, the neur­on will not emit any elec­tric­al sig­nal if the stim­u­lus it is exposed to is too weak. In oth­er words, it is the pres­ence or absence of an elec­tric­al sig­nal that con­sti­tutes inform­a­tion in itself, rather than the vari­ation in its amplitude.

“But elec­tric­al sig­nals are extremely rare,” says Simon Thorpe. “For me, what mat­ters in trans­mit­ting an image and its com­pos­i­tion from the ret­ina to the brain is, above all, the order in which the elec­tric­al sig­nals travel from the 100 mil­lion neur­ons in each ret­ina to brain cell.” The first elec­tric­al sig­nals will determ­ine the bright­est or highest-con­trast points, then the image will gradu­ally take shape, mov­ing from the bright­est to the dim­mest, until the stim­uli are too weak to trig­ger an elec­tric­al sig­nal from the neurons.

Image recognition

It was on this the­or­et­ic­al found­a­tion, developed as early as the 1980s by Simon Thorpe, that he built the basis for his own arti­fi­cial intel­li­gence mod­el for recog­nising objects with­in land­scapes. Indeed, the edges of an object with­in an image can be com­pared to abrupt changes in bright­ness. This makes it eas­ily iden­ti­fi­able by his mod­el. “And, if there is a face in this image, we can detect it, for example, by identi­fy­ing an upper edge – the top of the skull – a lower edge – the chin – and a middle edge – the mouth,” adds Simon Thorpe.

In 1999, Simon Thorpe co-foun­ded Spiken­et Tech­no­logy, an image-pro­cessing com­pany based on this tech­no­logy. One of its main aims was to identi­fy logos along the route of sport­ing events such as cyc­ling or For­mula 1, to help mar­ket­ing depart­ments optim­ise their ad place­ments. “It was a mar­ket that was grow­ing rap­idly at the time, both at phys­ic­al events and in digit­al spaces. Pro­fes­sion­als wanted to use this type of tech­no­logy to optim­ise the place­ment of their adverts, so that they would reach the audi­ence with s them feel­ing over­whelmed,” explains Jean-Gab­ri­el Ganas­cia, a pro­fess­or at the Sor­bonne and an expert on these issues.

Back to the future

From 2010 to 2017, a com­pet­i­tion well-known with­in the field, known as the ImageN­et Large Scale Visu­al Recog­ni­tion Chal­lenge (ILSVRC), brought togeth­er teams of research­ers from around the world, who sub­mit­ted their mod­els to clas­si­fy image data­bases. In 2012, a mod­el based on stat­ist­ic­al sci­ence won this com­pet­i­tion. This event helped to estab­lish this type of mod­el – which forms the basis of LLMs (Large Lan­guage Mod­els) – as the dom­in­ant approach on the glob­al stage. “Indeed, mod­els based on stat­ist­ic­al sci­ence were then con­sidered to have a lower mar­gin of error than those based on neur­os­cience,” explains Jean-Gab­ri­el Ganascia.

“For mod­els like mine, the momentum had passed,” says Simon Thorpe. “But – as recent events clearly show us – we are real­ising that these mod­els are very energy-intens­ive.” This is mainly due to the astro­nom­ic­al quant­it­ies of images they must pro­cess dur­ing the train­ing phase to learn to dis­tin­guish a sauce­pan from a drum. “But that’s not how human beings work. Oth­er­wise, 10 years of edu­ca­tion wouldn’t be enough to pass your A‑levels!” exclaims Simon Thorpe.

Simon Thorpe’s mod­el oper­ates in two suc­cess­ive phases: first, the iden­ti­fic­a­tion of an image based on the order of the dis­charges rather than their amp­litude, in order to con­struct the mean­ing of an image. Next, this image is stored in a sys­tem known as ‘silent neur­ons’. These are the digit­al equi­val­ents of inact­ive neur­al con­nec­tions in the brain that learn and then go into standby mode – to pre­vent the brain from becom­ing over­loaded. They can then be react­iv­ated when the image reappears before the eyes.

In fact, it fol­lows in the tra­di­tion of so-called ‘frugal’ AI mod­els, which seek to achieve their object­ives with the min­im­um of resources, rather than with max­im­um pro­cessing power. For example, the mul­tina­tion­al Microsoft and the French com­pany Mis­tral AI have also released their own frugal AI mod­els: Phi 3 and Mis­tral 7B. These mod­els train on lim­ited data­sets and are based on stat­ist­ic­al science.

Guénolé Boillot

Our world through the lens of science. Every week, in your inbox.

Get the newsletter