“INT/CONECT workshop : All models are wrong… but some are useful! when vision meets audition to reinvent neuroscience”
Audition and vision are two fields of research that have always faced one another. From psychophysics to cognitive neuroscience, the discoveries and models of one have often preceded those of the other, without it ever being entirely clear which came first, the chicken or the egg? This one-day workshop will focus on three different aspects of vision and auditition: the perception of textures, the perception of voices and faces, and the hierarchical (or non-hierarchical) nature of vision and audition. We will discuss how we can model both, what are common, what are different, what is better understood by one and less by the other, and in the end how insights from one can inform the other to advance our understanding of both.
Models are omnipresent in science, yet they are never more than imperfect analogies of natural phenomena. Models reduce, divide, simplify, or complexify a reality that is not always observable or measurable. Food for thought, models sometimes become theory, but they must never become dogma, lest they mislead us. In this sense, as George Box reminded us, “all models are wrong”! True—but the very nature of this falsifiability is that it helps us think more clearly about the world and better predict the phenomena that make it up. Whether biological, physical, ecological, or geological, models allow us to simulate, predict, and formulate new hypotheses, thereby pushing back the boundaries of knowledge. The history of science has been nourished by these constant back-and-forth movements between models and observations, and it is in this sense that, although they are all wrong, they are indeed all useful. In this lecture series, we wish to highlight modeling as a central aspect of the scientific approach. For this first edition, we wanted to bring into dialogue research on the modeling of vision and audition to foster an inter-disciplinary research program.
10h30 | Welcome coffee
11h00 | Introduction | Etienne Thoret
Moderation: Guillaume Masson
12h30 | Lunch break
Moderation: Anna Montagnini
15h05 | Coffee break
Moderation: Etienne Thoret
Moderation: Laurent Perrinet
17h50 | Closing words
18h00 | Refreshments
This workshop is open to researchers, students, and anyone interested in computational and cognitive neuroscience. No registration fee is required, but please contact the organizers if you wish to attend by filling this form: https://framaforms.org/intconect-workshop-all-models-are-wrong-but-some-are-useful-1781615391 For further inquiries, please contact the organizerts by e-mail.
This workshop has received support from the French government under the Programme « Investissements d’Avenir », Initiative d’Excellence d’Aix-Marseille Université via A*Midex funding (AMX-19-IET-004), and ANR (ANR-17-EURE-0029).
Segmentation is the process of grouping image features to form perceptual objects and segmenting those objects from each other. Visual cues related to image segmentation influence single neuron firing, as shown in primary visual cortex (V1) neurons. However, segment–specific effects on population V1 activity remain unknown. I will present a normative theory that build on existing frameworks to explain population firing through probabilistic inference about latent causes. My lab has recently proposed a probabilistic model of human segment perception, in which the inferences of image features and of image segments are coupled and proceeds iteratively (Biswas et al., 2026, bioRxiv). We have now extended the model to generate predictions for neural activity in visual cortical neurons. The model qualitatively explains single–neuron modulation by segments and predicts how image segments modulate pairwise covariability and population-level coordination. I will describe those predictions and provide initial empirical support with data from areas V1 and V4 of one awake macaque. This work offers a new theory of segment–based image encoding in visual cortex that predicts highly flexible, segment–dependent neuronal interactions. This is an important first step towards understanding the distributed neural population representation of natural scenes.
Little is known about how neural representations of speech differ across species, or how they are shaped by developmental exposure. To isolate sensitivity for higher-order structure, we synthesized sounds whose statistics are matched to natural sounds under a spectrotemporal model; because they are otherwise unconstrained, these synthetics lack the salient higher-order structure of speech. Ferrets distinguished natural from synthetic sounds behaviorally, through sound-evoked facial motion and pupil dynamics. Using functional ultrasound imaging, we then measured ferret auditory cortical responses to the same sounds previously tested in humans. Ferrets showed frequency and modulation tuning similar to humans, but whereas humans respond substantially more to natural than synthetic speech in non-primary regions, ferret responses were closely matched throughout primary and non-primary cortex, even for conspecific vocalizations. Finally, we ask whether higher-order sensitivity can be induced by experience: preliminary evidence shows that rearing ferret kits with early speech exposure paired with positive-valence experiences (feeding, water, play) enhances representations of natural over matched-synthetic sounds in non-primary ventral fields, scaling with exposure duration. Together, these results suggest that higher-order cortical sensitivity to speech can be shaped by early, behaviorally relevant experience.
How should we infer perceptual scales from behavioral judgments? Classical Thurstonian models explain discrimination through the geometry of an internal sensory scale: stimuli are encoded with noise, and choices reveal distances along that scale. Bayesian observer models add another ingredient: the brain may decode noisy measurements using prior expectations about the stimulus world. In this talk, I will compare these two views in a common probabilistic framework, using visual tasks such as two-alternative forced choice and maximum likelihood difference scaling. The goal is to clarify what each model assumes, what it predicts, and what can actually be inferred from behavioral data. I will show how random stimulus variability helps separate geometry from Bayesian decoding, and why this distinction matters for estimating perceptual representations.
Our ability to rapidly encode new auditory input is essential to identify meaningful sounds in complex acoustic environments, yet the neural mechanisms that support the formation of such auditory memories remain poorly understood. By combining human psychophysics and mouse neurophysiology, the presented body of works will examine how rapid auditory learning shapes cortical representations of complex sounds. Using a behavioural paradigm designed to mimic acoustic scenes, we presented listeners with a series of random acoustic patterns in which a specific sound pattern re-occurred intermittently across trials. Human listeners showed rapid and implicit learning of these recurring sounds, improving performance on an auditory task that did not require explicit detection of the recurrences, demonstrating that robust auditory memory can form after only a few exposures. To investigate the cortical basis of this process, we performed in vivo two-photon calcium imaging in the auditory cortex of passively listening, awake mice. These recordings revealed neural adaptation of population responses to recurring complex sounds, suggesting the emergence of sparse cortical representations for recurring sound patterns. To directly examine how functionally interconnected neurons coordinate during sound processing, we combined in vivo two-photon imaging with holographic optogenetic stimulation of a subset of co-tuned neurons. This approach revealed that functionally connected neurons dynamically adjust their activity to rebalance network-level responses during target sound processing, and that these ensembles are widely distributed across auditory cortex rather than tightly localised. Together, these findings suggest that effective auditory perception emerges from sparse cortical representations and is shaped by adaptive interactions across distributed neuronal networks.
Abilities for face recognition vary greatly even among neurologically typical individuals. At one end of the spectrum, developmental prosopagnosics show great difficulty recognizing faces, despite not having sustained any brain injuries. At the other end of the spectrum, super-recognizers easily recognize faces they have not seen in years, even if these faces have physically changed in a substantial manner. We recently characterised the brain computations of participants of various face recognition abilities using high-density electroencephalographic signals and a combination of behavioural tests, artificial neural network models, and machine learning analyses. We found that individual face recognition ability can be decoded from brain activity in an extended temporal interval for face and non-face objects. We show that both visual and semantic brain computations contribute to these individual differences. Understanding how perceptual mechanisms are linked with individual abilities can offer important and straightforward insights for improving face processing in both individuals with poor face recognition abilities and people whose jobs require strong face processing ability.
Voices are information rich “auditory faces” that our brain – and that of non-human primates – is expert at decoding. Functional MRI studies over two decades have provided evidence of areas of auditory cortex selectively activated by voices - the ‘temporal voice areas” (TVAs ) – analogous to the Face selective areas of visual cortex. More recently, homologous areas have been observed in the brain of macaques and marmosets, suggesting an evolutionary conserved ‘voice patch system’. How neuronal activity in these voice patches encode and transform voice information remains largely unexplored. It can be modelled using acoustical, theoretical, bio-inspired or AI-generated models. I will present examples drawn from primate fMRi and macaque fMRi-guided electrophysiology. This research sheds crucial light on the evolutionary foundations of human voice perception, informing future brain–machine interface technologies, including cortical implants designed to restore or enhance speech and voice perception.
Neurostimulation is the main strategy for treating profound deafness1. However, current devices, including cochlear or auditory brainstem implants, have a limited information throughput, which impacts restored perception and quality of life. These limits are due to the low number of electrical contacts providing independent information in the small target structures. An alternative strategy is to target the auditory cortex, a structure large enough to place hundreds to thousands of electrodes, and whose stimulation produces sound perceptions. However, efficient encoding algorithms have been missing to unlock the potential of cortical stimulation. Here, we introduce Braincodec, a deep neural network architecture that optimizes stimulation patterns of an auditory cortical implant while minimizing information loss, aligning with the spatial, temporal, and high-level properties of the auditory cortex neural code. Braincodec operates online and requires only a few hundred electrodes to transmit the level of acoustic details experienced in natural hearing which are inaccessible to cochlear implants. Moreover, when tested in mice with an auditory cortical implant prototype, Braincodec produced precise perceptions. Therefore, Braincodec enables the design of cortical implants with a high information throughput, overcoming the limitations of cochlear implants, with a generic framework applicable to other sensory modalities.
Backpropagation is the core learning mechanism underlying deep learning. However, whether and how this algorithm is implemented in the brain remains highly debated. In particular, while forward activations of pretrained models reliably map onto the cortical hierarchy of visual processing, it is unknown whether backpropagated gradients exhibit a similar correspondence. Here, we address this question using functional magnetic resonance imaging (fMRI) and magnetoencephalography (MEG) recordings of human brain responses to natural images. For this, we extend standard encoding analyses of forward activations to map backpropagated gradients onto neural data. Focusing on a recent self-supervised vision model (DINOv3) and reproducing results on eight vision models, we find that backpropagated gradients can reliably predict both fMRI and MEG signals, specifically in higher-level visual cortex and for later latencies. However, the spatial and temporal organization of these backpropagated gradients in the brain diverges from the patterns expected under a biologically plausible backpropagation mechanism: specifically, both the order in which gradients are computed and their spatial organization diverge from the temporal and spatial hierarchies of the human brain. Together, these results suggest that, although deep networks and the brain may share similar representational content, they likely rely on fundamentally different mechanisms to learn those representations.
Transforming natural sounds into a knowledge of the sound-generating objects and events in the environment requires converting acoustic signals into semantic representations. A key goal of research on the perceptual and cerebral representation of natural sounds is then to pinpoint the chain of sound-signal transformations involved in this process. Recent research has addressed this problem within a model-comparison framework, showing that computational modelling of behavioural and fMRI responses to natural sounds highlights the dominance of intermediate acoustic-semantic representations. Extensive model comparison approaches have however yet to be applied to time-resolved cerebral responses to natural sounds, leaving the temporal dynamics of this process insufficiently characterized. In this study, we investigated this issue by conducting a model-comparison analysis of magnetoencephalography (MEG) responses to a heterogeneous set of 2-second natural sounds. We evaluated acoustic models, text-based semantic models, and deep neural networks (DNNs) that learn semantics from sounds, to determine how well they capture time-resolved brain responses. Cross-validated linear regression and variance partitioning analyses were used to quantify and dissociate the contribution of different models to cerebral representations. In line with current research, our results show that DNNs outperform all model classes and, more importantly, capture all of the cerebrally predictive variance for text-based semantic models and most of all the cerebrally predictive acoustic variance. These results strongly suggest that the semantic representation in the brain is directly learned from the sound, and that linguistic only is insufficient. Finally, our findings reveal a temporal hierarchy of auditory representations, progressing from acoustic processing around 100 ms, to acoustic and sound-driven categorical semantic representations around 250 ms, and finally to sound-driven continuous semantic representations around 500 ms, followed by an acoustic rebound at sound offset. Although all models are necessarily simplified approximations, they remain valuable tools for investigating the complexity of human brain responses underlying auditory perception.