The representational hierarchy in human and artificial visual systems in the presence of object-scene regularities

PLoS Comput Biol. 2023 Apr 28;19(4):e1011086. doi: 10.1371/journal.pcbi.1011086. eCollection 2023 Apr.

Abstract

Human vision is still largely unexplained. Computer vision made impressive progress on this front, but it is still unclear to which extent artificial neural networks approximate human object vision at the behavioral and neural levels. Here, we investigated whether machine object vision mimics the representational hierarchy of human object vision with an experimental design that allows testing within-domain representations for animals and scenes, as well as across-domain representations reflecting their real-world contextual regularities such as animal-scene pairs that often co-occur in the visual environment. We found that DCNNs trained in object recognition acquire representations, in their late processing stage, that closely capture human conceptual judgements about the co-occurrence of animals and their typical scenes. Likewise, the DCNNs representational hierarchy shows surprising similarities with the representational transformations emerging in domain-specific ventrotemporal areas up to domain-general frontoparietal areas. Despite these remarkable similarities, the underlying information processing differs. The ability of neural networks to learn a human-like high-level conceptual representation of object-scene co-occurrence depends upon the amount of object-scene co-occurrence present in the image set thus highlighting the fundamental role of training history. Further, although mid/high-level DCNN layers represent the category division for animals and scenes as observed in VTC, its information content shows reduced domain-specific representational richness. To conclude, by testing within- and between-domain selectivity while manipulating contextual regularities we reveal unknown similarities and differences in the information processing strategies employed by human and artificial visual systems.

Publication types

  • Research Support, Non-U.S. Gov't

MeSH terms

  • Brain Mapping
  • Humans
  • Magnetic Resonance Imaging
  • Pattern Recognition, Visual*
  • Photic Stimulation
  • Visual Cortex*
  • Visual Perception

Grants and funding

S.B. was funded by FWO (Fonds Wetenschappelijk Onderzoek) through a postdoctoral fellowship (12S1317N) and a Research Grant (1505518N). J.M. was supported by the Ad futura Scholarship of the Public Scholarship, Development, Disability and Maintenance Fund of the Republic of Slovenia. H.O. B. was supported by the KU Leuven Research Council (C14/21/047; AKUL/19/05) and FWO grants EOS HumVisCat (nr. 30991544) and G0D3322N.