Back to results

Massachusetts Institute of Technology

Toward a more biologically plausible model of object recognition

Abstract

dc:description.abstract

Rapidly and reliably recognizing an object (is that a cat or a tiger?) is obviously an important skill for survival. However, it is a difficult computational problem, because the same object may appear differently under various conditions, while different objects may share similar features. A robust recognition system must have a capacity to distinguish between similar-looking objects, while being invariant to the appearance-altering transformation of an object. The fundamental challenge for any recognition system lies within this simultaneous requirement for both specificity and invariance. An emerging picture from decades of neuroscience research is that the cortex overcomes this challenge by gradually building up specificity and invariance with a hierarchical architecture. In this thesis, I present a computational model of object recognition with a feedforward and hierarchical architecture. The model quantitatively describes the anatomy, physiology, and the first few hundred milliseconds of visual information processing in the ventral pathway of the primate visual cortex. There are three major contributions. First, the two main operations in the model (Gaussian and maximum) have been cast into a more biologically plausible form, using monotonic nonlinearities and divisive normalization, and a possible canonical neural circuitry has been proposed. Second, shape tuning properties of visual area V4 have been explored using the corresponding layers in the model. It is demonstrated that the observed V4 selectivity for the shapes of intermediate complexity (gratings and contour features) can be explained by the combinations of orientation-selective inputs. Third, shape tuning properties in the higher visual area, inferior temporal (IT) cortex, have also been explored. It is demonstrated that the selectivity and invariance properties of IT neurons can be generated by the feedforward and hierarchical combinations of Gaussian-like and max-like operations, and their responses can support robust object recognition. Furthermore, experimentally-observed clutter effects and trade-off between selectivity and invariance in IT can also be observed and understood in this computational framework.

Degree

thesis:*
Department dc:contributor.department
Massachusetts Institute of Technology. Dept. of Physics.
Grantor dc:publisher
Massachusetts Institute of Technology
Year dc:date.issued
2007

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Kouh, Minjoon
Advisor dc:contributor.advisor
  • Tomaso Poggio and H. Sebastian Seung.

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission.
Language dc:language.iso
eng

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/1721.1/45156
OAI identifier oai:identifier
oai:dspace.mit.edu:1721.1/45156

Chain of custody

source
Harvested from
MIT
Base URL
dspace.mit.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
related terms
citation

Kouh, Minjoon. Toward a more biologically plausible model of object recognition. Massachusetts Institute of Technology, 2007. http://hdl.handle.net/1721.1/45156