物体识别技术发展状况( 三 )



更多阅读
【物体识别技术发展状况】 放养的深度学习-浅谈自编码器

■网友
物体识别是个难题,目前大家没有达成共识,没有公认的理论可以清晰地解释人脑的物体识别能力,既然物体识别的奥秘没有攻克,当然也没有一种机制可以达到和人脑一样级别的物体识别水平。这里贴一篇关于物体识别论文的核心部分:作者:MIT Dicarlo实验室的掌门人James J. DiCarlo,2012年发表于Cell-Neuron上的《How Does the Brain SolveVisual Object Recognition?》Our hypothesis is that each ventral stream cortical subpopulation uses at least three common, genetically encoded mechanisms (described below) to carry out that meta job description and that together, those mechanisms direct it to ‘‘choose’’ a set of input weights, a normalization pool, and a static nonlinearity that lead to improved subspace untangling. Specifically, we postulate the existence of the following three key conceptual mechanisms:(1) Each subpopulation sets up architectural nonlinearities that naturally tend to flatten object manifolds. Specifically, even with random (nonlearned) filter weights, NLN-like models tend to produce easier-to-decode object identity manifolds largely on the strength of the normalization operation (Jarrett et al., 2009; Lewicki and Sejnowski, 2000; Olshausen and Field, 2005; Pinto et al., 2008b), similar in spirit to the overcomplete approach of V1 (described above).(2) Each subpopulation embeds mechanisms that tune the synaptic weights to concentrate its dynamic response range to span regions of its input space where images are typically found (e.g., do not bother encoding things you never see). This is the basis of natural image statistics and compression (e.g., Hoyer and Hyva ?rinen, 2002; Olshausen and Field, 1996; Simoncelli and Olshausen, 2001) and its importance is supported by the observation that higher levels of the ventral stream are more tuned to natural feature conjunctions than lower levels (e.g., Rust and DiCarlo, 2010).(3) Each subpopulation uses an unsupervised algorithm to tune its parameters such that input patterns that occur close together in time tend to lead to similar output responses. This implements the theoretical idea that naturally occurring temporal contiguity cues can ‘‘instruct’’ the building of tolerance to identity-preserving transformations. More specifically, because each object’s identity is temporally stable, different retinal images of the same object tend to be temporally contig- uous (Fazl et al., 2009; Foldiak, 1991; Stryker, 1992; Wallis and Rolls, 1997; Wiskott and Sejnowski, 2002). In the geometrical, population-based description presented in Figure 2, response vectors that are produced by retinal images occurring close together in time tend to be the directions in the population response space that correspond to identity-preserving image variation, and thus attempts to produce similar neural responses for temporally contiguous stimuli achieve the larger goal of factorizing object identity and other object variables (position, scale, pose, etc.). For example, the ability of IT neurons to respond similarly to the same object seen at different retinal positions (‘‘position tolerance’’) could be bootstrapped by the large number of saccadic-driven image translation experiences that are spontaneously produced on the retinae ($100 million such translation experiences per year of life). Indeed, artificial manipula- tions of temporally contiguous experience with object images across different positions and sizes can rapidly and strongly reshape the position and size tolerance of IT neurons—destroying existing tolerance and building new tolerance, depending on the provided visual experi- ence statistics (Li and DiCarlo, 2008, 2010), and predict- ably modifying object perception (Cox et al., 2005). We refer the reader to computational work on how such learning might explain properties of the ventral stream (e.g., Foldiak, 1991; Hurri and Hyva ? rinen, 2003; Wiskott and Sejnowski, 2002; see section 4), as well as other potentially important types of unsupervised learning that do not require temporal cues (Karlinsky et al., 2008; Perry et al., 2010).


推荐阅读