DRFnet：用于目标检测和图像识别的动态感受野网络。

DRFnet: Dynamic receptive field network for object detection and image recognition.

作者信息

Tan Minjie, Yuan Xinyang, Liang Binbin, Han Songchen

机构信息

School of Aeronautics and Astronautics, Sichuan University, Chengdu, China.

出版信息

Front Neurorobot. 2023 Jan 10;16:1100697. doi: 10.3389/fnbot.2022.1100697. eCollection 2022.

DOI:10.3389/fnbot.2022.1100697

PMID:36704718

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC9871543/

Abstract

Biological experiments discovered that the receptive field of neurons in the primary visual cortex of an animal's visual system is dynamic and capable of being altered by the sensory context. However, in a typical convolution neural network (CNN), a unit's response only comes from a fixed receptive field, which is generally determined by the preset kernel size in each layer. In this work, we simulate the dynamic receptive field mechanism in the biological visual system (BVS) for application in object detection and image recognition. We proposed a Dynamic Receptive Field module (DRF), which can realize the global information-guided responses under the premise of a slight increase in parameters and computational cost. Specifically, we design a transformer-style DRF module, which defines the correlation coefficient between two feature points by their relative distance. For an input feature map, we first divide the relative distance corresponding to different receptive field regions between the target feature point and its surrounding feature points into N different discrete levels. Then, a vector containing N different weights is automatically learned from the dataset and assigned to each feature point, according to the calculated discrete level that this feature point belongs. In this way, we achieve a correlation matrix primarily measuring the relationship between the target feature point and its surrounding feature points. The DRF-processed responses of each feature point are computed by multiplying its corresponding correlation matrix with the input feature map, which computationally equals to accomplish a weighted sum of all feature points exploiting the global and long-range information as the weight. Finally, by superimposing the local responses calculated by a traditional convolution layer with DRF responses, our proposed approach can integrate the rich context information among neighbors and the long-range dependencies of background into the feature maps. With the proposed DRF module, we achieved significant performance improvement on four benchmark datasets for both tasks of object detection and image recognition. Furthermore, we also proposed a new matching strategy that can improve the detection results of small targets compared with the traditional IOU-max matching strategy.

摘要

生物学实验发现，动物视觉系统初级视觉皮层中神经元的感受野是动态的，并且能够被感觉环境改变。然而，在典型的卷积神经网络（CNN）中，一个单元的响应仅来自固定的感受野，该感受野通常由每层中预设的内核大小决定。在这项工作中，我们模拟生物视觉系统（BVS）中的动态感受野机制以应用于目标检测和图像识别。我们提出了一种动态感受野模块（DRF），它能够在参数和计算成本略有增加的前提下实现全局信息引导的响应。具体来说，我们设计了一种Transformer风格的DRF模块，它通过两个特征点之间的相对距离定义它们的相关系数。对于输入特征图，我们首先将目标特征点与其周围特征点之间不同感受野区域对应的相对距离划分为N个不同的离散级别。然后，从数据集中自动学习一个包含N个不同权重的向量，并根据该特征点所属的计算出的离散级别将其分配给每个特征点。通过这种方式，我们得到了一个主要测量目标特征点与其周围特征点之间关系的相关矩阵。每个特征点经过DRF处理后的响应是通过将其对应的相关矩阵与输入特征图相乘来计算的，这在计算上等同于利用全局和远距离信息作为权重对所有特征点进行加权求和。最后，通过将传统卷积层计算的局部响应与DRF响应叠加，我们提出的方法可以将邻居之间丰富的上下文信息和背景的远距离依赖整合到特征图中。通过所提出的DRF模块，我们在用于目标检测和图像识别这两项任务的四个基准数据集上实现了显著的性能提升。此外，我们还提出了一种新的匹配策略，与传统的IOU-max匹配策略相比，该策略可以改善小目标的检测结果。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/87aa/9871543/4ec970519b8b/fnbot-16-1100697-g0001.jpg

相似文献

DRFnet: Dynamic receptive field network for object detection and image recognition.

Front Neurorobot. 2023 Jan 10;16:1100697. doi: 10.3389/fnbot.2022.1100697. eCollection 2022.

TwinsReID: Person re-identification based on twins transformer's multi-level features.

Math Biosci Eng. 2023 Jan;20(2):2110-2130. doi: 10.3934/mbe.2023098. Epub 2022 Nov 14.

Research on Object Detection of PCB Assembly Scene Based on Effective Receptive Field Anchor Allocation.

Comput Intell Neurosci. 2022 Feb 14;2022:7536711. doi: 10.1155/2022/7536711. eCollection 2022.

A Mixed Visual Encoding Model Based on the Larger-Scale Receptive Field for Human Brain Activity.

Brain Sci. 2022 Nov 29;12(12):1633. doi: 10.3390/brainsci12121633.

Dynamic Serpentine Convolution with Attention Mechanism Enhancement for Beef Cattle Behavior Recognition.

Animals (Basel). 2024 Jan 31;14(3):466. doi: 10.3390/ani14030466.

A novel low light object detection method based on the YOLOv5 fusion feature enhancement.

Sci Rep. 2024 Feb 23;14(1):4486. doi: 10.1038/s41598-024-54428-8.

Multi-scale object detection in UAV images based on adaptive feature fusion.

PLoS One. 2024 Mar 27;19(3):e0300120. doi: 10.1371/journal.pone.0300120. eCollection 2024.

An Efficient Image Deblurring Network with a Hybrid Architecture.

Sensors (Basel). 2023 Aug 18;23(16):7260. doi: 10.3390/s23167260.

Learning Nonclassical Receptive Field Modulation for Contour Detection.

IEEE Trans Image Process. 2019 Sep 16. doi: 10.1109/TIP.2019.2940690.

A continuation method for image registration based on dynamic adaptive kernel.

Neural Netw. 2023 Aug;165:774-785. doi: 10.1016/j.neunet.2023.06.025. Epub 2023 Jun 30.

引用本文的文献

Building Segmentation in Urban and Rural Areas with MFA-Net: A Multidimensional Feature Adjustment Approach.

Sensors (Basel). 2025 Apr 19;25(8):2589. doi: 10.3390/s25082589.

本文引用的文献

Contextual Transformer Networks for Visual Recognition.

IEEE Trans Pattern Anal Mach Intell. 2023 Feb;45(2):1489-1500. doi: 10.1109/TPAMI.2022.3164083. Epub 2023 Jan 6.

Circuits and Mechanisms for Surround Modulation in Visual Cortex.

Annu Rev Neurosci. 2017 Jul 25;40:425-451. doi: 10.1146/annurev-neuro-072116-031418. Epub 2017 May 3.

Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

IEEE Trans Pattern Anal Mach Intell. 2017 Jun;39(6):1137-1149. doi: 10.1109/TPAMI.2016.2577031. Epub 2016 Jun 6.

Contrast-dependent variations in the excitatory classical receptive field and suppressive nonclassical receptive field of cat primary visual cortex.

Cereb Cortex. 2013 Feb;23(2):283-92. doi: 10.1093/cercor/bhs012. Epub 2012 Feb 2.

The "silent" surround of V1 receptive fields: theory and experiments.

J Physiol Paris. 2003 Jul-Nov;97(4-6):453-74. doi: 10.1016/j.jphysparis.2004.01.023.

Receptive fields, binocular interaction and functional architecture in the cat's visual cortex.

J Physiol. 1962 Jan;160(1):106-54. doi: 10.1113/jphysiol.1962.sp006837.

Discharge patterns and functional organization of mammalian retina.

J Neurophysiol. 1953 Jan;16(1):37-68. doi: 10.1152/jn.1953.16.1.37.

Nature and interaction of signals from the receptive field center and surround in macaque V1 neurons.

J Neurophysiol. 2002 Nov;88(5):2530-46. doi: 10.1152/jn.00692.2001.

文献AI研究员

20分钟写一篇综述，助力文献阅读效率提升50倍。

立即体验

用中文搜PubMed

大模型驱动的PubMed中文搜索引擎

马上搜索

文档翻译

学术文献翻译模型，支持多种主流文档格式。

立即体验

DRFnet：用于目标检测和图像识别的动态感受野网络。

DRFnet: Dynamic receptive field network for object detection and image recognition.

作者信息

机构信息

出版信息

相似文献

引用本文的文献

本文引用的文献

文献AI研究员

用中文搜PubMed

文档翻译

Suppr 超能文献

相似文献

引用本文的文献

本文引用的文献