基于改进DBNet-CRNN的声级计读数识别方法

汪嘉,祝海江,王寅初,何龙标,杨平,牛锋

计量学报 ›› 2024, Vol. 45 ›› Issue (8) : 1191-1199.

PDF(729 KB)
PDF(729 KB)
计量学报 ›› 2024, Vol. 45 ›› Issue (8) : 1191-1199. DOI: 10.3969/j.issn.1000-1158.2024.08.15
声学计量

基于改进DBNet-CRNN的声级计读数识别方法

  • 汪嘉1,祝海江1,王寅初2,何龙标2,杨平2,牛锋2
作者信息 +

Sound Level Meter Reading Recognition Method Based on Improved DBNet-CRNN

  • WANG Jia1,ZHU Haijiang1,WANG Yinchu2,HE Longbiao2,YANG Ping2,NIU Feng2
Author information +
文章历史 +

摘要

为提高声级计校准工作效率,提出了一种基于深度学习神经网络的声级计图像读数检测与识别方法。读数检测模型以DBNet(differentiable binarization network)为基础模型,将ShuffleNetV2作为主干网络,显著降低模型参数量;为提高读数区域检测精度,引入高效通道注意力ECA模块,提高网络对于通道特征的提取能力,优化后的模型在保持精度的同时参数量缩减为原来的15.4%,计算量缩减为原来的67.4%。读数识别模型以CRNN((convolutional recurrent neural network)(为基础模型,先加入批量规范化层,提高网络训练时的稳定性;然后,引入残差块替换原有的卷积块,提高了网络对于复杂特征的提取能力;将Dropout应用于网络中,提高网络的泛化能力;此外,在合成读数数据集上对读数识别模型进行预训练,有效增加了模型准确率。改进后的方法准确率达到了99.7%,相较原方法提高了2.4%。实验结果表明,该方法对声级计图像中存在的字体多样、光照不均、模糊等影响因素具有较强的鲁棒能力,对声级计图像中的读数具有较高的识别精度。

Abstract

In order to improve the work efficiency of the sound level meter calibration work, a reading detection and recognition method based on deep learning neural network for the image of the sound level meter is proposed. The reading detection model is based on DBNet and uses ShuffleNetV2 as the backbone network, significantly reducing the number of model parameters. To improve the accuracy of reading area detection, an efficient channel attention ECA module is introduced to enhance the network's ability to extract channel features. The optimized model reduces the number of parameters to 15.4% and the calculation amount to 67.4% while maintaining accuracy. The reading recognition model is based on the CRNN model, which first adds a batch normalization layer to improve the stability of network training. Then, residual blocks are introduced to replace the original convolutional blocks, improving the network's ability to extract complex features. Applying Dropout to the network to improve its generalization ability. In addition, pre-training the reading recognition model on the synthesized reading dataset effectively increases the accuracy of the model. The improved method achieves an accuracy of 99.7%, which is 2.4% higher than the original method. The experimental results show that this method has strong robustness against factors such as diverse fonts, uneven lighting, and blurring in sound level meter images, and has high recognition accuracy for readings in sound level meter images.

关键词

声学计量 / 图像识别;声级计;DBNet;CRNN;注意力模块

Key words

acoustic metrology;image recognition / sound level meter / DBNet / CRNN / attention module

引用本文

导出引用
汪嘉,祝海江,王寅初,何龙标,杨平,牛锋. 基于改进DBNet-CRNN的声级计读数识别方法[J]. 计量学报. 2024, 45(8): 1191-1199 https://doi.org/10.3969/j.issn.1000-1158.2024.08.15
WANG Jia,ZHU Haijiang,WANG Yinchu,HE Longbiao,YANG Ping,NIU Feng. Sound Level Meter Reading Recognition Method Based on Improved DBNet-CRNN[J]. Acta Metrologica Sinica. 2024, 45(8): 1191-1199 https://doi.org/10.3969/j.issn.1000-1158.2024.08.15
中图分类号: TB95   

参考文献

[1]ZHANG Q H, WAN C X, BIAN S F, et al. Research on Intelligent Instrument Character Location Technology Based on Computer Vision[C]//2018 3rd International Conference on Robotics and Automation Engineering (ICRAE). Guangzhou, China, 2018. [2]SUN Q S, WANG X J, ZHONG B, et al. Design and Realization of Reading Recognition System for Sound Level Meter[C]//2011 Third International Conference on Measuring Technology and Mechatronics Automation. Shanghai, China, 2011. [3]LI X J, GAO Y. Digitisation of Conventional Water Meters using Automated Number Recognition[C]//TENCON 2021—2021 IEEE Region 10 Conference (TENCON). Auckland, New Zealand, 2021. [4]ZHANG Z J, CHEN G H, LI J W, et al. The research on digit recognition algorithm for automatic meter reading system[C]//2010 8th World Congress on Intelligent Control and Automation. Jinan, China, 2010. [5]LIU J, WU H, CHEN Z. Automatic detection and recognition method of digital instrument representation[C]//2021 4th International Conference on Intelligent Autonomous Systems (ICoIAS). Wuhan, China, 2021. [6]LAROCA R, ARAUJO A B, ZANLORENSI L A, et al. Towards image-based automatic meter reading in unconstrained scenarios: A robust and efficient approach[J]. IEEE Access, 2021, 9: 67569-67584. [7]SRIPANUSKUL N, BUAYAI P, MAO X. Generative Data Augmentation for Automatic Meter Reading Using CNNs[J]. IEEE Access, 2022, 10: 28471-28486. [8]周逸, 潘永杲, 闵琪涛, 等. 基于机器视觉的多表位数字温湿度计自动校准系统研究[J]. 计量学报, 2023, 44(8): 1214-1221. ZHOU Y, PAN Y G, MIN Q T, et al. Research on Automatic Calibration System of Multiple Digital Thermo-hygrometers Based on Machine Vision [J]. Acta Metrologica Sinica, 2023, 44(8): 1214-1221. [9]伍开宇, 朱海清, 沈晓东, 等. 基于机器视觉的指针式压力表智能检定系统研究[J]. 计量学报, 2022, 43(11): 1450-1455. WU K Y, ZHU H Q, SHEN X D, et al. Research on Intelligent Verification System of Pointer Pressure Meter Based on Machine Vision[J]. Acta Metrologica Sinica, 2022, 43(11): 1450-1455. [10]LIAO M, WAN Z, YAO C, et al. Real-time scene text detection with differentiable binarization[C]//Proceedings of the AAAI conference on artificial intelligence. New York, USA, 2020. [11]SHI B, BAI X, YAO C. An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition[J]. IEEE transactions on pattern analysis and machine intelligence, 2016, 39(11): 2298-2304. [12]MA N, ZHANG X, ZHENG H T, et al. Shufflenet v2: Practical guidelines for efficient cnn architecture design[C]//Proceedings of the European conference on computer vision(ECCV). Munich, Germany, 2018. [13]WANG Q, WU B, ZHU P, et al. ECA-Net: Efficient channel attention for deep convolutional neural networks[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. Seattle, USA, 2020. [14]IOFFE S, SZEGEDY C. Batch normalization: Accelerating deep network training by reducing internal covariate shift[C]//International conference on machine learning. Lile, France, 2015. [15]HE K, ZHANG X, REN S, et al. Deep residual learning for image recognition[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. Las Vegas, USA, 2016. [16]HE K, ZHANG X, REN S, et al. Identity mappings in deep residual networks[C]//Proceedings of the European conference on computer vision(ECCV). Amsterdam, Netherlands, 2016. [17]SRIVASTAVA N, HINTON G, KRIZHEVSKY A, et al. Dropout: a simple way to prevent neural networks from overfitting[J]. The journal of machine learning research, 2014, 15(1): 1929-1958. [18]HOWARD A, SANDLER M, CHU G, et al. Searching for mobilenetv3[C]//Proceedings of the IEEE/CVF international conference on computer vision. Seoul, Korea, 2019. [19]HU J, SHEN L, SUN G. Squeeze-and-excitation networks[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. Salt Lake City, USA,2018. [20]WOO S, PARK J, LEE J Y, et al. Cbam: Convolutional block attention module[C]//Proceedings of the European conference on computer vision (ECCV). Munich, Germany,2018. [21]YANG L, ZHANG R Y, LI L, et al. Simam: A simple, parameter-free attention module for convolutional neural networks[C]//International conference on machine learning. 2021. [22]ROY A G, NAVAB N, WACHINGER C. Concurrent spatial and channel ‘squeeze & excitation’in fully convolutional networks[C]//Medical Image Computing and Computer Assisted Intervention-MICCAI 2018: 21st International Conference. Granada, Spain, 2018. [23]ZHU Y, CHEN J, LIANG L, et al. Fourier contour embedding for arbitrary-shaped text detection[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2021. [24]WANG W, XIE E, SONG X, et al. Efficient and accurate arbitrary-shaped text detection with pixel aggregation network[C]//Proceedings of the IEEE/CVF international conference on computer vision. Seoul, Korea,2019. [25]WANG W, XIE E, LI X, et al. Shape robust text detection with progressive scale expansion network[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. Long Beach, USA,2019. [26]oh-my-ocr. text_render [EB/OL]. [2023-07-12]. https: //github. com/ oh-my-ocr/text_renderer. [27]LEE C Y, OSINDERO S. Recursive recurrent nets with attention modeling for ocr in the wild[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. Las Vegas, USA,2016. [28]WANG J, HU X. Gated recurrent convolution neural network for ocr[C]//Advances in Neural Information Processing Systems.Long Beach, California,USA,2017. [29]BORISYUK F, GORDO A, SIVAKUMAR V. Rosetta: Large scale system for text detection and recognition in images[C]//Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining. London, UK,2018. [30]DU Y, CHEN Z, JIA C, et al. Svtr: Scene text recognition with a single visual model[J]. arXiv preprint,arXiv: 2205. 00159. [31]SHI B, WANG X, LYU P, et al. Robust scene text recognition with automatic rectification[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. 2016. [32]LIU W, CHEN C, WONG K Y K, et al. Star-net: a spatial attention residue network for scene text recognition[C]//Proceedings of the British Machine Vision Conference. York, UK, 2016.

基金

国家重点研发计划(2021YFF0600202)

PDF(729 KB)

Accesses

Citation

Detail

段落导航
相关文章

/