数字农科院2.0

Transfer large models to crop pest recognition—A cross-modal unified framework for parameters efficient fine-tuning

文献类型: 外文期刊

作者: Jianping Liu;Jialu Xing;Guomin Zhou;Jian Wang;Lulu Sun;Xi Chen

作者机构:

关键词: CLIP;Computer vision;Cross-modal fusion;Large pre-training model;Parameters efficient fine-tuning;Pest recognition

期刊名称: Computers and Electronics in Agriculture

ISSN: 0168-1699

年卷期: 2025 年 237 卷

页码:

收录情况: SCIE(2025版) ; ; EI(2025版)

摘要: Crop pest recognition is an important direction in agricultural research, which is of great significance for improving crop yield and scientifically classifying pests for precision agriculture. Traditional deep learning pest recognition usually trains proprietary models on single categories and scenes as well as unimodal information, achieving excellent performance. However, this scheme has a weak foundation of general knowledge, insufficient transferability, and unimodal information has limited effect on the recognition of pest background and different life stages. In recent years, transferring the general knowledge of Large pre-trained models (LPTM) to specific domains through full fine-tuning has become an effective solution. However, full fine-tuning requires massive data and operator resources to effectively adapt all parameters. Therefore, this paper proposes a cross-modal parameter efficient fine-tuning (PEFT) unified framework for crop pest recognition with the multimodal large model CLIP as the pre-training model. The proposed method employs CLIP as the encoder for both image and text modalities, introducing the Dual-(PAL)G model. Firstly, learnable Prompt sequences are embedded in the input or hidden layers of the encoder. Secondly, multimodal LoRA is parallelly replaced in the dimension expansion layer of the fully connected layer. Then, the Gate unit integrates three PEFT methods—Prompt, Adapter, and LoRA, to enhance learning ability. We designed the GSC-Adapter and the parameter-efficient Light-GCS-Adapter for cross-modal semantic information fusion. To verify the effectiveness of the method, we conducted a large number of experiments on public datasets for crop pest recognition. Firstly, on the public dataset IP102 (for fine-grained recognition), we surpassed ViT and Swin Transformer with 66% of the sample size. In wolfberry pest dataset WPIT9K, using only about 15% of the sample size, it surpasses the previous state-of-the-art model ITF-WPI, achieving 98% accuracy. It also shows excellent performance on eight general tasks. This study provides a new technical solution for the field of agricultural pest recognition. This solution can efficiently transfer the general knowledge of multimodal LPTM to the specific pest recognition field under the condition of a few samples, with only a minimal number of parameters introduced. At the same time, this method has universality in cross-modal recognition tasks. The code for this study will be posted on GitHub (https://github.com/VcRenOne/Dual–PAL-G)

分类号:

  • 相关文献

[1]ITF-WPI: Image and text based cross-modal feature fusion model for wolfberry pest recognition. Guowei Dai,Jingchao Fan,Christine Dewi. 2023

[2]基于参数高效微调的跨模态枸杞虫害识别模型D-PAG. 邢嘉璐,刘建平,周国民,刘立波,王健. 2024

[3]Developing a hybrid convolutional neural network for automatic aphid counting in sugar beet fields. Xumin Gao,Wenxin Xue,Callum Lennox,Mark Stevens,Junfeng Gao. 2024

[4]A novel cross-modal decoupling dual-spectral fusion system and method: A case study on the on-site estimation of fresh tobacco leaf curing characteristics. Zhongtao Huang,Shichang Wang,Rongguang Zhu,Yapeng Kang,Lingfeng Meng,Jie Ren. 2025

[5]Outdoor color rating of sweet cherries using computer vision. Wang, Qi,Wang, Hui,Xie, Lijuan,Zhang, Qin,Wang, Hui,Xie, Lijuan.

[6]Classification of weed seeds based on visual images and deep learning. Tongyun Luo,Jianye Zhao,Yujuan Gu,Shuo Zhang,Xi Qiao,Wen Tian,Yangchun Han. 2023

[7]Feasibility assessment of tree-level flower intensity quantification from UAV RGB imagery: A triennial study in an apple orchard. Chenglong Zhang,João Valente,Wensheng Wang,Leifeng Guo,Aina Tubau Comas,Pieter van Dalfsen,Bert Rijk,Lammert Kooistra. 2023

[8]Maize-IAS: a maize image analysis software using deep learning for high-throughput plant phenotyping. Zhou Shuo,Chai Xiujuan,Yang Zixuan,Wang Hongwu,Yang Chenxue,Sun Tan. 2021

[9]Maize-IAS: a maize image analysis software using deep learning for high-throughput plant phenotyping. Zhou Shuo,Chai Xiujuan,Yang Zixuan,Wang Hongwu,Yang Chenxue,Sun Tan. 2021

[10]Editorial: Remote sensing for field-based crop phenotyping. Jiangang Liu,Zhenjiang Zhou,Bo Li. 2024

[11]A survey of efficient fine-tuning methods for Vision-Language Models — Prompt and Adapter. Xing J.,Liu J.,Wang J.,Sun L.,Chen X.,Gu X.,Wang Y.. 2024

[12]SPP-extractor: Automatic phenotype extraction for densely grown soybean plants. Zhou, Wan,Chen, Yijie,Li, Weihao,Zhang, Cong,Xiong, Yajun,Zhan, Wei,Huang, Lan,Wang, Jun,Qiu, Lijuan. 2023

[13]Editorial: Advanced technologies for energy saving, plant quality control and mechanization development in plant factory. Yuxin Tong,Myung Min Oh,Wei Fang. 2023

[14]Vision-based measuring method for individual cow feed intake using depth images and a Siamese network. Xinjie Wang,Baisheng Dai,Xiaoli Wei,Weizheng Shen,Yonggen Zhang,Benhai Xiong. 2023

[15]Accurate recognition of the reproductive development status and prediction of oviposition fecundity in Spodoptera frugiperda (Lepidoptera: Noctuidae) based on computer vision. Chun yang LÜ,Shi shuai GE,Wei HE,Hao wen ZHANG,Xian ming YANG,Bo CHU,Kong ming WU. 2023

[16]Intelligent weight prediction of cows based on semantic segmentation and back propagation neural network. Beibei Xu,Yifan Mao,Wensheng Wang,Guipeng Chen. 2024

[17]Keypoint detection and diameter estimation of cabbage (Brassica oleracea L.) heads under varying occlusion degrees via YOLOv8n-CK network. Jinming Zheng,Xiaochan Wang,Yinyan Shi,Xiaolei Zhang,Yao Wu,Dezhi Wang,Xuekai Huang,Yanxin Wang,Jihao Wang,Jianfei Zhang. 2024

[18]A hybrid method for water stress evaluation of rice with the radiative transfer model and multidimensional imaging. Yufan Zhang,Xiuliang Jin,Liangsheng Shi,Yu Wang,Han Qiao,Yuanyuan Zha. 2025

[19]Lightweight deep learning model for embedded systems efficiently predicts oil and protein content in rapeseed. Mengshuai Guo,Huifang Ma,Xin Lv,Dan Wang,Li Fu,Ping He,Desheng Mei,Hong Chen,Fang Wei. 2025

[20]Recent Advances and Applications of Imaging and Spectroscopy Technologies for Tea Quality Assessment: A Review. Shujun Zhi,Ting An,Han Zhang,Yuhao Bai,Baohua Zhang,Guangzhao Tian. 2025

作者其他论文 更多>>