数字农科院2.0

A New Chinese Named Entity Recognition Method for Pig Disease Domain Based on Lexicon-Enhanced BERT and Contrastive Learning

文献类型: 外文期刊

作者: Cheng Peng;Xiajun Wang;Qifeng Li;Qinyang Yu;Ruixiang Jiang;Weihong Ma;Wenbiao Wu;Rui Meng;Haiyan Li;Heju Huai;Shuyan Wang;Longjuan He

作者机构:

关键词: Chinese named entity recognition;contrastive learning;lexicon-enhanced BERT;pig disease;small sample

期刊名称: Applied Sciences (Switzerland)

ISSN: 2076-3417

年卷期: 2024 年 14 卷 16 期

页码:

收录情况: SCIE(2024版) ; ; EI(2024版)

摘要: Featured Application: Our work provides reliable technical support for the information extraction of pig diseases in Chinese. It can be applied to other domain - specific fields, thereby facilitating seamless adaptation for named entity identification across diverse contexts. Named Entity Recognition (NER) is a fundamental and pivotal stage in the development of various knowledge-based support systems, including knowledge retrieval and question-answering systems. In the domain of pig diseases, Chinese NER models encounter several challenges, such as the scarcity of annotated data, domain-specific vocabulary, diverse entity categories, and ambiguous entity boundaries. To address these challenges, we propose PDCNER, a Pig Disease Chinese Named Entity Recognition method leveraging lexicon-enhanced BERT and contrastive learning. Firstly, we construct a domain-specific lexicon and pre-train word embeddings in the pig disease domain. Secondly, we integrate lexicon information of pig diseases into the lower layers of BERT using a Lexicon Adapter layer, which employs char–word pair sequences. Thirdly, to enhance feature representation, we propose a lexicon-enhanced contrastive loss layer on top of BERT. Finally, a Conditional Random Field (CRF) layer is employed as the model’s decoder. Experimental results show that our proposed model demonstrates superior performance over several mainstream models, achieving a precision of 87.76%, a recall of 86.97%, and an F1-score of 87.36%. The proposed model outperforms BERT-BiLSTM-CRF and LEBERT by 14.05% and 6.8%, respectively, with only 10% of the samples available, showcasing its robustness in data scarcity scenarios. Furthermore, the model exhibits generalizability across publicly available datasets. Our work provides reliable technical support for the information extraction of pig diseases in Chinese and can be easily extended to other domains, thereby facilitating seamless adaptation for named entity identification across diverse contexts.

分类号:

  • 相关文献
作者其他论文 更多>>