Impact of Training Instance Selection on Domain-Specific Entity Extraction using BERT

Eileen Salhofer; Xing Lan Liu; Roman Kern

doi:10.18653/v1/2022.naacl-srw.11

Impact of Training Instance Selection on Domain-Specific Entity Extraction using BERT

Eileen Salhofer, Xing Lan Liu, Roman Kern

Publikation: Beitrag in Buch/Bericht/Konferenzband › Beitrag in einem Konferenzband › Begutachtung

Abstract

State of the art performances for entity extraction tasks are achieved by supervised learning, specifically, by fine-tuning pretrained language models such as BERT. As a result, annotating application specific data is the first step in many use cases. However, no practical guidelines are available for annotation requirements. This work supports practitioners by empirically answering the frequently asked questions (1) how many training samples to annotate? (2) which examples to annotate? We found that BERT achieves up to 80% F1 when fine-tuned on only 70 training examples, especially on biomedical domain. The key features for guiding the selection of high performing training instances are identified to be pseudo-perplexity and sentence-length. The best training dataset constructed using our proposed selection strategy shows F1 score that is equivalent to a random selection with twice the sample size. The requirement of only a small number of training data implies cheaper implementations and opens door to wider range of applications.

Originalsprache	englisch
Titel	Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Student Research Workshop
Herausgeber (Verlag)	Association for Computational Linguistics
Seiten	83-88
DOIs	https://doi.org/10.18653/v1/2022.naacl-srw.11
Publikationsstatus	Veröffentlicht - 2022
Veranstaltung	2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics: NAACL 2022 - Seattle, Hybrider Event, USA / Vereinigte Staaten Dauer: 10 Juli 2022 → 15 Juli 2022

Konferenz

Konferenz	2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics
Kurztitel	NAACL 2022
Land/Gebiet	USA / Vereinigte Staaten
Ort	Hybrider Event
Zeitraum	10/07/22 → 15/07/22

Zugriff auf Dokument

10.18653/v1/2022.naacl-srw.11Lizenz: CC BY 4.0

Andere Dateien und Links

http://dx.doi.org/10.18653/v1/2022.naacl-srw.11

Dieses zitieren

Salhofer, E., Liu, X. L., & Kern, R. (2022). Impact of Training Instance Selection on Domain-Specific Entity Extraction using BERT. in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Student Research Workshop (S. 83-88). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.naacl-srw.11

Impact of Training Instance Selection on Domain-Specific Entity Extraction using BERT. / Salhofer, Eileen; Liu, Xing Lan; Kern, Roman.
Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Student Research Workshop. Association for Computational Linguistics, 2022. S. 83-88.

Publikation: Beitrag in Buch/Bericht/Konferenzband › Beitrag in einem Konferenzband › Begutachtung

Salhofer, E, Liu, XL & Kern, R 2022, Impact of Training Instance Selection on Domain-Specific Entity Extraction using BERT. in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Student Research Workshop. Association for Computational Linguistics, S. 83-88, 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics, Hybrider Event, USA / Vereinigte Staaten, 10/07/22. https://doi.org/10.18653/v1/2022.naacl-srw.11

Salhofer E, Liu XL, Kern R. Impact of Training Instance Selection on Domain-Specific Entity Extraction using BERT. in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Student Research Workshop. Association for Computational Linguistics. 2022. S. 83-88 doi: 10.18653/v1/2022.naacl-srw.11

@inproceedings{cb3614ef25064013a61f7c43a3e68e89,

title = "Impact of Training Instance Selection on Domain-Specific Entity Extraction using BERT",

abstract = "State of the art performances for entity extraction tasks are achieved by supervised learning, specifically, by fine-tuning pretrained language models such as BERT. As a result, annotating application specific data is the first step in many use cases. However, no practical guidelines are available for annotation requirements. This work supports practitioners by empirically answering the frequently asked questions (1) how many training samples to annotate? (2) which examples to annotate? We found that BERT achieves up to 80% F1 when fine-tuned on only 70 training examples, especially on biomedical domain. The key features for guiding the selection of high performing training instances are identified to be pseudo-perplexity and sentence-length. The best training dataset constructed using our proposed selection strategy shows F1 score that is equivalent to a random selection with twice the sample size. The requirement of only a small number of training data implies cheaper implementations and opens door to wider range of applications.",

author = "Eileen Salhofer and Liu, {Xing Lan} and Roman Kern",

year = "2022",

doi = "10.18653/v1/2022.naacl-srw.11",

language = "English",

pages = "83--88",

booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Student Research Workshop",

publisher = "Association for Computational Linguistics",

note = "2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics : NAACL 2022, NAACL 2022 ; Conference date: 10-07-2022 Through 15-07-2022",

}

TY - GEN

T1 - Impact of Training Instance Selection on Domain-Specific Entity Extraction using BERT

AU - Salhofer, Eileen

AU - Liu, Xing Lan

AU - Kern, Roman

PY - 2022

Y1 - 2022

N2 - State of the art performances for entity extraction tasks are achieved by supervised learning, specifically, by fine-tuning pretrained language models such as BERT. As a result, annotating application specific data is the first step in many use cases. However, no practical guidelines are available for annotation requirements. This work supports practitioners by empirically answering the frequently asked questions (1) how many training samples to annotate? (2) which examples to annotate? We found that BERT achieves up to 80% F1 when fine-tuned on only 70 training examples, especially on biomedical domain. The key features for guiding the selection of high performing training instances are identified to be pseudo-perplexity and sentence-length. The best training dataset constructed using our proposed selection strategy shows F1 score that is equivalent to a random selection with twice the sample size. The requirement of only a small number of training data implies cheaper implementations and opens door to wider range of applications.

AB - State of the art performances for entity extraction tasks are achieved by supervised learning, specifically, by fine-tuning pretrained language models such as BERT. As a result, annotating application specific data is the first step in many use cases. However, no practical guidelines are available for annotation requirements. This work supports practitioners by empirically answering the frequently asked questions (1) how many training samples to annotate? (2) which examples to annotate? We found that BERT achieves up to 80% F1 when fine-tuned on only 70 training examples, especially on biomedical domain. The key features for guiding the selection of high performing training instances are identified to be pseudo-perplexity and sentence-length. The best training dataset constructed using our proposed selection strategy shows F1 score that is equivalent to a random selection with twice the sample size. The requirement of only a small number of training data implies cheaper implementations and opens door to wider range of applications.

UR - http://dx.doi.org/10.18653/v1/2022.naacl-srw.11

U2 - 10.18653/v1/2022.naacl-srw.11

DO - 10.18653/v1/2022.naacl-srw.11

M3 - Conference paper

SP - 83

EP - 88

BT - Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Student Research Workshop

PB - Association for Computational Linguistics

T2 - 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics

Y2 - 10 July 2022 through 15 July 2022

ER -

Impact of Training Instance Selection on Domain-Specific Entity Extraction using BERT

Abstract

Konferenz

Zugriff auf Dokument

Andere Dateien und Links

Fingerprint

Dieses zitieren