☆ 4.1 Article

CustFRE: An annotated dataset for extraction of family relations from English text

DATA IN BRIEF (2022)

期刊

DATA IN BRIEF

卷 41, 期 -, 页码 -

出版社

ELSEVIER

DOI: 10.1016/j.dib.2022.107980

关键词

Natural language processing; Relation classification; Machine learning; Family relations

类别

Multidisciplinary Sciences

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

智能总结 New
摘要

Meaningful Information extraction is a crucial task, and it requires annotated datasets which are scarce. This manuscript presents a dataset, CustFRE, for extracting family relations from text, which can be used as a benchmark for evaluating and training family relation extraction systems.

Meaningful Information extraction is an extremely important and challenging task due to the ever growing size of data. Training and evaluating automated systems for the task requires annotated datasets which are rarely available because of the great amount of human effort and time required for annotating data. The dataset described in this manuscript, CustFRE, is meant for systems that learn extracting family relations from text. Sentences having at least two persons have been collected from the internet. The texts are first processed using Stanford's NLP pipeline for basic NLP tagging. Next, a team of natural language processing experts annotated the dataset. All family relations among persons in the texts have been annotated, or a no_relation is annotated if no family relation between two persons can be inferred from the text. After annotation, the dataset was verified by an NLP expert for completeness and correctness. CustFRE contains in total 2,716 annotations. The dataset can be used by information extraction researchers as a benchmark for evaluating their systems, and can also be used for training and evaluating family relation extraction systems. (C) 2022 The Authors. Published by Elsevier Inc.

CustFRE: An annotated dataset for extraction of family relations from English text

期刊

DATA IN BRIEF

出版社

ELSEVIER

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

CustFRE: An annotated dataset for extraction of family relations from English text

期刊

DATA IN BRIEF

出版社

ELSEVIER

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文