CALBC silver standard corpus

Rebholz-Schuhmann, Dietrich; Yepes, Antonio José Jimeno; van Mulligen, Erik; Kang, Ning; Kors, Jan; Milward, David; Corbett, Peter; Buyko, Ekaterina; Beisswanger, Elena; Hahn, Udo

doi:10.1142/S0219720010004562

The CALBC initiative aims to provide a large-scale biomedical text corpus that contains semantic annotations for named entities of different kinds. The generation of this corpus requires that the annotations from different automatic annotation systems be harmonized. In the first phase, the annotation systems from five participants (EMBL-EBI, EMC Rotterdam, NLM, JULIE Lab Jena, and Linguamatics) were gathered. All annotations were delivered in a common annotation format that included concept identifiers in the boundary assignments and that enabled comparison and alignment of the results. During the harmonization phase, the results produced from those different systems were integrated in a single harmonized corpus ("silver standard" corpus) by applying a voting scheme. We give an overview of the processed data and the principles of harmonization formal boundary reconciliation and semantic matching of named entities. Finally, all submissions of the participants were evaluated against that silver standard corpus. We found that species and disease annotations are better standardized amongst the partners than the annotations of genes and proteins. The raw corpus is now available for additional named entity annotations. Parts of it will be made available later on for a public challenge. We expect that we can improve corpus building activities both in terms of the numbers of named entity classes being covered, as well as the size of the corpus in terms of annotated documents.

Additional Metadata
Keywords	Biomedical text mining, Corpus generation, Named entity recognition
Persistent URL	doi.org/10.1142/S0219720010004562, hdl.handle.net/1765/67943
Journal	Journal of Bioinformatics and Computational Biology
Organisation	Department of Medical Informatics
Citation APA Style AAA Style APA Style Cell Style Chicago Style Harvard Style IEEE Style MLA Style Nature Style Vancouver Style American-Institute-of-Physics Style Council-of-Science-Editors Style BibTex Format Endnote Format RIS Format CSL Format DOIs only Format	Rebholz-Schuhmann, D., Yepes, A. J. J., van Mulligen, E., Kang, N., Kors, J., Milward, D., … Hahn, U. (2010). CALBC silver standard corpus. Journal of Bioinformatics and Computational Biology, 8(1), 163–179. doi:10.1142/S0219720010004562

CALBC silver standard corpus

Publication

Publication

About

CALBC silver standard corpus

Publication

Publication

Workflow

Workflow

Add Content