Parallel Creation of Gigaword Corpora for Medium Density Languages - an Interim Report

Presented at: The Sixth International Language Resources and Evaluation Conference (LREC2008)

by Péter Halácsy, András Kornai, Péter Németh, Dániel Varga

Webpage: http://www.lrec-conf.org/proceedings/lrec2008/pdf/858_paper.pdf
Webpage: http://www.lrec-conf.org/proceedings/lrec2008/summaries/858.html

For increased speed in developing gigaword language resources for medium resource density languages we integrated several FOSS tools in the HUN* toolkit. While the speed and efficiency of the resulting pipeline has surpassed our expectations, our experience in developing LDC-style resource packages for Uzbek and Kurdish makes clear that neither the data collection nor the subsequent processing stages can be fully automated.

Keywords: Corpus (creation, annotation, etc.), Multilinguality, Tools, systems, applications, Linguistics


Resource URI on the dog food server: http://data.semanticweb.org/conference/lrec/2008/papers/858


Explore this resource elsewhere: