Producing an Encyclopedic Dictionary using Patent Documents

Presented at: The Sixth International Language Resources and Evaluation Conference (LREC2008)

by Atsushi Fujii

Webpage: http://www.lrec-conf.org/proceedings/lrec2008/pdf/519_paper.pdf
Webpage: http://www.lrec-conf.org/proceedings/lrec2008/summaries/519.html

Although the World Wide Web has late become an important source to consult for the meaning of words, a number of technical terms related to high technology are not found on the Web. This paper describes a method to produce an encyclopedic dictionary for high-tech terms from patent information. We used a collection of unexamined patent applications published by the Japanese Patent Office as a source corpus. Given this collection, we extracted terms as headword candidates and retrieved applications including those headwords. Then, we extracted paragraph-style descriptions and categorized them into technical domains. We also extracted related terms for each headword. We have produced a dictionary including approximately 400,000 Japanese terms as headwords. We have also implemented an interface with which users can explore our dictionary by reading text descriptions and viewing a related-term graph.

Keywords: Corpus (creation, annotation, etc.), Information Extraction, Information Retrieval, Question Answering, Linguistics


Resource URI on the dog food server: http://data.semanticweb.org/conference/lrec/2008/papers/519


Explore this resource elsewhere: