Download:
|
by Jing-shin Chang, Keh-yih Su
Department of Electrical Engineering, National Tsing-Hua University
http://www.bdc.com.tw/~shin/doc/rocling/cpnr.RocX.final.ps.Z
Add To MetaCart
Abstract:
An improved statistical model is proposed in this paper for extracting compound words from a text corpus. Traditional terminology extraction methods rely heavily on simple filtering-and-thresholding methods, which are unable to minimize the error counts objectively. Therefore, a method for minimizing the error counts is very desirable. In this paper, an improved statistical model is developed to integrate parts of speech information as well as other frequently used word association metrics to jointly optimize the extraction tasks. The features are modelled with a multivariate Gaussian mixture for handling the inter-feature correlations properly. With a training (resp. testing) corpus of 20715 (resp. 2301) sentences, the weighted precision & recall (WPR) can achieve about 84 % for bigram compounds, and 86 % for trigram compounds. The F-measure performances are about 82 % for bigrams and 84 % for trigrams.
Citations
|
2961
|
Pattern Classification and Scene Analysis
– Duda, Hart
- 1973
|
|
2217
|
J.: Introduction to Modern Information Retrieval
– Salton, Macgill
- 1983
|
|
880
|
and B-H Juang, Fundamentals of Speech Recognition
– Rabiner
- 1993
|
|
428
|
Word association norms, mutual information, and lexicography
– Church, Hanks
- 1990
|
|
209
|
Principles and Practice of Information Theory
– Blahut
- 1988
|
|
162
|
Retrieving collocations from text: Xtract
– Smadja
- 1993
|
|
113
|
FASTUS: a finite-state processor for information extraction from real-world text
– Appelt, Hobbs, et al.
- 1993
|
|
103
|
FASTUS: A Cascaded Finite-State Transducer for Extracting Information from Natural-Language Text
– Hobbs, Appelt, et al.
- 1997
|
|
95
|
Discriminative learning for minimum error classification
– Juang, Katagiri
- 1992
|
|
50
|
Probability and Statistics
– Papoulis
- 1990
|
|
23
|
Acquisition of Lexical Information from a Large Textual Italian Corpus
– Calzolari, Bindi
- 1990
|
|
21
|
An Unsupervised Iterative Method for Chinese New Lexicon Extraction
– Chang, Su
- 1997
|
|
11
|
A Corpus-based Approach to Automatic Compound Extraction
– Su, Wu, et al.
- 1994
|
|
10
|
Speech Recognition Using Weighted HMM and Subspace Projection Approaches
– Su, Lee
- 1994
|
|
8
|
Discrimination oriented probabilistic tagging
– LIN, CHIANG
- 1992
|
|
4
|
The Identification and Classification of Unknown Words in Chinese: An N-Grams-Based Approach
– Wang, Huang, et al.
- 1995
|
|
3
|
Automatic Lexicon Acquisition and Precision-Recall Maximization for Untagged Text Corpora
– Chang
- 1997
|
|
3
|
A First Course
– Roussas
- 1973
|
|
2
|
Yi-Chung Lin and Keh-Yih Su, "Automatic Construction of a Chinese Electronic Dictionary
– Chang
- 1995
|
|
2
|
The Processing of English Compound and Complex Words in an English-Chinese Machine Translation System
– Chen, Su
- 1988
|
|
2
|
Extracting Information from the
– Hirschman, Vilain
- 1995
|
|
2
|
The effects of learning, parameter tying and model refinement for improving probabilistic tagging
– Lin, Chiang, et al.
- 1995
|
|
2
|
and Vasileios Hatzivassiloglou, "Translating Collocations for Bilingual Lexicons: A Statistical Approach
– Smadja, McKeown
- 1996
|
|
2
|
Constructing a Phrase Structure Grammar by Incorporating Linguistic Knowledge and Statistical Log-Likelihood Ratio
– Su, Hsu, et al.
- 1991
|
|
1
|
Surface Grammar Analysis for the Extraction of Terminological Noun Phrases
– Bourigault
- 1992
|