※ 國立清華大學語言所101學年度第二學期專題演講 ※
演講者: 高照明 (國立台灣大學)
演講題目: Applications of Corpus-based Computer-Aided Chinese Lexical Semantic Researches
時間: 2013 年 3 月 5日(星期二) 中午 12 時 30 分
地點: 語言所研討室(人社院 B305 室)
講題摘要
In recent years, corpus-based computational approaches to lexical semantic researches have shed new light on thorny problems such as word sense discrimination/disambiguation, the identification of new senses, and the extraction of semantically related words. Following this line of research, this talk presents a pilot study on Chinese lexical semantics based on a corpus-based computer-aided approach. Exploiting the assumption that lexical semantics can be revealed by lexical syntax, we develop a Chinese dependency parser based on the Sinica Chinese Treebank. The Chinese dependency parser can automatically analyze dependency relations such as subject/object, verb/noun, and modifier/noun of a sentence. The context of a verb can then be modeled by the collocations which bear dependency relations with it. By grouping the collocations of a verb in terms of its dependency relations and semantic information derived from HowNet, we are able to study a number of Chinese lexical semantic problems in a systematic way. Exploring the syntagmatic and paradigmatic relations of words, we can identify different senses of a word, differences in synonyms, as well as sets of semantically related words. We extend our approach to the analyses of words which have new meanings. We find that semantic change of words is generally accompanied by two clues, namely, significant change of word frequency, and the change of collocations. Based on the information pertaining to the distributions of word frequency for a given word in a newspaper corpus, we can predict the possible period in which new meanings of this word begin to emerge. By comparing the change of lexical, syntactic, and semantic information of the co-occurring words before and after the predicted period, semantic change of words can be identified. The applications of this approach to phrasal semantics, lexicography, E-learning, and the differences between the Mandarin used in Taiwan and China will also be discussed.
Keywords: word sense disambiguation, semantic change, word frequency distributions, contextual similarity, collocations, Chinese newspaper corpus
講者簡介
Dr. Gao Zhao-Ming received his Ph.D. in language engineering from the University of Manchester in 1998. Prior to joining the faculty of the Department of Foreign Languages and Literatures at National Taiwan University in 1999, where he currently works as associate professor, he was a postdoctoral research fellow at the Institute of Linguistics at Academia Sinica and assistant professor at National Chi Nan University. He has served as a reviewer for numerous conferences and journals, including COLING, PACLIC, Lexical Resources and Evaluations, and Computer-Assisted Language Learning. He was the linguistic section editor of the Journal of Computational Linguistics and Chinese Language Processing during 2008-2010. From 2008 to 2012, he was head of the Division of Information Processing as well a board member of the National Languages Committee at Ministry of Education in Taiwan. He has been a board member of the Association of Computational Linguistics and Chinese Language Processing in Taiwan since 2007. Dr. Gao has a keen interest in developing corpus-based computational tools and has published extensively on corpus linguistics, computer-assisted translation, and intelligent computer-assisted language learning.