| 研究生: |
彭北定 Peng, Bei-Ding |
|---|---|
| 論文名稱: |
基因符號辨識器─智慧型基因名稱識別 Gene Symbol Tagger–An intelligent gene names detector |
| 指導教授: |
蔣榮先
Chiang, Jung-Hsien |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2004 |
| 畢業學年度: | 92 |
| 語文別: | 中文 |
| 論文頁數: | 43 |
| 中文關鍵詞: | 基因名稱識別 、名稱識別 |
| 外文關鍵詞: | gene symbols recognition |
| 相關次數: | 點閱:175 下載:1 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
專有名詞的識別在資訊萃取的步驟中扮演著重要的角色,文件中資訊的描述,皆由專有名詞引述出資訊的所在。在生物醫學的領域的文件中,基因及蛋白質的名稱就是必須優先擷取的專有名詞。
傳統的基因名稱識別,是以基因名稱字彙典比對的方式進行文件中基因名稱的識別,其中存在比對耗時,與無法識別字彙典中不存在的基因名稱等缺點。本論文提出一套有別於傳統基因名稱識別的系統,系統中利用單字字型規則、文句中的單字特徵、詞性特徵與詞性序列規則來識別文件中之基因名稱,單字字型規則將選出候選的單一字基因名稱,接著利用詞意特徵與詞性特徵建立篩選候選單一字基因名稱的分類器,以提高識別的正確性。此外,也利用有限狀態機判斷詞性序列以識別多單字的基因名稱。
經由實驗證明,本論文提出的基因名稱識別系統具有80%的基因名稱識別精確率,並克服傳統字彙典比對所面臨的問題。
The rapid increasing of biological literature makes protein related information such as biological process, molecular function and protein-protein interactions more abundant and profuse. The gene symbols (or protein names) play definitely the leading roles while describing a biological incident by sentences.
As we known, gene symbols have various ways to be represented by different authors. Except spelling variation, some gene symbols have named as a single word, and some may be named more than one word such as “pituitary-specific factor Pit-1”. For this reason, it’s more effective to identify gene symbols by separating single word gene symbols from compound gene symbols. In single word gene symbols identification, rule-based method can recognize most of potential names and Multilayer perceptrons(MLP) classifier is employed for increasing recognition precision. Compound gene symbols were be extracted from finite state machines ( FSM )-based framework.
Our approach provide efficiency method to tagging gene symbols. We achieve 80% precision and 60% recall rates in gene symbols recognition.
[1]. Altschul, S.F., Gish, W., Miller, W., Myers, E.W., Lipman, D.J., "Basic local alignment search tool.", J. Mol. Biol. 215, pp 403–410., 1990
[2]. Blaschke,C., Andrade,M., Ouzounis,C. and Valencia,A., "Automatic extraction of biological information from scientific text: protein–protein interactions.", In Proceedings of the International Conference on Intelligent Systems for Molecular Biology, pp 60–67, 1999
[3]. Ciria R, Abreu-Goodger C., Morett E., and Merino E., "GeConT: gene context analysis.", Bioinformatics, Advance Access published April 8, 2004
[4]. Collier, N., Nobata, C., and Tsujii, J., "Extracting the names of genes and gene products with a hidden markov model.", In Proceedings of the 18th International Conference on Computational Linguistics, pp 201-207, 2002
[5]. E. Wiener, J.O. Pedersen, and A.S. Weigend., "A neural network approach to topic spotting.", In Proceedings of the Fourth Annual Symposium on Document Analysis and Information Retrieval (SDAIR'95), 1995
[6]. Fukuda,K., Tamura,A., Tsunoda,T. and Takagi,T., "Toward information extraction: identifying protein names from biological papers.", Pac. Symp. Biocomp., 3, pp 707–718, 1998
[7]. GuoDong Zhou, Jie Zhang1, Jian Su, Dan Shen and ChewLim Tan, "Recognizing Names in Biomedical Texts: a Machine Learning Approach.", Bioinformatics, Advance Access published February 10, 2004
[8]. Jenssen,T., Laegreid,A., Komorowski,J. and Hovig,E., "A literature network of human genes for high-throughput analysis of gene expression.", Nat. Genet., 28, pp 21–28., 2001
[9]. Joshua M. Temkin and Mark R., Gilder. "Extraction of protein interaction information from unstructured text using a context-free grammar.", Bioinformatics, 19, pp 2046–2053, 2003
[10]. Jung-Hsien Chiang and Hsu-Chun Yu, "MeKE: discovering the functions of gene products from biomedical literature via sentence alignment.", Bioinformatics, 19, pp 1417–1422., 2003
[11]. Jung-Hsien Chiang, Hsu-Chun Yu and Huai-Jen Hsu, "GIS: a biomedical text-mining system for gene information discovery.", Bioinformatics, 20,
pp 120–121., 2004
[12]. Kazuhiro Seki and Javed Mostafa, "An Approach to Protein Name Extraction using Heuristics and a Dictionary", ASIST 2003 Annual Meeting - Humanizing Information Technology: From Ideas to Bits and Back, 2003
[13]. Kenneth Ward Church and Patrick Hanks., "Word association norms, mutual information and lexicography.", In Proceedings of ACL 27, pp 76-83,Vancouver, Canada,1989
[14]. Kyungsook Han, Byungkyu Park, Hyongguen Kim, Jinsun Hong and Jong Park, "HPID: The Human Protein Interaction Database.", Bioinformatics, Advance Access published April 29, 2004
[15]. Langley P., Elements of Machine Learning., San Francisco:Morgan Kaufmann Publishers, Inc., 1996
[16]. Lorraine Tanabe and W. John Wilbur., "Tagging gene and protein names in biomedical text.", Bioinformatics , 18, pp 1124-1132, 2002
[17]. Lorraine Tanabe and W. John Wilbur., "Tagging gene and protein names in full text articles.", Proceedings of the ACL-02 Workshop on Natural Language Processing in the Biomedical Domain, pp 9-13, 2002
[18]. Michael Krauthammer, Andrey Rzhetsky, Pavel Morozov, Carol Friedman, "Using BLAST for identifying gene and protein names in journal articles", Gene, 259, pp 245–252, 2000
[19]. Mitchell TM., Machine Learning., Boston: WCB/McGraw-Hill,1997
[20]. Ng,S. and Wong,M., "Toward routine automatic pathway discovery from on-line scientific text abstracts.", Genome Inform., 10, pp 104–112., 1999
[21]. Nikolai Daraselia., Anton Yuryev, Sergei Egorov, Svetalana Novichkova, Alexander Nikitin and Ilya Mazo, "Extracting human protein interactions from MEDLINE using a full-sentence parser.", Bioinformatics, 20, pp 604–611, 2004
[22]. Nobata, C., Collier, N., and Tsujii, J., "Automatic term identification and classification in biology texts.", In Proceedings of the 5th Natural Language Processing Pacific Rim Symposium, pp 369-374, 1999
[23]. Olsson, F., Eriksson, G., Franzen, K., Asker, L., and Liden, P., "Notions of correctness when evaluating protein name taggers.", In Proceedings of the 19th International Conference on Computational Linguistics., 2002
[24]. Ono,T., Hishigaki,H., Tanigami,A. and Takagi,T., "Automated extraction of information on protein–protein interactions from the biological literature.", Bioinformatics, 17, pp 155–161., 2001
[25]. R. Fano., Transmission of Information., MIT Press, Cambridge, MA, 1961
[26]. Simon Haykin, NEURAL NETWORKS a comprehensive foundation(second edition), Prentice Hall Interactional, Incs., Upper Saddle River, NJ, 1999
[27]. Thomas,J., Milward,D., Ouzounis,C., Pulman,S. and Carroll,M., "Automatic extraction of protein interactions from scientific abstracts.", Pac. Symp. Biocomp., 5, pp 541–552., 2000
[28]. Wilbur WJ., "Boosting Naive Bayesian Learning on a Large Subset of MEDLINE.", American Medical Informatics 2000 Annual Symposium., Los Angeles, CA: American Medical Informatics Association, pp 918-922, 2000
[29]. Wong,L., "Pies, a protein interaction extraction system.", Pac. Symp. Biocomp., 6, pp 520–531., 2001
[30]. Yakushiji,A., Tateisi,Y., Miyao,Y. and ichi Tsujii,J., "Event extraction from biomedical papers using a full parser.", Pac. Symp. Biocomp., 6, pp 408–419., 2001
[31]. Yoshimasa Tsuruoka and Jun chi Tsujii, "Boosting precision and recall of dictionary-based protein name recognition.", Proceedings of the ACL 2003 Workshop on Natural Language Processing in Biomedicine, pp 41-48, 2003.
[32]. Andrew McCallum, “Information Extraction : Coreference and Relation Extraction.”
[33]. Richard Hughey and Kevin Karplus, “Bioinformatics : A new field in engineering education.”
[34]. Language Modeling of Biological Data, University of Pennsylvania, February 2001: http://www.ircs.upenn.edu/modeling2001/.
[35]. Workshops on Natural Language Processing in the Biomedical Domain, Association of Computational Linguistics,
July 2002: http://www.ccs.neu.edu/home/futrelle/bionlp/acl02/BIO/
July 2003: http://www-tsujii.is.s.u-tokyo.ac.jp/ACL03/bionlp.htm