簡易檢索 / 詳目顯示

研究生: 彭北定
Peng, Bei-Ding
論文名稱: 基因符號辨識器─智慧型基因名稱識別
Gene Symbol Tagger–An intelligent gene names detector
指導教授: 蔣榮先
Chiang, Jung-Hsien
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2004
畢業學年度: 92
語文別: 中文
論文頁數: 43
中文關鍵詞: 基因名稱識別名稱識別
外文關鍵詞: gene symbols recognition
相關次數: 點閱:175下載:1
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  •   專有名詞的識別在資訊萃取的步驟中扮演著重要的角色,文件中資訊的描述,皆由專有名詞引述出資訊的所在。在生物醫學的領域的文件中,基因及蛋白質的名稱就是必須優先擷取的專有名詞。
      傳統的基因名稱識別,是以基因名稱字彙典比對的方式進行文件中基因名稱的識別,其中存在比對耗時,與無法識別字彙典中不存在的基因名稱等缺點。本論文提出一套有別於傳統基因名稱識別的系統,系統中利用單字字型規則、文句中的單字特徵、詞性特徵與詞性序列規則來識別文件中之基因名稱,單字字型規則將選出候選的單一字基因名稱,接著利用詞意特徵與詞性特徵建立篩選候選單一字基因名稱的分類器,以提高識別的正確性。此外,也利用有限狀態機判斷詞性序列以識別多單字的基因名稱。
      經由實驗證明,本論文提出的基因名稱識別系統具有80%的基因名稱識別精確率,並克服傳統字彙典比對所面臨的問題。

      The rapid increasing of biological literature makes protein related information such as biological process, molecular function and protein-protein interactions more abundant and profuse. The gene symbols (or protein names) play definitely the leading roles while describing a biological incident by sentences.
      As we known, gene symbols have various ways to be represented by different authors. Except spelling variation, some gene symbols have named as a single word, and some may be named more than one word such as “pituitary-specific factor Pit-1”. For this reason, it’s more effective to identify gene symbols by separating single word gene symbols from compound gene symbols. In single word gene symbols identification, rule-based method can recognize most of potential names and Multilayer perceptrons(MLP) classifier is employed for increasing recognition precision. Compound gene symbols were be extracted from finite state machines ( FSM )-based framework.
      Our approach provide efficiency method to tagging gene symbols. We achieve 80% precision and 60% recall rates in gene symbols recognition.

    第一章 導論.....1 1.1 前言.....1 1.2 研究動機.....1 1.3 解決方法.....2 1.4 系統概述.....3 1.5 論文架構.....4 第二章 相關研究.....5 2.1 生物資訊學.....5 2.2 資訊萃取(Information Extraction).....5 2.3 資訊萃取運用在生物醫學文件.....6 2.4 基因名稱識別.....6 2.4.1 以字彙典為基礎之基因名稱識別.....6 2.4.2 以規則為基礎之基因名稱識別.....7 2.4.3 統計學理論之基因名稱識別.....7 2.5 類神經網路.....8 2.5.1 類神經網路應用.....8 2.5.2 類神經網路特性.....8 2.5.3 類神經網路架構與學習演算法.....9 第三章 可容錯式多型態基因名稱標記系統.....10 3.1 系統架構.....10 3.2 文件前處理.....12 3.2.1 文件斷句.....12 3.2.2 詞性標記.....12 3.2.3 特殊單字類別之建立.....13 3.3 單一字基因名稱特性與識別規則.....13 3.4 單一字基因名稱分類程序.....14 3.4.1 特徵選取(Feature Selection).....14 3.4.2 多層式感知器(Multilayer Perceptrons).....17 3.5 多單字基因名稱識別.....17 3.5.1 縮寫字還原規則.....17 3.5.2 有限狀態機建立與執行.....18 3.6 基因名稱合併程序.....21 第四章 實驗與結果分析.....22 4.1 實驗資料集.....22 4.1.1 BioCreAtIvE資料集.....22 4.1.2 GENIA資料集.....26 4.2 實驗流程與結果.....32 4.2.1 BioCreAtIvE資料集實驗結果與分析.....32 4.2.2 GENIA資料集實驗結果與分析.....34 第五章 結論與未來研究方向.....36 5.1. 結論.....36 5.2. 未來研究方向.....37 參考文獻.....38 附錄一 相關術語.....42

    [1]. Altschul, S.F., Gish, W., Miller, W., Myers, E.W., Lipman, D.J., "Basic local alignment search tool.", J. Mol. Biol. 215, pp 403–410., 1990

    [2]. Blaschke,C., Andrade,M., Ouzounis,C. and Valencia,A., "Automatic extraction of biological information from scientific text: protein–protein interactions.", In Proceedings of the International Conference on Intelligent Systems for Molecular Biology, pp 60–67, 1999

    [3]. Ciria R, Abreu-Goodger C., Morett E., and Merino E., "GeConT: gene context analysis.", Bioinformatics, Advance Access published April 8, 2004

    [4]. Collier, N., Nobata, C., and Tsujii, J., "Extracting the names of genes and gene products with a hidden markov model.", In Proceedings of the 18th International Conference on Computational Linguistics, pp 201-207, 2002

    [5]. E. Wiener, J.O. Pedersen, and A.S. Weigend., "A neural network approach to topic spotting.", In Proceedings of the Fourth Annual Symposium on Document Analysis and Information Retrieval (SDAIR'95), 1995

    [6]. Fukuda,K., Tamura,A., Tsunoda,T. and Takagi,T., "Toward information extraction: identifying protein names from biological papers.", Pac. Symp. Biocomp., 3, pp 707–718, 1998

    [7]. GuoDong Zhou, Jie Zhang1, Jian Su, Dan Shen and ChewLim Tan, "Recognizing Names in Biomedical Texts: a Machine Learning Approach.", Bioinformatics, Advance Access published February 10, 2004

    [8]. Jenssen,T., Laegreid,A., Komorowski,J. and Hovig,E., "A literature network of human genes for high-throughput analysis of gene expression.", Nat. Genet., 28, pp 21–28., 2001

    [9]. Joshua M. Temkin and Mark R., Gilder. "Extraction of protein interaction information from unstructured text using a context-free grammar.", Bioinformatics, 19, pp 2046–2053, 2003

    [10]. Jung-Hsien Chiang and Hsu-Chun Yu, "MeKE: discovering the functions of gene products from biomedical literature via sentence alignment.", Bioinformatics, 19, pp 1417–1422., 2003

    [11]. Jung-Hsien Chiang, Hsu-Chun Yu and Huai-Jen Hsu, "GIS: a biomedical text-mining system for gene information discovery.", Bioinformatics, 20,
    pp 120–121., 2004

    [12]. Kazuhiro Seki and Javed Mostafa, "An Approach to Protein Name Extraction using Heuristics and a Dictionary", ASIST 2003 Annual Meeting - Humanizing Information Technology: From Ideas to Bits and Back, 2003

    [13]. Kenneth Ward Church and Patrick Hanks., "Word association norms, mutual information and lexicography.", In Proceedings of ACL 27, pp 76-83,Vancouver, Canada,1989

    [14]. Kyungsook Han, Byungkyu Park, Hyongguen Kim, Jinsun Hong and Jong Park, "HPID: The Human Protein Interaction Database.", Bioinformatics, Advance Access published April 29, 2004

    [15]. Langley P., Elements of Machine Learning., San Francisco:Morgan Kaufmann Publishers, Inc., 1996

    [16]. Lorraine Tanabe and W. John Wilbur., "Tagging gene and protein names in biomedical text.", Bioinformatics , 18, pp 1124-1132, 2002

    [17]. Lorraine Tanabe and W. John Wilbur., "Tagging gene and protein names in full text articles.", Proceedings of the ACL-02 Workshop on Natural Language Processing in the Biomedical Domain, pp 9-13, 2002

    [18]. Michael Krauthammer, Andrey Rzhetsky, Pavel Morozov, Carol Friedman, "Using BLAST for identifying gene and protein names in journal articles", Gene, 259, pp 245–252, 2000

    [19]. Mitchell TM., Machine Learning., Boston: WCB/McGraw-Hill,1997

    [20]. Ng,S. and Wong,M., "Toward routine automatic pathway discovery from on-line scientific text abstracts.", Genome Inform., 10, pp 104–112., 1999

    [21]. Nikolai Daraselia., Anton Yuryev, Sergei Egorov, Svetalana Novichkova, Alexander Nikitin and Ilya Mazo, "Extracting human protein interactions from MEDLINE using a full-sentence parser.", Bioinformatics, 20, pp 604–611, 2004

    [22]. Nobata, C., Collier, N., and Tsujii, J., "Automatic term identification and classification in biology texts.", In Proceedings of the 5th Natural Language Processing Pacific Rim Symposium, pp 369-374, 1999

    [23]. Olsson, F., Eriksson, G., Franzen, K., Asker, L., and Liden, P., "Notions of correctness when evaluating protein name taggers.", In Proceedings of the 19th International Conference on Computational Linguistics., 2002

    [24]. Ono,T., Hishigaki,H., Tanigami,A. and Takagi,T., "Automated extraction of information on protein–protein interactions from the biological literature.", Bioinformatics, 17, pp 155–161., 2001

    [25]. R. Fano., Transmission of Information., MIT Press, Cambridge, MA, 1961

    [26]. Simon Haykin, NEURAL NETWORKS a comprehensive foundation(second edition), Prentice Hall Interactional, Incs., Upper Saddle River, NJ, 1999

    [27]. Thomas,J., Milward,D., Ouzounis,C., Pulman,S. and Carroll,M., "Automatic extraction of protein interactions from scientific abstracts.", Pac. Symp. Biocomp., 5, pp 541–552., 2000

    [28]. Wilbur WJ., "Boosting Naive Bayesian Learning on a Large Subset of MEDLINE.", American Medical Informatics 2000 Annual Symposium., Los Angeles, CA: American Medical Informatics Association, pp 918-922, 2000

    [29]. Wong,L., "Pies, a protein interaction extraction system.", Pac. Symp. Biocomp., 6, pp 520–531., 2001

    [30]. Yakushiji,A., Tateisi,Y., Miyao,Y. and ichi Tsujii,J., "Event extraction from biomedical papers using a full parser.", Pac. Symp. Biocomp., 6, pp 408–419., 2001

    [31]. Yoshimasa Tsuruoka and Jun chi Tsujii, "Boosting precision and recall of dictionary-based protein name recognition.", Proceedings of the ACL 2003 Workshop on Natural Language Processing in Biomedicine, pp 41-48, 2003.

    [32]. Andrew McCallum, “Information Extraction : Coreference and Relation Extraction.”

    [33]. Richard Hughey and Kevin Karplus, “Bioinformatics : A new field in engineering education.”

    [34]. Language Modeling of Biological Data, University of Pennsylvania, February 2001: http://www.ircs.upenn.edu/modeling2001/.

    [35]. Workshops on Natural Language Processing in the Biomedical Domain, Association of Computational Linguistics,
    July 2002: http://www.ccs.neu.edu/home/futrelle/bionlp/acl02/BIO/
    July 2003: http://www-tsujii.is.s.u-tokyo.ac.jp/ACL03/bionlp.htm

    下載圖示
    2005-07-30公開
    QR CODE