| 研究生: |
王泓閔 Wang, Hung-Min |
|---|---|
| 論文名稱: |
應用於智慧藥局之全方位AI處方箋辨識系統與友善應用介面 Comprehensive AI Prescription Recognition System and Friendly User Interface Applied for Intelligent Pharmacy |
| 指導教授: |
王駿發
Wang, Jhing-Fa |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 英文 |
| 論文頁數: | 51 |
| 中文關鍵詞: | 處方箋辨識 、資訊萃取 、深度學習 、雙向長短期記憶模型 |
| 外文關鍵詞: | Prescription Recognition, Information Extraction, Deep Learning, Bidirectional Long Short-Term Memory |
| 相關次數: | 點閱:197 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
由於處方上有許多欄位,藥師在申報處方箋時會花很多時間。為了提升藥師的工作效率,本研究提出了一個全方位的處方簽辨識系統與友善使用介面。處方辨識系統主要包含三個部分,分別為前處理、資訊萃取與後處理。其中前處理包含了輸入圖像進行對比增強、光學文字辨識,以及對光學文字辨識後的文字進行合併與排序。資訊萃取的部分我們混合兩個不同的策略分別是基於規則的方法與基於深度學習的方法。在基於深度學習的方法中,我們設計三種特徵並使用雙向長短期記憶模型對輸入的文字框進行分類。最後,後處理對一些常見的光學文字辨識錯誤進行校正與對特定的欄位進行格式轉換。在使用者介面上我們以Django建構了一個網頁應用伺服器,提供藥師上傳,編輯,與管理處方。為了使整個申報流程更為流暢,我們也建立了一個使用者端的應用程式,該程式會自動監測掃描完的檔案並與申報系統做介接。在深度學習的資訊萃取方法上,平均的F1-score達到93.8%。在整個辨識系統的實驗中我們測試202張不同的處方箋,並分別比較了基於規則、基於深度學習、與混和兩種的資訊萃取方法。混和的方法在多數欄位上都比基於規則的方法有所提升,在總共15個欄位均達到超過93%的正確率。最後,本系統在平均意見分數上得到4.38分。
Since there are many fields on the prescription, the pharmacists spend a lot of time when declaring prescriptions. In order to improve the work efficiency of pharmacists, this work proposes a comprehensive prescription recognition system and friendly user interface. The prescription recognition system mainly consists of three parts, namely pre-processing, information extraction and post-processing. The pre-processing includes contrast enhancement of the input image, optical character recognition, and the merging and sorting of the characters after optical character recognition. In the information extraction part, we hybrid two different strategies, the rule-based method, and the deep-learning-based method. In the deep-learning-based method, we design three features and use a bidirectional long short-term memory model to classify the input text box. Finally, post-processing corrects some common optical character recognition errors and performs format conversion on specific fields. On the user interface, we build a web application server with Django, which provides pharmacists to upload, edit, and manage prescriptions. In order to make the entire declaration process smoother, we also create a user-side application that will automatically monitor the scanned files and interface with the declaration system. In the deep-learning-based information extraction, the average F1-score is 93.8%. In the experiment of the overall recognition system, we test 202 different prescriptions, and compare the three different information extraction strategies, which are rules-based method, deep-learning-based method, and hybrid method respectively. The hybrid method is improved in most fields, and the accuracy of all the fields in our system is greater than 93%. Finally, the system gets 4.38 point on the average opinion score.
[1] M. T. Qadri and M. Asif, "Automatic number plate recognition system for vehicle identification using optical character recognition," in 2009 International Conference on Education Technology and Computer, 2009: IEEE, pp. 335-338.
[2] N. Ezaki, K. Kiyota, B. T. Minh, M. Bulacu, and L. Schomaker, "Improved text-detection methods for a camera-based text reading system for blind persons," in Eighth International Conference on Document Analysis and Recognition (ICDAR'05), 2005: IEEE, pp. 257-261.
[3] X. Liu, "A camera phone based currency reader for the visually impaired," in Proceedings of the 10th international ACM SIGACCESS conference on Computers and accessibility, 2008, pp. 305-306.
[4] C. Yi and Y. Tian, "Localizing text in scene images by boundary clustering, stroke segmentation, and string fragment classification," IEEE Transactions on Image Processing, vol. 21, no. 9, pp. 4256-4268, 2012.
[5] S. Long, X. He, and C. Yao, "Scene text detection and recognition: The deep learning era," International Journal of Computer Vision, vol. 129, no. 1, pp. 161-184, 2021.
[6] R. Girshick, J. Donahue, T. Darrell, and J. Malik, "Rich feature hierarchies for accurate object detection and semantic segmentation," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 580-587.
[7] R. Girshick, "Fast r-cnn," in Proceedings of the IEEE international conference on computer vision, 2015, pp. 1440-1448.
[8] S. Ren, K. He, R. Girshick, and J. Sun, "Faster r-cnn: Towards real-time object detection with region proposal networks," Advances in neural information processing systems, vol. 28, pp. 91-99, 2015.
[9] W. Liu et al., "Ssd: Single shot multibox detector," in European conference on computer vision, 2016: Springer, pp. 21-37.
[10] Z. Tian, W. Huang, T. He, P. He, and Y. Qiao, "Detecting text in natural image with connectionist text proposal network," in European conference on computer vision, 2016: Springer, pp. 56-72.
[11] Y. Baek, B. Lee, D. Han, S. Yun, and H. Lee, "Character region awareness for text detection," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 9365-9374.
[12] D. Deng, H. Liu, X. Li, and D. Cai, "Pixellink: Detecting scene text via instance segmentation," in Proceedings of the AAAI Conference on Artificial Intelligence, 2018, vol. 32, no. 1.
[13] P. He, W. Huang, Y. Qiao, C. C. Loy, and X. Tang, "Reading scene text in deep convolutional sequences," in Thirtieth AAAI conference on artificial intelligence, 2016.
[14] B. Shi, X. Bai, and C. Yao, "An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition," IEEE transactions on pattern analysis and machine intelligence, vol. 39, no. 11, pp. 2298-2304, 2016.
[15] C.-Y. Lee and S. Osindero, "Recursive recurrent nets with attention modeling for ocr in the wild," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 2231-2239.
[16] R. Grishman and B. M. Sundheim, "Message understanding conference-6: A brief history," in COLING 1996 Volume 1: The 16th International Conference on Computational Linguistics, 1996.
[17] S. R. Eddy, "Hidden markov models," Current opinion in structural biology, vol. 6, no. 3, pp. 361-365, 1996.
[18] J. R. Quinlan, "Induction of decision trees," Machine learning, vol. 1, no. 1, pp. 81-106, 1986.
[19] J. N. Kapur, Maximum-entropy models in science and engineering. John Wiley & Sons, 1989.
[20] M. A. Hearst, S. T. Dumais, E. Osuna, J. Platt, and B. Scholkopf, "Support vector machines," IEEE Intelligent Systems and their applications, vol. 13, no. 4, pp. 18-28, 1998.
[21] J. Lafferty, A. McCallum, and F. C. Pereira, "Conditional random fields: Probabilistic models for segmenting and labeling sequence data," 2001.
[22] D. Nadeau and S. Sekine, "A survey of named entity recognition and classification," Lingvisticae Investigationes, vol. 30, no. 1, pp. 3-26, 2007.
[23] J. Li, A. Sun, J. Han, and C. Li, "A survey on deep learning for named entity recognition," IEEE Transactions on Knowledge and Data Engineering, 2020.
[24] D. Schuster et al., "Intellix--End-User Trained Information Extraction for Document Archiving," in 2013 12th International Conference on Document Analysis and Recognition, 2013: IEEE, pp. 101-105.
[25] M. Rusinol, T. Benkhelfallah, and V. Poulain dAndecy, "Field extraction from administrative documents by incremental structural templates," in 2013 12th International Conference on Document Analysis and Recognition, 2013: IEEE, pp. 1100-1104.
[26] V. P. d'Andecy, E. Hartmann, and M. Rusinol, "Field extraction by hybrid incremental and a-priori structural templates," in 2018 13th IAPR International Workshop on Document Analysis Systems (DAS), 2018: IEEE, pp. 251-256.
[27] R. B. Palm, O. Winther, and F. Laws, "Cloudscan-a configuration-free invoice analysis system using recurrent neural networks," in 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), 2017, vol. 1: IEEE, pp. 406-413.
[28] C. Sage, A. Aussem, H. Elghazel, V. Eglin, and J. Espinas, "Recurrent neural network approach for table field extraction in business documents," in 2019 International Conference on Document Analysis and Recognition (ICDAR), 2019: IEEE, pp. 1308-1313.
[29] P. Zhang et al., "TRIE: End-to-End Text Reading and Information Extraction for Document Understanding," in Proceedings of the 28th ACM International Conference on Multimedia, 2020, pp. 1413-1422.
[30] R. Achkar, K. Ghayad, R. Haidar, S. Saleh, and R. Al Hajj, "Medical Handwritten Prescription Recognition Using CRNN," in 2019 International Conference on Computer, Information and Telecommunication Systems (CITS), 2019: IEEE, pp. 1-5.
[31] E. Hassan, H. Tarek, M. Hazem, S. Bahnacy, L. Shaheen, and W. H. Elashmwai, "Medical Prescription Recognition using Machine Learning," in 2021 IEEE 11th Annual Computing and Communication Workshop and Conference (CCWC), 2021: IEEE, pp. 0973-0979.
[32] P. Wu, F. Wang, and J. Liu, "An Integrated Multi-Classifier Method for Handwritten Chinese Medicine Prescription Recognition," in 2018 IEEE 9th International Conference on Software Engineering and Service Science (ICSESS), 2018: IEEE, pp. 1-4.
[33] Y.-Y. Ou, S.-P. Tseng, J. Lin, X.-P. Zhou, J.-F. Wang, and T.-W. Kuan, "Automatic Prescription Recognition System," in 2018 International Conference on Orange Technologies (ICOT), 2018: IEEE, pp. 1-4.
[34] K. Zuiderveld, "Contrast limited adaptive histogram equalization," in Graphics gems IV: Academic Press Professional, Inc., 1994, pp. 474–485.
[35] 全民健康保險藥品編碼原則. Available: https://www.nhi.gov.tw/Resource/webdata/4168_2_10_%E5%85%A8%E6%B0%91%E5%81%A5%E5%BA%B7%E4%BF%9D%E9%9A%AA%E8%97%A5%E5%93%81%E7%B7%A8%E7%A2%BC%E5%8E%9F%E5%89%87.pdf.
[36] A. Graves and J. Schmidhuber, "Framewise phoneme classification with bidirectional LSTM and other neural network architectures," Neural networks, vol. 18, no. 5-6, pp. 602-610, 2005.
[37] S. Hochreiter and J. Schmidhuber, "Long short-term memory," Neural computation, vol. 9, no. 8, pp. 1735-1780, 1997.
[38] QRcode二維條碼處方箋資料欄位建置規則. Available: https://www.nhi.gov.tw/resource/Webdata/QRcode%E4%BA%8C%E7%B6%AD%E6%A2%9D%E7%A2%BC%E8%99%95%E6%96%B9%E7%AE%8B%E8%B3%87%E6%96%99%E6%AC%84%E4%BD%8D%E5%BB%BA%E7%BD%AE%E8%A6%8F%E5%89%87-1040213.pdf.