| 研究生: |
賴彥儒 Lai, Yan-Ru |
|---|---|
| 論文名稱: |
使用雙向長短期記憶網路偵測GTTM局部規則及邊界 Detection of GTTM Local Boundary and Rules by Using Bidirectional Long Short-Term Memory Neural Network |
| 指導教授: |
蘇文鈺
Su, Wen-Yu |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 英文 |
| 論文頁數: | 35 |
| 中文關鍵詞: | 調性音樂生成理論 、機器學習 、分組偏好規則 、雙向長短期記憶網路 、人工數據生成 、樂譜局部邊界偵測 |
| 外文關鍵詞: | A Generative Theory of Tonal Music, Machine Learning, Grouping Preference Rules, Bidirectional Long Short-Term Memory Network, Synthetic Data Generation, Music Score Local Boundary Detection |
| 相關次數: | 點閱:200 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
在調性音樂生成理論(generative theory of tonal music, GTTM)中,分組結構被視為基礎且重要的結構。它代表著聽眾對音樂理解最基本的組成,同時包含大量可用於生成表情音樂以及偵測樂句分段等應用的資訊。局部邊界(Local Boundary)代表著分組結構中界於不同群組之間音符的間隔,而分組偏好規則(grouping preference rules, GPRs)可以用來尋找局部邊界。因此本篇論文使用兩個架構不同的雙向長短期記憶(bidirectional long short-term memory, BLSTM)網路來分別訓練出偵測分組偏好規則以及局部邊界的模型。
為了訓練深度學習模型,首先我們提出可以生成大量具有標記規則的樂譜的方法,以減少專家收集資料和人工標記所花費的時間與精力。我們藉由自動生成的樂譜訓練出能偵測GPR6的BLSTM網路。另一方面,為了增加局部邊界模型的訓練資料,我們透過合乎樂理的方式擴增人工標記的樂譜,並且使用已訓練好GPRs的BLSTM網路輔助我們訓練出偵測局部邊界的BLSTM網路。
實驗結果顯示,使用生成資料所訓練出偵測GPR6的BLSTM網路已經比利用人工標記資料訓練出的BLSTM表現更優異且穩定,而偵測local boundary的BLSTM網路不只表現良好,它也改善了現存的模型輸入長度被限制的問題。
In the generative theory of tonal music (GTTM), the grouping structure is regarded as a basic and important structure. It represents the basic composition of listeners' understanding of music, and contains a lot of information that can be used for applications such as generating expression music and detecting phrase segmentation. Local boundary (Local Boundary) represents the interval of notes between different groups in the grouping structure, and grouping preference rules (GPRs) can be used to find local boundaries. Therefore, this paper uses two bidirectional long short-term memory (BLSTM) networks with different architectures to separately train the detection group preference rules and local boundaries.
To train the deep learning model, first, we propose a method that can generate a large amount of music scores with labeling rules to reduce the time and effort spent by experts in collecting data and manual labeling. We train a BLSTM network that can detect GPR6 by using automatically generated music scores. On the other hand, we use reasonable music theory methods to augment manually labeled scores as training data for local boundary model, and use the trained GPR BLSTM network to help us train the BLSTM network that can detect local boundaries.
According to the experimental results, the BLSTM network trained with the generated data to detect GPR6 has performed better and more stable than the BLSTM trained with manual labeling data. The BLSTM network that detects the local boundary not only performs well, but also improves the limitation of the length of the input in the existing model.
[1] ChingYeh Chen. Automated phrase analysis of sonatas. Master's thesis, National Cheng Kung University, 2019.
[2] Tsung-Ping Chen, Li Su, et al. Functional harmony recognition of symbolic music data with multi-task recurrent neural networks. In ISMIR, pages 90–97, 2018.
[3] David Cope. The algorithmic composer, volume 16. AR Editions, Inc., 2000.
[4] Michael Scott Cuthbert and Christopher Ariza. music21: A toolkit for computer-aided musicology and symbolic music data. 2010.
[5] Pavel Filonov, Andrey Lavrentyev, and Artem Vorontsov. Multivariate industrial time series with cyber-attack simulation: Fault detection using an lstm-based predictive data model. arXiv preprint arXiv:1612.06676, 2016.
[6] Christos Garoufis, Athanasia Zlatintsi, and Petros Maragos. An lstm-based dynamic chord progression generation system for interactive music performance. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 4502–4506. IEEE, 2020.
[7] Felix A Gers, Douglas Eck, and Jürgen Schmidhuber. Applying lstm to time series predictable through time-window approaches. In Neural Nets WIRN Vietri-01, pages 193–200. Springer, 2002.
[8] Michael Good. Musicxml for notation and analysis. The virtual score: representation, retrieval, restoration, 12(113-124):160, 2001.
[9] Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio. Deep learning, volume 1. MIT press Cambridge, 2016.
[10] Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman. Synthetic data for text localisation in natural images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2315–2324, 2016.
[11] M Hamanaka, K Hirata, and S Tojo. deepgttm-i: Local boundaries analyzer based on deep learning technique. In 12th International Symposium on Computer Music Multidisciplinary Research (CMMR), 2016.
[12] Masatoshi Hamanaka, Keiji Hirata, and Satoshi Tojo. Implementing “a generative theory of tonal music". Journal of New Music Research, 35(4):249–277, 2006.
[13] Masatoshi Hamanaka, Keiji Hirata, and Satoshi Tojo. Musical structural analysis database based on gttm. 2014.
[14] Masatoshi Hamanaka and Satoshi Tojo. Interactive gttm analyzer. In ISMIR, pages 291–296, 2009.
[15] Geoffrey E Hinton. Deep belief networks. Scholarpedia, 4(5):5947, 2009.
[16] Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
[17] Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Synthetic data and artificial neural networks for natural scene text recognition. arXiv preprint arXiv:1406.2227, 2014.
[18] Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur. Audio augmentation for speech recognition. In Sixteenth Annual Conference of the International Speech Communication Association, 2015.
[19] Fred Lerdahl and Ray S Jackendoff. A Generative Theory of Tonal Music, reissue, with a new preface. MIT press, 1996.
[20] Agnieszka Mikołajczyk and Michał Grochowski. Data augmentation for improving deep learning in image classification problem. In 2018 international interdisciplinary PhD workshop (IIPhDW), pages 117–122. IEEE, 2018.
[21] Marius Miron, Jordi Janer Mestres, and Emilia Gómez Gutiérrez. Generating data to train convolutional neural networks for classical music source separation. In Lokki T, Pätynen J, Välimäki V, editors. Proceedings of the 14th Sound and Music Computing Conference; 2017 Jul 5-8; Espoo, Finland. Aalto: Aalto University; 2017. p. 227-33. Aalto University, 2017.
[22] Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le. Specaugment: A simple data augmentation method for automatic speech recognition. arXiv preprint arXiv:1904.08779, 2019.
[23] Mike Schuster and Kuldip K Paliwal. Bidirectional recurrent neural networks. IEEE transactions on Signal Processing, 45(11):2673–2681, 1997.
[24] Connor Shorten and Taghi M Khoshgoftaar. A survey on image data augmentation for deep learning. Journal of Big Data, 6(1):1–48, 2019.
[25] YouCheng Siao. Detection of GTTM Local Boundary Rules by Using Algorithmically Generated Music Scores. PhD thesis, 2020.
[26] Geraint A Wiggins. Computer models of musical creativity: A review of computer models of musical creativity by david cope. Literary and Linguistic Computing, 23(1):109– 116, 2008.
[27] Adrien Ycart, Emmanouil Benetos, et al. A study on lstm networks for polyphonic music sequence modelling. ISMIR, 2017.