| 研究生: |
王維俊 WANG, WEI-JUN |
|---|---|
| 論文名稱: |
基於深度學習之小提琴音訊與圖形記譜的雙向轉換 Deep Learning-Based Bidirectional Conversion Between Violin Audio and Graphic Notation |
| 指導教授: |
蘇文鈺
Su, Alvin Wen-Yu |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 78 |
| 中文關鍵詞: | 小提琴音訊 、圖形記譜 、雙向轉換 、深度學習 、希爾伯特轉換 、循環一致性 |
| 外文關鍵詞: | violin audio, graphic notation, bidirectional conversion, deep learning, Hilbert transform, cycle consistency |
| 相關次數: | 點閱:31 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
傳統五線譜難以完整涵蓋演奏的特性,以小提琴而言,包含運弓力度、顫音等細微的演奏變化與動態特徵。相較之下,圖形記譜具備記錄多元表現手法的潛力,能提供不同於傳統記譜的呈現方式,以更具象的視覺符號捕捉音樂細節。隨著深度學習技術在音樂資訊檢索與合成領域的蓬勃發展,建立符號與實際音訊之間的直接聯繫成為可行的研究方向。
本論文提出一個雙向轉換的神經網路系統,旨在探索並建立小提琴音訊與圖形記譜之間的映射關係,整體系統架構主要包含音訊轉圖像與圖像轉音訊兩個核心模組。在音訊特徵處理階段,本研究主要基於希爾伯特轉換所取得之資訊來萃取聲學特徵,並將其映射至包含一種圖形譜所使用之筆觸軌跡與視覺紋理的畫布表徵上,藉此將時間軸上的聲音變化轉化為空間中的圖像資訊。核心目標在於維持雙向資訊轉換的一致性:其一為原始小提琴錄音音訊轉換為圖形譜後,再還原回音訊的過程;其二為圖形譜轉換為音訊後,再轉回圖形譜的過程。
在理想情況下,期望經由系統還原之結果能與原始輸入保持一致,以證明圖形記譜確實涵蓋了重構音訊所需的充足資訊。雖然,在實務的訊號重建過程中,要達成完全無損的轉換存在其難度,我們透過訊噪比(SNR)與峰值訊雜比(PSNR)等量化評估指標,客觀分析系統還原結果與原始輸入之間在音訊能量及影像像素上的差異。實驗結果顯示,本研究所提出之架構能建立小提琴音訊與圖形譜的雙向聯繫,實現演奏表情的視覺記錄與初步的音響還原。如何減低轉換的失真以及將此一技術應用在其他樂器是未來努力的方向。
Traditional staff notation struggles to fully capture the characteristics of a performance; in the case of the violin, these include subtle performance variations and dynamic features such as bowing dynamics and vibrato. Graphic notation, by contrast, has the potential to record diverse expressive techniques, offering a mode of presentation different from traditional notation and capturing musical detail with more concrete visual symbols. With the rapid development of deep learning in music information retrieval and synthesis, establishing a direct link between symbols and actual audio has become a feasible research direction.
This thesis proposes a bidirectional neural-network conversion system that aims to explore and establish the mapping between violin audio and graphic notation. The overall architecture comprises two core modules: audio-to-graphic and graphic-to-audio. In the audio-feature processing stage, this research extracts acoustic features primarily from the information obtained through the Hilbert transform and maps them onto a canvas representation that carries the stroke trajectories and visual textures of a kind of graphic score, thereby transforming audio variations along the time axis into image information in space. The core objective is to maintain the consistency of the bidirectional conversion: first, converting an original violin recording into a graphic score and then restoring it back to audio; and second, converting a graphic score into audio and then back into a graphic score.
Ideally, the result restored by the system should remain consistent with the original input, demonstrating that graphic notation indeed contains enough information to reconstruct the audio. Although completely lossless conversion is difficult to achieve in practical signal reconstruction, we use quantitative metrics such as the signal-to-noise ratio (SNR) and the peak signal-to-noise ratio (PSNR) to objectively analyze the differences in audio energy and image pixels between the system's restored result and the original input. Experimental results show that the proposed architecture can establish a bidirectional link between violin audio and graphic scores, achieving both a visual record of performance expression and a preliminary acoustic restoration. Reducing the conversion distortion and applying this technique to other instruments are directions for future work.
[1] Hsin-Lei Sun, Wei-Jun Wang, Ping-Yi Chen, Alvin Su, and Wen-Hsiang Lu. Graphics is our MIDI. Unpublished manuscript, submitted to the International Conference on Digital Audio Effects (DAFx), 2025.
[2] Gerard Marino, Marie-Hélène Serra, and Jean-Michel Raczinski. The UPIC system: Origins and innovations. Perspectives of New Music, 31(1):258–269, 1993.
[3] Jesse Engel, Lamtharn Hantrakul, Chenjie Gu, and Adam Roberts. DDSP: Differentiable digital signal processing. In International Conference on Learning Representations (ICLR), 2020.
[4] Hsin-Lei Sun, Wei-Jun Wang, Ping-Yi Chen, Alvin Su, and Wen-Hsiang Lu. Multiband discrete Hilbert transform for music tones analysis/manipulation and its VST implementation. In Audio Engineering Society (AES) 159th Convention, Long Beach, CA, USA, 2025.
[5] Earle Brown. Folio and 4 Systems. Earle Brown Music Foundation, 2026. [Online]. Available: https://earle-brown.org/work/folio-and-4-systems/. Accessed: July 26, 2026.
[6] Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016.
[7] Radoslaw Weychan, Tomasz Marciniak, Agnieszka Stankiewicz, and Adam Dabrowski. Real time recognition of speakers from internet audio stream. Foundations of Computing and Decision Sciences, 40(3):223–233, 2015.