簡易檢索 / 詳目顯示

研究生: 王維俊
WANG, WEI-JUN
論文名稱: 基於深度學習之小提琴音訊與圖形記譜的雙向轉換
Deep Learning-Based Bidirectional Conversion Between Violin Audio and Graphic Notation
指導教授: 蘇文鈺
Su, Alvin Wen-Yu
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 78
中文關鍵詞: 小提琴音訊圖形記譜雙向轉換深度學習希爾伯特轉換循環一致性
外文關鍵詞: violin audio, graphic notation, bidirectional conversion, deep learning, Hilbert transform, cycle consistency
相關次數: 點閱:31下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 傳統五線譜難以完整涵蓋演奏的特性,以小提琴而言,包含運弓力度、顫音等細微的演奏變化與動態特徵。相較之下,圖形記譜具備記錄多元表現手法的潛力,能提供不同於傳統記譜的呈現方式,以更具象的視覺符號捕捉音樂細節。隨著深度學習技術在音樂資訊檢索與合成領域的蓬勃發展,建立符號與實際音訊之間的直接聯繫成為可行的研究方向。
    本論文提出一個雙向轉換的神經網路系統,旨在探索並建立小提琴音訊與圖形記譜之間的映射關係,整體系統架構主要包含音訊轉圖像與圖像轉音訊兩個核心模組。在音訊特徵處理階段,本研究主要基於希爾伯特轉換所取得之資訊來萃取聲學特徵,並將其映射至包含一種圖形譜所使用之筆觸軌跡與視覺紋理的畫布表徵上,藉此將時間軸上的聲音變化轉化為空間中的圖像資訊。核心目標在於維持雙向資訊轉換的一致性:其一為原始小提琴錄音音訊轉換為圖形譜後,再還原回音訊的過程;其二為圖形譜轉換為音訊後,再轉回圖形譜的過程。
    在理想情況下,期望經由系統還原之結果能與原始輸入保持一致,以證明圖形記譜確實涵蓋了重構音訊所需的充足資訊。雖然,在實務的訊號重建過程中,要達成完全無損的轉換存在其難度,我們透過訊噪比(SNR)與峰值訊雜比(PSNR)等量化評估指標,客觀分析系統還原結果與原始輸入之間在音訊能量及影像像素上的差異。實驗結果顯示,本研究所提出之架構能建立小提琴音訊與圖形譜的雙向聯繫,實現演奏表情的視覺記錄與初步的音響還原。如何減低轉換的失真以及將此一技術應用在其他樂器是未來努力的方向。

    Traditional staff notation struggles to fully capture the characteristics of a performance; in the case of the violin, these include subtle performance variations and dynamic features such as bowing dynamics and vibrato. Graphic notation, by contrast, has the potential to record diverse expressive techniques, offering a mode of presentation different from traditional notation and capturing musical detail with more concrete visual symbols. With the rapid development of deep learning in music information retrieval and synthesis, establishing a direct link between symbols and actual audio has become a feasible research direction.
    This thesis proposes a bidirectional neural-network conversion system that aims to explore and establish the mapping between violin audio and graphic notation. The overall architecture comprises two core modules: audio-to-graphic and graphic-to-audio. In the audio-feature processing stage, this research extracts acoustic features primarily from the information obtained through the Hilbert transform and maps them onto a canvas representation that carries the stroke trajectories and visual textures of a kind of graphic score, thereby transforming audio variations along the time axis into image information in space. The core objective is to maintain the consistency of the bidirectional conversion: first, converting an original violin recording into a graphic score and then restoring it back to audio; and second, converting a graphic score into audio and then back into a graphic score.
    Ideally, the result restored by the system should remain consistent with the original input, demonstrating that graphic notation indeed contains enough information to reconstruct the audio. Although completely lossless conversion is difficult to achieve in practical signal reconstruction, we use quantitative metrics such as the signal-to-noise ratio (SNR) and the peak signal-to-noise ratio (PSNR) to objectively analyze the differences in audio energy and image pixels between the system's restored result and the original input. Experimental results show that the proposed architecture can establish a bidirectional link between violin audio and graphic scores, achieving both a visual record of performance expression and a preliminary acoustic restoration. Reducing the conversion distortion and applying this technique to other instruments are directions for future work.

    中文摘要 i Abstract iii Contents v List of Tables vii List of Figures viii 1 Introduction 1 2 Related Work 7 2.1 The Development of Graphic Notation and Its Parameterization Method 7 2.2 Multiband Discrete Hilbert Transform (MDHT) 14 3 System Architecture 18 3.1 The Audio-to-Graph Subsystem 19 3.1.1 The Definition of Audio and Audio Feature 20 3.1.2 The Definition of Graph Feature and Graph 22 3.2 The Graph-to-Audio Subsystem 24 3.2.1 The Relationship Between Graph and Graph Feature 24 3.2.2 The Restoration of Audio Feature and Audio 25 3.3 The Design of Bidirectional Round-Trip and Cycle Consistency 26 3.3.1 The Audio → Graph → Audio Cycle Route 27 3.3.2 The Graph → Audio → Graph Cycle Route 28 4 Experimental Results and Analysis 29 4.1 Data Preparation and Preprocessing 29 4.1.1 Preprocessing and Standardization of the Audio Feature 31 4.1.2 Preprocessing and Standardization of the Graph Feature 32 4.2 Evaluation Criteria 34 4.3 Modular Decomposition, Separate Pre-training, and Concatenated Fine-tuning 35 4.3.1 The Audio-to-Graph Direction 36 4.3.2 The Graph-to-Audio Direction 41 4.4 Experimental Results 43 4.4.1 The Audio → Graph → Audio Cycle Result 44 4.4.2 The Graph → Audio → Graph Cycle Result 47 5 Conclusion and Future Outlook 56 5.1 Research Summary 56 5.2 Research Limitations 58 5.3 Future Outlook 59 References 64

    [1] Hsin-Lei Sun, Wei-Jun Wang, Ping-Yi Chen, Alvin Su, and Wen-Hsiang Lu. Graphics is our MIDI. Unpublished manuscript, submitted to the International Conference on Digital Audio Effects (DAFx), 2025.
    [2] Gerard Marino, Marie-Hélène Serra, and Jean-Michel Raczinski. The UPIC system: Origins and innovations. Perspectives of New Music, 31(1):258–269, 1993.
    [3] Jesse Engel, Lamtharn Hantrakul, Chenjie Gu, and Adam Roberts. DDSP: Differentiable digital signal processing. In International Conference on Learning Representations (ICLR), 2020.
    [4] Hsin-Lei Sun, Wei-Jun Wang, Ping-Yi Chen, Alvin Su, and Wen-Hsiang Lu. Multiband discrete Hilbert transform for music tones analysis/manipulation and its VST implementation. In Audio Engineering Society (AES) 159th Convention, Long Beach, CA, USA, 2025.
    [5] Earle Brown. Folio and 4 Systems. Earle Brown Music Foundation, 2026. [Online]. Available: https://earle-brown.org/work/folio-and-4-systems/. Accessed: July 26, 2026.
    [6] Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016.
    [7] Radoslaw Weychan, Tomasz Marciniak, Agnieszka Stankiewicz, and Adam Dabrowski. Real time recognition of speakers from internet audio stream. Foundations of Computing and Decision Sciences, 40(3):223–233, 2015.

    下載圖示
    校外:立即公開
    QR CODE