簡易檢索 / 詳目顯示

研究生: 許庸袁
Hsu, Yong-Yuan
論文名稱: 基於可微數位訊號處理之吉他破音與等化器自動音色匹配系統
Automatic Tone Matching of Guitar Overdrive and Equalizer using Differentiable DSP
指導教授: 賀保羅
Horton, Paul
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 74
中文關鍵詞: 神經音訊效果可微數位訊號處理灰箱建模自動音色匹配人類參與任務解耦
外文關鍵詞: Differentiable DSP (DDSP), Neural Audio Effects, Tone Matching, Grey-box Modeling, Parameter Estimation, Human-in-the-loop
相關次數: 點閱:16下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 現代音樂製作高度依賴電吉他音色設計,但無論是操作實體設備或數位效果器,精準復現特定音色往往需要耗費龐大的調校成本。近期神經音訊效果技術(如虛擬類比與 Neural Amp Modeler)多採用端到端黑箱架構,這類作法雖具備高保真度,卻完全捨棄了效果器實體旋鈕的控制維度。為突破現有技術無法支援「人類參與(Human-in-the-loop)」的實務限制,本研究基於可微數位訊號處理(DDSP)架構,開發一套專注於非線性破音與線性等化器灰箱參數估算之自動音色匹配系統。本研究首先透過系統化均勻採樣,建構涵蓋 5,865 筆 TS9 與 EQ 高質量成對音訊之連續參數資料集。在系統實作上,採用短時傅立葉變換(STFT)萃取聲學特徵,並運用卷積神經網路(CNN)控制器預測對應之實體旋鈕參數(Drive、Tone、Level)。實驗數據表明,基線模型在同時處理線性與非線性任務時,易引發梯度衝突與捷徑學習;為此,本研究導入任務解耦機制實施物理特徵分流,成功將參數預測的平均絕對誤差(MAE)大幅降低 40.8%。在此基礎上,導入卷積區塊注意力模組(CBAM)探究早期融合架構的極限,測得全場最佳之參數預測 MAE(0.1623)。若針對複雜非線性頻譜進行解構,孿生網路(Siamese Network)的晚期融合策略則在頻域音色擬真度上展現優勢,將 Log-Spectral Distance (LSD) 壓低至 1.6462 dB。上述模型的推論即時因子(RTF)皆落在 0.0003 級距,確認了部署於即時系統的工程可行性。實驗末期,我們診斷出數位渲染器與類比設備間存在「多對一映射」的反問題特性。過強的控制器表徵能力會誘發「補償性過擬合」,導致模型傾向輸出帶有物理偏差的參數。儘管測試了動態損失函數排程進行探索性約束,時頻域指標間的拉扯現象依然指出了現有架構的理論限制。此一診斷確立了開發高擬真灰箱模型時必須介入更嚴格之正則化機制,也為具備物理可解釋性的 AI 音訊工作流提供了明確的後續優化方向。

    Modern music production relies heavily on electric guitar tone design. However, accurately replicating specific tones using physical or digital effects often incurs significant manual tuning costs. Recent neural audio effects, such as virtual analog and Neural Amp Modeler (NAM), predominantly utilize end-to-end black-box architectures. While these methods achieve high fidelity, they completely eliminate the control dimensions provided by physical hardware knobs. To overcome the practical limitations of existing frameworks that exclude "human-in-the-loop" workflows, this research develops an automatic tone matching system based on Differentiable Digital Signal Processing (DDSP), explicitly targeting grey-box parameter estimation for nonlinear overdrive and linear equalizers.The study first constructed a continuous parameter dataset comprising 5,865 high-quality paired audio samples of TS9 and EQ through systematic uniform sampling. For system implementation, Short-Time Fourier Transform (STFT) was utilized to extract acoustic features, combined with a Convolutional Neural Network (CNN) controller to predict corresponding physical knob configurations (Drive, Tone, Level). Experimental data indicate that baseline models are prone to gradient conflicts and shortcut learning when simultaneously handling linear and nonlinear tasks; therefore, this research introduced a task decoupling mechanism to separate physical features, successfully reducing the Mean Absolute Error (MAE) of parameter prediction by a substantial 40.8%. Building upon this framework, integrating a Convolutional Block Attention Module (CBAM) to explore the limits of early fusion architectures yielded the lowest parameter prediction MAE of 0.1623. When decomposing complex nonlinear spectra, the late fusion strategy of a Siamese Network demonstrated superior frequency-domain fidelity, bringing the Log-Spectral Distance (LSD) down to 1.6462 dB. The Real-Time Factor (RTF) for these models consistently remained at the 0.0003 level, confirming the engineering feasibility for real-time deployment. Diagnostically, the experiments revealed an ill-posed inverse problem characterized by a many-to-one mapping between the digital renderer and analog devices. Excessive controller representational capacity triggers "compensatory overfitting," causing the model to output parameters with physical deviations. Although dynamic loss scheduling was tested as an exploratory constraint, the persistent trade-off between time and frequency domain metrics highlights current theoretical limitations. This diagnosis establishes the absolute necessity of integrating stricter regularization mechanisms when designing high-fidelity grey-box models and provides a clear optimization trajectory for physically interpretable AI audio workflows.

    中文摘要 i Abstract iii 誌謝 v Contents vi List of Tables ix List of Figure x 1 Introduction 1 2 Related Work 4 2.1 Background Knowledge 4 2.1.1 Physical Nature of Linear and Nonlinear Systems in Electric Guitar Effects 4 2.1.2 Short-Time Fourier Transform and Phase Invariance 7 2.1.3 Physical Meaning of Convolutional Neural Networks in 2D Time- Frequency Representations 9 2.2 Related Work 11 2.2.1 Breakthroughs and Limitations of End-to-End Black-Box Neural Au-dio Models 11 2.2.2 Evolution of Differentiable Digital Signal Processing (DDSP) and Grey-Box Modeling 12 2.2.3 Limitations of Existing Audio Effect Datasets 14 2.3 Positioning of This Work 15 3 Methods 16 3.1 System Architecture 16 3.2 Dataset 18 3.2.1 Uniform Sampling Strategy 18 3.2.2 Signal Preprocessing and Tensor Definitions 19 3.3 Neural Parameter Controller Design 19 3.3.1 Convolutional Encoder and Dual-Channel Feature Fusion 19 3.3.2 Parameter Mapping and Multi-Layer Perceptrons (MLP) 24 3.4 DDSP Renderer 25 3.4.1 Signal Processing Chain Architecture 25 3.4.2 3.4.2 Implementation of Differentiable Frequency-domain Filtering and Numerical Stability 27 3.5 Training Strategy 28 3.5.1 Phase 1: DDSP Pre-training 28 3.5.2 Phase 2: Controller End-to-End Training 30 4 Results 33 4.1 Experimental Setup 33 4.1.1 Data Loading and Dynamic Augmentation 33 4.1.2 Hardware and Software Environment 34 4.2 Controller Objective Metrics 34 4.2.1 Parameter Estimation Capability 35 4.2.2 Frequency-Domain Metrics: Tonal Similarity 36 4.2.3 Time-Domain Waveform Fitting Metrics 39 4.3 Baseline Model and DDSP Evaluation 42 4.3.1 Physical Mapping and Fitting the Upper Limit of the DDSP White-Box Renderer 42 4.3.2 Baseline Model Performance Evaluation and System Bottleneck Di-agnosis 43 4.4 Ablation Study Results 45 4.4.1 Performance Improvements from Attention Mechanisms and Param-eter Heads 45 4.4.2 Analysis of Decoupled Architectures and Convergence Trends 46 4.4.3 Real-Time Processing Capability Evaluation 46 5 Disscusion 47 5.1 Baseline Model 48 5.2 Task Decoupling 49 5.2.1 Substantial Improvement in Parameter Prediction Accuracy and Cor-relation 49 5.2.2 Breaking ”Shortcut Learning” and Resolving Acoustic ”Over-compensation” 49 5.3 CBAM Architecture Discusion 50 5.4 Siamese Architecture Discusion 52 5.5 Summary 53 5.6 Future Work 54 6 Conclusion 56 Bibliography 58

    [1] S. Atkinson, “Generic Neural Acoustic Networks: An Overview and a Case Study in Black-Box Modeling of Audio Effects,” Proc. 23rd Int. Conf. Digit. Audio Effects (DAFx-20), Vienna, Austria, 2020.
    [2] S. Atkinson, Neural Amp Modeler, https://github.com/sdatkinson/neural-amp-modeler, Accessed: 2026-06-02, 2019.
    [3] J. Colonel and J. D. Reiss, “Reverse engineering of a recording mix with neural net-works,” Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), Toronto, ON, Canada, 2021.
    [4] M. Comunità et al., “Neural networks for guitar pedal parameter estimation,” Proc. 150th Audio Eng. Soc. Conv. (AES), 2021.
    [5] M. Comunità, C. J. Steinmetz, and J. D. Reiss, Differentiable Black-box and Gray-box Modeling of Nonlinear Audio Effects, 2025. arXiv: 2502.14405 [cs.SD].
    [6] M. Comunità, C. J. Steinmetz, and J. D. Reiss, “NablAFx: A Framework for Differen-tiable Black-box and Gray-box Modeling of Audio Effects,” arXiv preprint arXiv:2502.11668, 2025.
    [7] E. P. Damskägg, L. Juvela, E. Thuillier, and V. Välimäki, “Deep learning for tube amplifier emulation,” Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), Brighton, UK, 2019, pp. 471–475.
    [8] F. Eichas and U. Zölzer, “Black-box modeling of distortion circuits with block-oriented models,” Proc. 21st Int. Conf. Digit. Audio Effects (DAFx-18), Aveiro, Portugal, Sep. 2018.
    [9] J. Engel et al., “DDSP: Differentiable Digital Signal Processing,” Proc. Int. Conf. Learn. Represent. (ICLR), Addis Ababa, Ethiopia, 2020.
    [10] R. Geirhos, J. H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann, “Shortcut learning in deep neural networks,” Nature Machine Intelligence, vol. 2, no. 11, pp. 665–673, 2020.
    [11] S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, et al., “CNN ar-chitectures for large-scale audio classification,” Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), New Orleans, LA, USA, Mar. 2017, pp. 131–135.
    [12] S. Ioffe and C. Szegedy, “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” Proc. 32nd Int. Conf. Mach. Learn. (ICML), Lille, France, 2015, pp. 448–456.
    [13] B. Kuznetsov, J. D. Parker, and V. Välimäki, “Differentiable gray-box modeling of guitar amplifiers,” Proc. 23rd Int. Conf. Digit. Audio Effects (DAFx-20), Vienna, Aus-tria, Sep. 2020.
    [14] J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR – half-baked or well done?” Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), Brighton, UK, 2019, pp. 626–630.
    [15] P. Micikevicius et al., “Mixed precision training,” Proc. Int. Conf. Learn. Represent. (ICLR), 2018.
    [16] M. Müller, Fundamentals of Music Processing: Audio, Analysis, Algorithms, Applica-tions. Cham, Switzerland: Springer, 2015.
    [17] A. van den Oord et al., “WaveNet: A generative model for raw audio,” arXiv preprint arXiv:1609.03499, 2016.
    [18] A. V. Oppenheim and R. W. Schafer, Discrete-Time Signal Processing, 3rd. Upper Saddle River, NJ, USA: Pearson Education, 2010.
    [19] J. Pons, T. Lidy, and X. Serra, “Experimenting with musically motivated convolu-tional neural networks,” Proc. 14th Int. Workshop Content-Based Multimedia Indexing (CBMI), Bucharest, Romania, Jun. 2016, pp. 1–6.
    [20] V. Sitzmann, J. Martel, A. Bergman, D. Lindell, and G. Wetzstein, “Implicit neural representations with periodic activation functions,” Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, 2020, pp. 7462–7473.
    [21] C. G. Snoek, M. Worring, and A. W. Smeulders, “Early versus late fusion in semantic video analysis,” Proc. 13th ACM Int. Conf. Multimedia, Singapore, 2005, pp. 399–402.
    [22] M. Stein, J. Macos, and G. Schuller, “Automatic detection of audio effects in guitar and bass recordings,” Proc. 128th Audio Eng. Soc. Conv. (AES), London, UK, 2010.
    [23] C. J. Steinmetz and J. D. Reiss, “auraloss: Audio-focused loss functions in PyTorch,”Proc. Digit. Music Res. Network One-day Event (DMRN+ 15), London, UK, 2020.
    [24] S. Vandenhende et al., “Multi-task learning for dense prediction tasks: A survey,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 7, pp. 3614–3633, Jul. 2021.
    [25] Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, et al., “Tacotron: Towards end-to-end speech synthesis,” Proc. Interspeech, Stockholm, Sweden, 2017, pp. 4006–4010.
    [26] S. Woo, J. Park, J. Y. Lee, and I. S. Kweon, “CBAM: Convolutional block attention module,” Proc. Eur. Conf. Comput. Vis. (ECCV), Munich, Germany, Sep. 2018, pp. 3–19.
    [27] A. Wright, E. P. Damskägg, L. Juvela, and V. Välimäki, “Real-Time Guitar Amplifier Emulation with Deep Learning,” Applied Sciences, vol. 10, no. 3, p. 766, 2020.
    [28] A. Wright, E. P. Damskägg, and V. Välimäki, “Real-time black-box modelling with recurrent neural networks,” Proc. 22nd Int. Conf. Digit. Audio Effects (DAFx-19), Birmingham, UK, Sep. 2019.
    [29] R. Yamamoto, E. Song, and J. M. Kim, “Parallel WaveGAN: A fast waveform gener-ation model based on generative adversarial networks with multi-resolution spectro-gram,” Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), Barcelona, Spain, 2020, pp. 6199–6203.
    [30] D. T. Yeh, J. S. Abel, and J. O. Smith, “Automated physical modeling of nonlinear audio circuits for real-time audio effects—Part I: Theoretical development,” IEEE Trans. Audio, Speech, Lang. Process., vol. 18, no. 4, pp. 728–737, May 2010.
    [31] U. Zölzer, Ed., DAFX: Digital Audio Effects, 2nd. Chichester, U.K.: John Wiley & Sons, 2011.

    QR CODE