Download Eigensystem Realization of Violin Bridge Admittances
Modeling violin bridge admittance is a long-standing problem in musical acoustics, with applications in sound analysis, synthesis, and virtual instrument design. In this work, we investigate the use of the Eigensystem Realization Algorithm (ERA) for deriving reduced-order state-space models directly from measured impulse responses. The proposed approach allows us to extract dominant system dynamics and obtain compact realizations without requiring explicit modal parameterization. We evaluate ERA on a dataset of modern and historical violins and compare it against established modal and state-space identification methods. Experimental results demonstrate that ERA outperforms existing approaches by achieving lower reconstruction errors in both the time and frequency domains while preserving perceptually relevant characteristics of the bridge response. Furthermore, we show that the state-space realizations obtained using ERA reproduce the target frequency-dependent energy decay more accurately than models obtained using the baseline methods. These findings support the use of ERA as an efficient and flexible alternative for modeling violin bridge admittances, with applications that span from audio synthesis and processing to instrument virtualization.
Download Measurement-Informed Nonlinear Modal Synthesis of 65 Classical Guitars
When a classical guitar string is plucked, vibration energy flows through the bridge into the body and is radiated as sound. Synthesising this process for a large collection of instruments requires both an efficient nonlinear string model and a robust method for extracting instrument-specific parameters from measurements. This paper addresses both issues. Starting from the publicly available dataset of Mores, which provides impulse-response measurements on 65 classical guitars, modal parameters of the bridge compliance and of the bridge-to-air radiation path are extracted for each instrument. These feed a nonlinear string model in which transverse vibration is governed by a geometrically exact elastic potential coupled at an interior bridge point to the measured body data. The nonlinear potential is quadratised via the Scalar Auxiliary Variable (SAV) method, so that the equations of motion become linear in a scalar variable and a known gradient vector, even at the continuous level. After time discretisation, the coupled system is inverted through two sequential Sherman–Morrison rank-one updates (one for the bridge coupling, one for the SAV nonlinearity), yielding an O(N) algorithm per time step. Two regularisation techniques prevent long-term drift of the auxiliary variable. The complete pipeline is demonstrated by synthesising plucked notes across all frets and strings for each of the 65 guitars.
Download Differentiable Articulatory Copy-Synthesis of Biphonic Singing
Sygyt is a Tuvan style of biphonic singing in which a low vocal drone is sustained while a high harmonic is selectively amplified in the 1–3 kHz region. Copy-synthesizing this effect remains challenging for articulatory models, since it requires fine control of narrowly focused resonances that standard low-dimensional tract parameterizations cannot easily reproduce. We address this problem with a differentiable Kelly–Lochbaum waveguide augmented with a sublingual second source, cubic B-spline tract parameterization, and spatially varying learnable damping, optimized end-to-end by gradient descent from audio. On 20 segments from two independent sygyt datasets (5 singers, 10 pitches), the proposed model reduces log-spectral distance by 30–38% relative to an articulatory baseline, with the largest gains concentrated in the overtone region. Cepstral-envelope analysis further shows more accurate recovery of the merged formant structure characteristic of sygyt production. The model also outperforms a DDSP harmonic-plus-noise baseline with direct per-harmonic spectral control, suggesting that explicit acoustic structure is a useful inductive bias for overtone-singing copy-synthesis.
Download Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
Recent diffusion-based generative models have achieved strong results in domain-specific audio generation tasks such as speech, singing, and instrumental music synthesis. However, these models are typically specialized and do not generalize well to mixed or intermediate audio types. In this work, we adapt a diffusion-based model originally designed for multi-instrument music synthesis to voice conversion, covering both speech and singing within a unified framework. Specifically, we extend musical note-based conditioning to include phonetic posteriorgrams (PPGs) and pitch contours, and reinterpret timbre conditioning as speaker or singer identity via feature-wise linear modulation. Experiments show that the adapted model matches or surpasses a dedicated voice conversion system in terms of naturalness and performer similarity, while maintaining accurate pitch control across speech and singing. At the same time, we observe limitations in phonetic fidelity and a degradation in vocal quality when incorporating instrumental training data. Furthermore, we demonstrate that off-the-shelf feature extractors provide effective conditioning signals, enabling large-scale self-supervised training without manual annotations. These results highlight the potential of cross-domain model transfer towards unified audio generation systems capable of handling speech, singing, and music.
Download Explicit Wave Digital Model of the Fulltone OCD Pedal Based on Canonical Piecewise-Linear Functions
Virtual Analog (VA) modeling aims at digitally emulating analog audio equipment while preserving its characteristic nonlinear behavior and musical expressiveness. In the context of guitar effects, overdrive pedals represent a cornerstone of many signal chains, as they strongly contribute to the perceived dynamics, articulation, and timbral identity of the instrument. Among these, the Fulltone OCD overdrive is considered a standard in both studio and live environments, being widely adopted across rock and metal genres. In this article, we present an explicit Wave Digital (WD) model of the Fulltone OCD (v2) pedal. By exploiting the circuit topology, the MOSFETs and the germanium diode composing the asymmetric clipping stage are grouped into a single equivalent nonlinear element, enabling an explicit WD realization that avoids costly iterative solvers. The resulting nonlinear characteristic is approximated by means of a Canonical Piecewise-Linear (CPWL) function, yielding a compact and efficient explicit model suitable for real-time implementation. The proposed model is validated against reference simulations and implemented both in MATLAB and as a real-time audio plug-in using the JUCE framework.
Download Fourier Neural Operators for Sample-Rate-Independent Virtual Analog Modeling
Neural networks that operate directly on time-domain signals are widely used for virtual analog (VA) modeling. A key limitation of these models is their dependence on the sampling rate used during training, which becomes implicitly encoded in the learned parameters, so that changing it generally alters the realized dynamics. Although architectural modifications to recurrent neural networks have been proposed to enable sample-rate independent operation, these approaches are inherently tailored to upsampling and do not accommodate downsampling scenarios. In this manuscript, we present a VA modeling framework based on Fourier Neural Operators (FNOs) adapted to process fixed-duration audio frames. The proposed formulation defines the learned mapping over a fixed temporal support and evaluates it on uniform grids of different densities, so that a model trained at a single sampling rate can be applied at unseen sampling resolutions. Numerical results on a nonlinear transistor circuit show that the proposed model achieves competitive accuracy in upsampling scenarios while remaining directly applicable to downsampling, unlike a sample-rate independent baseline recurrent architecture.
Download Deep Regularized RNNs for Virtual Analog
Virtual analog (VA) modeling methods seek to emulate analog audio hardware using digital signal processing (DSP). Modeling approaches fall into three broad categories: white-box methods, which use detailed device knowledge for accurate simulation; gray-box methods that use generic DSP blocks to model the system; and black-box methods, which rely solely on opaque models learned from input–output data. A category of architectures used widely in black-box modeling are recurrent neural networks (RNNs). To model device controls, the control values can be provided as conditioning input to the network. However, when the conditioning is time-varied, the models are susceptible to producing noise artifacts. Regularization of the RNN dynamics significantly reduces these artifacts, though at a loss in modeling accuracy. This paper closes the dynamics regularization quality gap by introducing deep control-conditioned LSTMs and a gammatone filterbank (GFB) loss. Experiments indicate that the proposed method achieves comparable modeling performance as unregularized baselines while avoiding the noise artifacts caused by time-varying control inputs.
Download FM Parameter Estimation with Low-Order Rational Constraints on Wasserstein Loss Landscape
Frequency modulation (FM) synthesis has been widely used in music production and sound design due to its ability to generate rich timbres with few control parameters. However, estimating the frequency parameters from a target sound remains challenging because different parameter configurations can yield similar spectra, creating numerous local minima in the loss landscape. In this paper, we analyze the Wasserstein distance loss landscape for two-operator FM synthesis under practical FFT-based spectral representations and show that it exhibits non-differentiable ridges at rational frequency ratios, arising from negative-frequency folding and spectral ordering transitions. Exploiting this structure, we propose a constrained gradient-based optimization strategy that constrains the frequency ratio in each optimization run to an interval bounded by consecutive low-order rational ratios and retains the lowest-loss candidate across intervals. Experimental results from controlled ablations show that maintaining the constraint throughout optimization improves reliability over random initialization and initialization-only constraints, particularly for more complex spectra at higher modulation indices.
Download Enhancing Automatic Chord Recognition via Pseudo-Labeling and Knowledge Distillation
Automatic Chord Recognition (ACR) is constrained by the scarcity of aligned chord annotations, which are costly to acquire. At the same time, open-weight pre-trained models are more accessible than their proprietary training data. In this work, we present a two-stage training pipeline that leverages pre-trained models together with unlabeled audio. The proposed method decouples training into two stages. In the first stage, we use the pre-trained BTC model as a teacher to generate pseudo-labels for over 1,000 hours of diverse unlabeled audio and train a student model solely on these pseudo-labels. In the second stage, the student is continually trained on ground-truth labels as they become available. To prevent catastrophic forgetting of the representations learned in the first stage, we apply selective knowledge distillation (KD) from the teacher as a regularizer. In our experiments, two models (BTC, 2E1D) were used as students. In Stage 1, using only pseudo-labels, the BTC student achieves about 99% of the teacher's performance, while the 2E1D model achieves about 97% of the teacher's performance across seven standard mir_eval metrics. After continual training with labeled data in Stage 2, the resulting BTC student model consistently surpasses both the traditional supervised learning baseline and the original pre-trained teacher model across all metrics. The resulting 2E1D student model also outperforms the supervised baseline and approaches teacher-level performance, with both models demonstrating substantial gains on rare chord qualities.
Download A Perceptually Inspired Single Parameter Auditory Distance Renderer for Music Production
Conveying auditory distance in a digital audio workstation requires balancing several uncoupled tools (reverb, gain, equalization, pre-delay) by hand, a workflow that is cognitively demanding and easily produces spatially incoherent results. We demonstrate a real-time VST3 plugin that derives five correlated distance cues from a single normalized control and keeps them mutually coherent by construction, grounded in the psychoacoustics of auditory distance perception. A headphone listening test with 18 participants showed that the coupled renderer roughly halves distance placement error relative to an uncoupled manual mix and was unanimously preferred on composite spatial quality. Users sweep one knob and hear sources move convincingly from near to far on multitrack material, compare the result against a manual uncoupled mix, and toggle an optional binaural externalization stage. Source code is available on GitHub.