Download SCHAEFFER: A Dataset of Human-Annotated Sound Objects for Machine Learning Applications
Machine learning for sound generation is rapidly expanding within the computer music community. However, most datasets used to train models are built from field recordings, foley sounds, instrumental notes, or commercial music. This presents a significant limitation for composers working in acousmatic and electroacoustic music, who require datasets tailored to their creative processes. To address this gap, we introduce the SCHAEFFER Dataset (Spectromorphological Corpus of Human-annotated Audio with Electroacoustic Features For Experimental Research), a curated collection of 1000 sound objects designed and annotated by composers and students of electroacoustic composition. The dataset, distributed under Creative Commons licenses, features annotations combining technical and poetic descriptions, alongside classifications based on pre-defined spectromorphological categories.
Download Pitch-Conditioned Instrument Sound Synthesis From an Interactive Timbre Latent Space
This paper presents a novel approach to neural instrument sound synthesis using a two-stage semi-supervised learning framework capable of generating pitch-accurate, high-quality music samples from an expressive timbre latent space. Existing approaches that achieve sufficient quality for music production often rely on highdimensional latent representations that are difficult to navigate and provide unintuitive user experiences. We address this limitation through a two-stage training paradigm: first, we train a pitchtimbre disentangled 2D representation of audio samples using a Variational Autoencoder; second, we use this representation as conditioning input for a Transformer-based generative model. The learned 2D latent space serves as an intuitive interface for navigating and exploring the sound landscape. We demonstrate that the proposed method effectively learns a disentangled timbre space, enabling expressive and controllable audio generation with reliable pitch conditioning. Experimental results show the model’s ability to capture subtle variations in timbre while maintaining a high degree of pitch accuracy. The usability of our method is demonstrated in an interactive web application, highlighting its potential as a step towards future music production environments that are both intuitive and creatively empowering: https://pgesam.faresschulz.com/.
Download Neural Sample-Based Piano Synthesis
Piano sound emulation has been an active topic of research and development for several decades. Although comprehensive physicsbased piano models have been proposed, sample-based piano emulation is still widely utilized for its computational efficiency and relative accuracy despite presenting significant memory storage requirements. This paper proposes a novel hybrid approach to sample-based piano synthesis aimed at improving the fidelity of sound emulation while reducing memory requirements for storing samples. A neural network-based model processes the sound recorded from a single example of piano key at a given velocity. The network is trained to learn the nonlinear relationship between the various velocities at which a piano key is pressed and the corresponding sound alterations. Results show that the method achieves high accuracy using a specific neural architecture that is computationally efficient, presenting few trainable parameters, and it requires memory only for one sample for each piano key.
Download Piano-SSM: Diagonal State Space Models for Efficient Midi-to-Raw Audio Synthesis
Deep State Space Models (SSMs) have shown remarkable performance in long-sequence reasoning tasks, such as raw audio classification, and audio generation. This paper introduces PianoSSM, an end-to-end deep SSM neural network architecture designed to synthesize raw piano audio directly from MIDI input. The network requires no intermediate representations or domainspecific expert knowledge, simplifying training and improving accessibility. Quantitative evaluations on the MAESTRO dataset show that Piano-SSM achieves a Multi-Scale Spectral Loss (MSSL) of 7.02 at 16kHz, outperforming DDSP-Piano v1 with a MSSL of 7.09. At 24kHz, Piano-SSM maintains competitive performance with an MSSL of 6.75, closely matching DDSP-Piano v2’s result of 6.58. Evaluations on the MAPS dataset achieve an MSSL score of 8.23, which demonstrates the generalization capability even when training with very limited data. Further analysis highlights Piano-SSM’s ability to train on high sampling-rate audio while synthesizing audio at lower sampling rates, explicitly linking performance loss to aliasing effects. Additionally, the proposed model facilitates real-time causal inference through a custom C++17 header-only implementation. Using an Intel Core i712700 processor at 4.5GHz, with single core inference, allows synthesizing one second of audio at 44.1kHz in 0.44s with a workload of 23.1GFLOPS/s and an 10.1µs input/output delay with the largest network. While the smallest network at 16kHz only needs 0.04s with 2.3GFLOP/s and 2.6µs input/output delay. These results underscore Piano-SSM’s practical utility and efficiency in real-time audio synthesis applications.
Download A Statistics-Driven Differentiable Approach for Sound Texture Synthesis and Analysis
In this work, we introduce TexStat, a novel loss function specifically designed for the analysis and synthesis of texture sounds characterized by stochastic structure and perceptual stationarity. Drawing inspiration from the statistical and perceptual framework of McDermott and Simoncelli, TexStat identifies similarities between signals belonging to the same texture category without relying on temporal structure. We also propose using TexStat as a validation metric alongside Frechet Audio Distances (FAD) to evaluate texture sound synthesis models. In addition to TexStat, we present TexEnv, an efficient, lightweight and differentiable texture sound synthesizer that generates audio by imposing amplitude envelopes on filtered noise. We further integrate these components into TexDSP, a DDSP-inspired generative model tailored for texture sounds. Through extensive experiments across various texture sound types, we demonstrate that TexStat is perceptually meaningful, time-invariant, and robust to noise, features that make it effective both as a loss function for generative tasks and as a validation metric. All tools and code are provided as open-source contributions and our PyTorch implementations are efficient, differentiable, and highly configurable, enabling its use in both generative tasks and as a perceptually grounded evaluation metric.
Download Ambisonic Decoder Equalization in Reverberant Environments via Closed-Hull Crosstalk Inversion
In this work, the authors develop a higher order Ambisonic (HOA) decoder that compensates for listening room reverberation by cascading a conventional decode matrix with a crosstalk matrix derived from room impulse responses (RIR) constrained over the listener area, namely by sampling over a spherical boundary according to HOA convention. A decoder is generated for simulated RIRs of mixed specular and diffuse reverberation, and spectral and spatiotemporal energy distributions are shown for directional and diffuse HOA signals rendered through the reverberant room. Results demonstrate that the decoder renders temporally-variant directional sources with higher directivity as compared to a conventional decoder.ing Ambisonic decoders in reverberant environments using closed-hull crosstalk convolution. By modeling the acoustic path from each loudspeaker to the listener's ears via measured room impulse responses, we formulate equalization filters that compensate for room-induced distortions in the decoded signals. The proposed closed-hull approach constrains the solution to preserve perceptually relevant spatial cues while minimizing spectral coloration. Listening tests demonstrate improved externalization, timbral transparency, and localization accuracy compared to standard free-field decoding in reverberant conditions.
Download Perceptual Optimisation of Loudspeaker-Based Reproduction
This paper proposes POLAR, a framework for the optimisation of loudspeaker signals using end-to-end differentiable perceptual loss functions. The framework optimises multiple perceptual attributes across multiple listeners, offering a versatile method for a range of problems. This versatility stems from the ability to customise the number of loudspeakers, listeners, and the weighting applied to different perceptual attributes. Here, we apply the method to four problems: (a) source panning for a single listener in stereo reproduction, (b) single-listener colouration matching in stereo reproduction, (c) extended sweet spot using stereo pairs beamforming, and (d) multi-attribute perceptually driven panning in stereo reproduction. The first three problems are evaluated against solutions traditionally used for these tasks: solutions of (a) are shown to be similar to those obtained with tangent panning law and vector-base amplitude panning (VBAP), solutions of (b) are shown to be similar to those obtained for cross-talk cancellation, and solutions of (c) are shown to be similar to those obtained in earlier work on directivity pattern optimisation for sweet spot widening. Each of these solutions was previously obtained using fundamentally different methodologies, demonstrating the flexibility and broad applicability of the proposed framework.
Download Multi-Source Extension and Hyperparameter Optimization of the DiffRIR Framework for Room Impulse Response Synthesis
Efficient prediction of Room Impulse Responses (RIRs) is a cornerstone for immersive virtual acoustics and scalable room acoustic modeling. This study extends the DiffRIR framework – proposed by Wang et al. in Hearing Anything Anywhere – by introducing a multi-source training logic and systematically optimizing its convergence behavior to overcome the inherent limitations of the original framework. Our results reveal that multi-source training acts as implicit data augmentation, where the resulting increase in spatial entropy enhances the model's spectral accuracy. Furthermore, we demonstrate that the model exhibits remarkable robustness against geometric inaccuracies, maintaining numerical stability even with source positional offsets of up to 4 m in single-source baseline evaluations. By identifying a learning rate of 3×10⁻², we were able to reduce the training duration to 23% of the original baseline without compromising prediction accuracy. While the increased complexity of multi-source fields necessitates a trade-off in temporal precision – quantified via our newly integrated Energy Decay Convergence (EDC) metric – this research provides an efficient and resilient solution for acoustic simulations in complex environments.
Download Parametric Resynthesis of Measured Spatial Room Impulse Responses
Spatial Room Impulse Responses (SRIRs) are fundamental to immersive audio rendering and have become a key focus of recent machine learning research in acoustics and auralization. Due to the high computational cost of direct convolution, spatial audio systems commonly employ artificial reverberation algorithms. However, these approaches often fail to accurately reproduce the spatial, temporal, and spectral characteristics of early reflections, leading to notable deviations from measured SRIRs. This paper presents a comprehensive framework for the analysis and efficient resynthesis of SRIRs captured with Spherical Microphone Arrays (SMAs). The proposed method accounts for hardware-induced artifacts, including scattering and spatial aliasing. Early reflections are reconstructed using a parametric approach based on the Herglotz analysis method, while late reverberation is synthesized using a Directional Feedback Delay Network (DFDN) with optimized filter-attenuation and correlation-matching. The proposed framework produces signals whose spatial correlation and Energy Decay Relief (EDR) closely match those of measured SRIRs, demonstrating its effectiveness for both real-time spatial audio rendering and realistic dataset generation for machine learning applications.
Download Physical Model of the Chinese Yehu for Sound Synthesis
The yehu is a Chinese bowed string instrument featuring a resonator carved from a coconut shell, a seashell-based bridge, and two silk strings. This paper proposes a physical model of the yehu and reports on simulations using a finite-difference scheme with measurement-based physical characterization. The proposed model consists of two stiff strings coupled at the bridge, a bow with elastic bow hairs, a stopping finger, and a modal model of the bridge. A non-iterative solver based on energy quadratization is used to model the finger–string contact force, while an iterative solver is used for elasto-plastic bow-string friction force. The bridge-body model is based on a modal characterization obtained from the measured bridge admittance. The measured radiation transfer function is represented as a bank of parallel second-order filters and is applied to the simulated bridge force to incorporate body radiation characteristics. Finally, computational performance tests are conducted, showing that the proposed model is capable of real-time computation.