We propose the use of Non-Negative Autoencoders (NAEs) for sound deconstruction and user-guided manipulation of sounds for creative purposes. NAEs offer a versatile and scalable extension of traditional Non-Negative Matrix Factorization (NMF)-based approaches for interpretable audio decomposition. By enforcing non-negativity constraints through projected gradient descent, we obtain decompositions where internal weights and activations can be directly interpreted as spectral shapes and temporal envelopes, and where components can themselves be listened to as individual sound events. In particular, multi-layer Deep NAE architectures enable hierarchical representations with an adjustable level of granularity, allowing sounds to be deconstructed at multiple levels of abstraction—from high-level note envelopes down to fine-grained spectral details. This framework enables a wide new range of expressive, controllable, and randomized sound transformations. We introduce novel manipulation operations including cross-component and cross-layer synthesis, hierarchical deconstructions, and several randomization strategies that control timbre and event density. Through visualizations and resynthesis of practical examples, we demonstrate how NAEs can serve as flexible and interpretable tools for object-based sound editing.
Sound deconstruction with NAEs
Example of a 2-layer full deconstruction of a synthetic bells sound. [1]
Original sound
Hierarchical deconstruction
Same input sound than above (synth bells), but deconstructed hierarchically by selecting in turn every inner activation and setting the rest to zero.
Click on the activation tabs to select the corresponding hierarchical deconstruction.
Original sound
Mix of inner activation 0
Original sound
Mix of inner activation 1
Original sound
Mix of inner activation 2
Original sound
Mix of inner activation 3
Original sound
Mix of inner activation 4
Original sound
Mix of inner activation 5
Sound manipulation
Transient removal
In this example, the above bells sound has been manipulated by removing all outer components associated to inner activation 1, which
corresponds to the transient/wideband part of the sound.
Original sound
Without transients
Cross-components
In this example, new timbral contents is created via cross-components. In particular, we create a mixture of all cross-components by combining spectrum 19 above
(wideband noise) with all other activations.
Original sound
Mixture of all cross-components with spectrum 19
Random permutations
This is an example of randomly permuting all the outer activation/spectrum pairs. It results in a sound
with randomized timbral structure (often containing swapped notes), but similar temporal structure and density than the original.
[2]
Original sound
Full random permutation of the outer layer.
Weight randomization
This results in a similar timbre randomization, but an overall denser spectral content is apparent,
especially towards the end of the sound.
Original sound
Full weight randomization of the inner layer.
Controlled weight randomization
For this example, only 3 out of the 8 original inner weights are randomized. The result
is a subtler timbral change, with many of the original sound events clearly audible.