The International Audio Laboratories Erlangen (AudioLabs) are a joint institution of Fraunhofer IIS and Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU). The AudioLabs were founded in 2010 and are unique worldwide in both its mission and international approach: A team of globally-renowned scientists is working to shape the future of audio and multimedia technologies in research areas such as audio coding, audio signal analysis and perceptual spatial audio signal processing.
Notice: Starting the videos transfers usage data to YouTube.
A small neural network receives sounds that differ only in pitch and loudness, and has to describe each one with two numbers before reconstructing it. Nobody tells the network what pitch or loudness is, so whatever representation it ends up with was learned only from examples. In one chapter of his PhD thesis, Simon Schwär shows that the outcome depends on a single design choice, namely how the similarity between the original and reconstructed sounds is measured during training. Comparing the sounds sample by sample leads to a patchwork of disconnected islands in the latent space (A). A standard comparison of the frequency spectra leaves the latent space without much recognizable structure (B). Only when the comparison responds smoothly to changes in pitch, the latent space is neatly organized (C). This simple example illustrates a more general point: What a machine learns depends on the way we measure similarity.
The AudioLabs professors offer lectures in different research areas in which they share their expertise. As part of our education activities, we create interactive and multimodal lecture material. Check out an example video on time-frequency representations: