Enhancing Advanced Audio Understanding Through Large Language Models – UROP Spring Symposium 2025

Enhancing Advanced Audio Understanding Through Large Language Models

Antonio Capdevielle

Research Mentor(s): Alanson Sample
Mentor Department: Computer Science and Engineering
Authors: Antonio Capdevielle, Kaylee Yixuan, Alanson Sample
Session: Session 1 (9:00am – 9:50am)
Presentation Type: Poster 68

Abstract

The ability to accurately interpret and describe sound events has significant implications for various fields, including accessibility technology, smart home systems, and automated content generation. Traditional audio classification models typically assign simple labels to sound events, limiting their contextual understanding. This is especially apparent in complex acoustic environments with overlapping sounds and background noise. Our research aims to bridge this gap by integrating advanced audio encoders with large language models (LLMs) to generate rich, context-aware descriptions of audio events. To achieve this, we are developing, training, and evaluating multiple audio encoders, including AST, CavMAE, and GAMA, to determine the most effective approach for extracting meaningful representations of sound. These embeddings are then processed by an LLM to generate natural language descriptions that capture not only individual sound events but also their relationships and contextual significance. Our models are trained and fine-tuned on datasets such as AudioSet, ESC-50, and custom-labeled recordings to ensure robust performance across diverse real-world conditions. Preliminary results indicate that AST models trained on curated datasets outperform other approaches in structured audio representation, achieving high accuracy and coherence in generated descriptions. Looking forward, we aim to refine our human evaluation methodology, optimize inference efficiency, and enhance model generalization in noisy and dynamic environments. By pushing the boundaries of audio-language integration, this research contributes to a more advanced and intuitive understanding of sound, enabling more sophisticated applications in assistive technology, security, and human-computer interaction.

lsa logoum logo