Abstract
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.
Community
GigaChat Audio 10B (A1.8B) — an open-source, MIT-licensed audio-native LLM for speech and long-form audio understanding. It combines the GigaAM encoder with a 10B MoE decoder using 1.8B active parameters. The model supports ASR, translation, audio QA, emotion recognition, and temporal grounding for recordings up to two hours. With inter-timings—periodic temporal anchors embedded into the audio stream—it reaches 48.3 mIoU on temporal localization, compared with approximately 0.0 for other open-source models.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- MOSS-Audio Technical Report (2026)
- VoiceGiraffe: A Benchmark for Extreme Long-Context Audio-Language Understanding (2026)
- MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs (2026)
- REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing (2026)
- Empowering Long-form Omni-modal Understanding with Robust Audio Perception (2026)
- Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding (2026)
- MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend