Skip to main content
archive
Search Submit Donate Log in
Press Enter to search · Advanced search

Sound

Authors and titles for recent submissions

  • Tue, 21 Jul 2026
  • Mon, 20 Jul 2026
  • Fri, 17 Jul 2026
  • Thu, 16 Jul 2026
  • Wed, 15 Jul 2026

See today's new changes

Total of 61 entries : 1-50 51-61
Showing up to 50 entries per page: fewer | more | all

Tue, 21 Jul 2026 (showing 16 of 16 entries )

[1] arXiv:2607.18190 [pdf, html, other]
Title: Audio Cross Verification Using Dual Alignment Likelihood Ratio Test
Heidi Lei, Arm Wonghirundacha, Irmak Bukey, TJ Tsai
Comments: Published at ICASSP 2023
Journal-ref: Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1-5
Subjects: Sound (cs.SD)
[2] arXiv:2607.18189 [pdf, html, other]
Title: Dense-Sparse Dynamic Time Warping for Customizing Piano Concerto Accompaniments
TJ Tsai, Kavi Dey, Yigitcan Ozer, Meinard Muller
Comments: Published at ICASSP 2025
Journal-ref: Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1-5
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[3] arXiv:2607.17900 [pdf, html, other]
Title: Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer
Shengfan Shen, Di Wu, Xingchen Song, Dinghao Zhou, Pengyu Cheng, Sixiang Lyu, Jian Luan, Shuai Wang
Subjects: Sound (cs.SD)
[4] arXiv:2607.17761 [pdf, html, other]
Title: Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
Jun Xue, Zhuolin Yi, Yanzhen Ren, Yihuan Huang, Jiayu Xiong, Yi Chai, Guanxiang Feng, Jiajun Liu, Tong Zhang
Comments: Accepted by ACM MM 2026
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI)
[5] arXiv:2607.17615 [pdf, html, other]
Title: Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture
Yuxuan Wu, Yifan Xu, Junkun Wang, Jiayong Jiang, Xin Zhao, Zhaojie Luo
Comments: Accepted by NCMMSC 2025
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI)
[6] arXiv:2607.17592 [pdf, html, other]
Title: SSTMark: Robust Training-Free Semantic-Level Speech Watermarking
Kuan-Lin Chu, Jun-Cheng Chen, Chun-Shien Lu
Subjects: Sound (cs.SD)
[7] arXiv:2607.17526 [pdf, html, other]
Title: FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration
Ali Boudaghi, Hadi Zare
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Signal Processing (eess.SP); Numerical Analysis (math.NA)
[8] arXiv:2607.17098 [pdf, html, other]
Title: Multi-Level Privacy-Preserving Dementia Detection from Speech via Targeted Adversarial Obfuscation and Representation Learning
Henriette Flore Kenne, Raphael Anaadumba, Mohammad Arif Ul Alam
Comments: Accepted
Journal-ref: Interspeech 2026
Subjects: Sound (cs.SD); Cryptography and Security (cs.CR)
[9] arXiv:2607.16870 [pdf, html, other]
Title: Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
Ye Lu, Yihan Yan, Zhaoyang Zhang, Zhitao Ou, Runze Liu, Li Liu, Shen Wang
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
[10] arXiv:2607.16803 [pdf, html, other]
Title: Explainable Lightweight Compact Deep Models for Speech Emotion Recognition
Nelly Elsayed
Comments: Accepted in the IEEE ICMLA 2026 Conference
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET); Machine Learning (cs.LG)
[11] arXiv:2607.16657 [pdf, html, other]
Title: HARP: Harmonic-Aware Residual Partitioning for Neural Audio Codecs
Qiaoyu Yang, Lixing He, Binyue Deng, Weifeng Zhao
Comments: Accepted to Interspeech 2026
Subjects: Sound (cs.SD)
[12] arXiv:2607.16599 [pdf, html, other]
Title: Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves
Yishan Lv, Jing Luo, Xinyu Yang, Zhizheng Wu
Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[13] arXiv:2607.16369 [pdf, html, other]
Title: Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection
André Runewicz, Karla Schäfer, Martin Steinebach
Comments: Accepted to 2026 ICME workshop
Subjects: Sound (cs.SD)
[14] arXiv:2607.16980 (cross-list from eess.SP) [pdf, html, other]
Title: Efficient Audio-Visual Event Recognition via Knowledge Distillation and Dynamic INT8 Quantization of a Hybrid Cross-Attention Network
Parinaz Binandeh Dehaghani, Danilo Pena, A. Pedro Aguiar
Comments: 15 pages, 4 figures
Subjects: Signal Processing (eess.SP); Sound (cs.SD)
[15] arXiv:2607.16736 (cross-list from eess.AS) [pdf, html, other]
Title: RealDESED: A Real-World Domestic Sound Event Detection Benchmark
Florian Schmid, Paul Primus, Alexander Fichtinger, Tara Jadidi, Tobias Morocutti, Gerhard Widmer
Comments: Submitted to the DCASE 2026 Workshop (Detection and Classification of Acoustic Scenes and Events). Resources: Dataset (Zenodo): this https URL code and baseline implementation (GitHub): this https URL
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Sound (cs.SD)
[16] arXiv:2607.16220 (cross-list from cs.CY) [pdf, html, other]
Title: Comparing Spectrogram Front-Ends for Abnormal Heart-Sound Detection with a Convolutional Neural Network
Abhinav Pala, Dhanush Pala
Subjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)

Mon, 20 Jul 2026 (showing 10 of 10 entries )

[17] arXiv:2607.15755 [pdf, html, other]
Title: AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
Zhenqi Jia, Yuan Zhao, Aruukhan, Rui Liu, Haizhou Li
Comments: Accepted by ACM MM 2026
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI)
[18] arXiv:2607.15697 [pdf, html, other]
Title: SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models
Jinwen Xin, Xixiang Lv
Comments: 8 pages
Journal-ref: 2024 International Joint Conference on Neural Networks (IJCNN), pp. 1-8, 2024
Subjects: Sound (cs.SD); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
[19] arXiv:2607.15634 [pdf, html, other]
Title: StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems
Yuan-Chiao Cheng, Jui-Te Wu, Brian Chen, Yen-Tung Yeh, Yu-Hua Chen, Yi-Hsuan Yang
Comments: Accepted to ISMIR 2026. 8 pages, 4 figures
Subjects: Sound (cs.SD)
[20] arXiv:2607.15475 [pdf, html, other]
Title: Segmental DTW: A Parallelizable Alternative to Dynamic Time Warping
TJ Tsai
Comments: Published at ICASSP 2021
Journal-ref: Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 106-110
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[21] arXiv:2607.15443 [pdf, html, other]
Title: Estimating the Reliability of Dynamic Time Warping Alignments Using Circumstantial Evidence
Aanya Pratapneni, Alice Yuan, TJ Tsai
Comments: Accepted at ISMIR 2026
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[22] arXiv:2607.15724 (cross-list from cs.CR) [pdf, html, other]
Title: Natural Backdoor Attacks on Speech Recognition Models
Jinwen Xin, Xixiang Lyu, Jing Ma
Comments: This is the authors' manuscript of a chapter published in Machine Learning for Cyber Security, Lecture Notes in Computer Science, vol. 13655, pp. 597-610 (2023)
Journal-ref: Machine Learning for Cyber Security, Lecture Notes in Computer Science, vol. 13655, pp. 597-610 (2023)
Subjects: Cryptography and Security (cs.CR); Machine Learning (cs.LG); Sound (cs.SD)
[23] arXiv:2607.15694 (cross-list from eess.AS) [pdf, html, other]
Title: A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors
Shuhei Kato
Comments: 24 pages, 6 figures, 8 tables. Submitted to IEEE Access
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[24] arXiv:2607.15478 (cross-list from cs.DS) [pdf, html, other]
Title: A Study of Parallelizable Alternatives to Dynamic Time Warping for Aligning Long Sequences
Daniel Yang, Thaxter Shaw, TJ Tsai
Comments: Published in IEEE/ACM Transactions on Audio, Speech, and Language Processing
Journal-ref: IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 2117-2127, 2022
Subjects: Data Structures and Algorithms (cs.DS); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[25] arXiv:2607.15298 (cross-list from eess.IV) [pdf, html, other]
Title: Data-driven Video Codec with Implicit Neural Representations
Nishan Khanal, Saugat Neupane, Abhinav Chalise, Nimesh Gopal Pradhan, Dinesh Baniya Kshatri
Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Sound (cs.SD)
[26] arXiv:2607.15295 (cross-list from cs.MM) [pdf, html, other]
Title: AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin
Subjects: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD)

Fri, 17 Jul 2026 (showing 9 of 9 entries )

[27] arXiv:2607.14846 [pdf, html, other]
Title: RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr Cłapa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis
Comments: Benchmark and leaderboard: this https URL
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI)
[28] arXiv:2607.14753 [pdf, html, other]
Title: Large Audio Language Models for Spoofing-Aware Speaker Verification
Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir, Artem Dvirniak, Dmitrii Korzh, Oleg Y. Rogov
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI)
[29] arXiv:2607.14537 [pdf, html, other]
Title: MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music
Scott H. Hawley
Comments: 8 pages, 8 figures
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[30] arXiv:2607.14474 [pdf, html, other]
Title: Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026
Anthony Miyaguchi, Murilo Gustineli, Adrian Cheung
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
[31] arXiv:2607.14148 [pdf, html, other]
Title: ITGPT: A Transformer Based Architecture for the Generation of Dance Dance Revolution and In the Groove Charts
Miguel O'Malley
Comments: 14 pages, 11 figures, 2 tables + appendix
Subjects: Sound (cs.SD); Machine Learning (cs.LG)
[32] arXiv:2607.15265 (cross-list from cs.CV) [pdf, html, other]
Title: SceneBind: Binding What and Where Across Vision, Audio and Language
Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman
Comments: Project website: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Sound (cs.SD)
[33] arXiv:2607.15243 (cross-list from eess.AS) [pdf, html, other]
Title: What does the model actually see? Evaluation protocols and input availability in data-driven prediction of room acoustic parameters
Akın Oktav
Comments: 12 pages, 4 figures. Submitted to Acta Acustica
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[34] arXiv:2607.15198 (cross-list from eess.AS) [pdf, html, other]
Title: SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
Shuai Wang, Zihan Qian, Ke Zhang, Jiangyu Han, Zikai Liu, Xiaoyang Yu, Haoyu Li, Marc Delcroix, Kai Yu, Lei Xie, Ming Li, Haizhou Li
Comments: Overview paper of Real-TSE Challenge
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[35] arXiv:2607.14189 (cross-list from cs.CV) [pdf, html, other]
Title: MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang, Yiting He, Lizhuo Shao, Weiming Zhu, Tengfei Liu, Yang Shi, Jialu Chen, Yuanxing Zhang, Huaxiong Li
Comments: 32 pages
Subjects: Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)

Thu, 16 Jul 2026 (showing 11 of 11 entries )

[36] arXiv:2607.13903 [pdf, html, other]
Title: Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation
Yizhou Zhang, Wangjin Zhou, Yi Zhao, Wei Tan, Keisuke Imoto, Zhi Gong
Comments: Accept by ISMIR 2026
Subjects: Sound (cs.SD)
[37] arXiv:2607.13864 [pdf, html, other]
Title: Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?
Wangjin Zhou, Yizhou Zhang, Yichi Wang, Tatsuya Kawahara
Comments: Accept by Interspeech 2026
Subjects: Sound (cs.SD)
[38] arXiv:2607.13840 [pdf, html, other]
Title: From Continuous Deployment to Queryable Dataset: Terabyte-Scale AIS-Aligned Passive Acoustic Labelling
Wayne Renaud, Priyanka Aravindan, Gabriel Spadon
Comments: OCEANS'26 - Monterey
Subjects: Sound (cs.SD); Databases (cs.DB)
[39] arXiv:2607.13587 [pdf, html, other]
Title: From Prediction to Collaboration: Interactive Symbolic Music Analysis
Emmanouil Karystinaios, Johannes Hentschel, Markus Neuwirth, Gerhard Widmer
Comments: in Proceedings of the 27th International Society for Music Information Retrieval Conference (ISMIR) 2026
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI)
[40] arXiv:2607.13477 [pdf, html, other]
Title: Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation
Joonyong Park, David M. Chan, Yuki Saito, Hiroshi Saruwatari
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[41] arXiv:2607.13278 [pdf, html, other]
Title: Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
Ben Maman, Frank Zalkow, Hans-Ulrich Berendes, Paolo Sani, Christian Dittmar, Meinard Müller
Comments: Accepted to International Conference on Digital Audio Effects (DAFx) 2026
Subjects: Sound (cs.SD)
[42] arXiv:2607.14072 (cross-list from cs.LG) [pdf, other]
Title: MetaPerch: Learning from metadata for bioacoustics foundation models
Mustafa Chasmai, Vincent Dumoulin, Jenny Hamer
Comments: Accepted to ICML 26
Subjects: Machine Learning (cs.LG); Sound (cs.SD)
[43] arXiv:2607.13571 (cross-list from eess.AS) [pdf, html, other]
Title: Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification
Shiqi Zhang, Tuomas Virtanen
Comments: submitted to DCASE workshop 2026, under reviewing
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Sound (cs.SD)
[44] arXiv:2607.13555 (cross-list from eess.AS) [pdf, html, other]
Title: Greedy Volume Maximization of Gradient Embeddings for Long-Tailed Frame-Level Bioacoustic Active Learning
Shiqi Zhang, Marius Faiß, Ariana Strandburg-Peshkin, Tuomas Virtanen
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Sound (cs.SD)
[45] arXiv:2607.13471 (cross-list from cs.CV) [pdf, html, other]
Title: Bring Music The Horizon: Music-Driven 360$^\circ$ Video Generation
Kai Hsu Tsai, Yong Wei Fu, Hung I Yang, Yu-Chih Chen
Comments: 5 pages, 1 figure
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)
[46] arXiv:2607.13408 (cross-list from eess.AS) [pdf, html, other]
Title: Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim, Suyoun Kim, Bo-Ru Lu, Qinming Tang, Ankur Gandhe, Hung-yi Lee, Chieh-Chi Kao, Chao Wang
Comments: Accepted to the Long Paper Track at Interspeech 2026
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)

Wed, 15 Jul 2026 (showing first 4 of 15 entries )

[47] arXiv:2607.12872 [pdf, html, other]
Title: Low-Latency Neural Models for Real-Time Music Enhancement
Emmanouil Karystinaios, Jonathan Greif, David Nadrchal, Paul Primus, Gerhard Widmer
Subjects: Sound (cs.SD)
[48] arXiv:2607.12857 [pdf, html, other]
Title: ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation
Jhen-Ke Lin
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI)
[49] arXiv:2607.12725 [pdf, html, other]
Title: Neural Morphing: Sequence-Optimized Token-Level Morphing in Neural Audio Codecs
Emmanouil Karystinaios
Comments: In proceedings of the 29th International Conference on Digital Audio Effects (DAFx) 2026
Subjects: Sound (cs.SD)
[50] arXiv:2607.12706 [pdf, html, other]
Title: AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling
Haowei Lou, Junda Wu, Chengkai Huang, Tong Yu, Hye-young Paik, Wen Hu, Lina Yao
Subjects: Sound (cs.SD)
Total of 61 entries : 1-50 51-61
Showing up to 50 entries per page: fewer | more | all
We gratefully acknowledge support from our major funders, member institutions, , and all contributors.
About · Help · Contact · Subscribe · Copyright · Privacy · Accessibility · Operational Status (opens in new tab)
Major funding support from
Simons Foundation Simons Foundation International Schmidt Sciences