From Speech to Underwater Acoustics: A Transfer Learning Framework for Real-Time Passive Diver Detection Using Keyword Spotting Models
Abstract
Passive acoustic detection of divers faces challenges such as low signal-to-noise ratios (SNRs), data scarcity, and latency in conventional methods. This paper proposes keyword spotting for diver detection (KWS-DD) – a transfer learning framework that repurposes speech-oriented KWS models for data-efficient diver detection. Diver inhalation signatures are treated as acoustic ‘keywords,’ enabling adaptation of the transformer-based HuBERT architecture (pre-trained on speech) to identify quasi-periodic respiratory events in underwater audio. The core innovation of this work lies in adapting the state-of-the-art speech model HuBERT for accurate diver detection via non-speech inhalation acoustics. This approach eliminates the need for accumulation during the respiratory cycle, enabling real-time detection using a minimal amount of domain-specific data (120 inhalation samples). The solution, deployed in diverse marine conditions, achieved 94.4 % accuracy and 94.6 % F1-score for inhalation sounds. This represents a more than 50 % range extension over conventional methods, which proved unreliable at distances greater than 10 meters in low-SNR environments. The framework reduces false alarms caused by boat noise and generalizes to external datasets, validating cross-domain transferability. This work bridges AI-based speech processing and passive sonar signal processing, offering a resource-efficient solution for real-time underwater surveillance.
Keywords:
passive diver detection, keyword spotting, transfer learning, real-time detection, passive sonar, hydrophoneReferences
- Alvaro A., Schwock F., Ragland J., Abadi S. (2021), Ship detection from passive underwater acoustic recordings using machine learning, The Journal of the Acoustical Society of America, 150(4 Supplement): A124, https://doi.org/10.1121/10.0007848
- Chen X., Wang R., Tureli U. (2006), Passive acoustic detection of divers under strong interference, [in:] OCEANS 2006, https://doi.org/10.1109/OCEANS.2006.306869
- Chen Y., Shang J. (2019), Underwater target recognition method based on convolution autoencoder, [in:] 2019 IEEE International Conference on Signal, Information and Data Processing (ICSIDP), https://doi.org/10.1109/ICSIDP47821.2019.9173362
- Chung K.W., Li H., Sutin A. (2007), A frequency-domain multi-band matched-filter approach to passive diver detection, [in:] 2007 Conference Record of the Forty-First Asilomar Conference on Signals, Systems and Computers, https://doi.org/10.1109/ACSSC.2007.4487426
- Cole A.M. (2019), Automated open circuit scuba diver detection with low cost passive sonar and machine learning, MA Thesis, Massachusetts Institute of Technology, https://hdl.handle.net/1721.1/122269
- Davi-sh (2022), Underwater scuba diver, Pixabay, https://pixabay.com/sound-effects/underwater-scuba-diver-28467/ (access: 21.06.2025).
- Deeb O., Jafar A., Al Dakkak O. (2025), A deep learning framework for Arabic continuous speech keyword spotting in low-resource settings using isolated-word keyword spotting and posterior probability functions, Advances in Artificial Intelligence and Machine Learning, 5(3): 4074–4093, https://doi.org/10.54364/AAIML.2025.53229
- Domingos L.C.F., Santos P.E., Skelton P.S.M., Brinkworth R.S.A., Sammut K. (2022), A survey of underwater acoustic data classification methods using deep learning for shoreline surveillance, Sensors, 22(6): 2181, https://doi.org/10.3390/s22062181
- Dong Y., Shen X., Wang H. (2022), Bidirectional denoising autoencoders-based robust representation learning for underwater acoustic target signal denoising, IEEE Transactions on Instrumentation and Measurement, 71: 1–8, https://doi.org/10.1109/TIM.2022.3210979
- Donskoy D.M., Sedunov N.A., Sedunov A.N., Tsionskiy M.A. (2008), Variability of SCUBA diver’s acoustic emission, [in:] Proceedings of SPIE – The International Society for Optical Engineering, 6945, https://doi.org/10.1117/12.783500
- Feng S., Ma S., Zhu X., Yan M. (2024), Artificial intelligence-based underwater acoustic target recognition: a survey, Remote Sensing, 16(17): 3333, https://doi.org/10.3390/rs16173333
- Fuchs L.R., Larsson C., Gallstrom A. (2019), Deep learning based technique for enhanced sonar imaging, [in:] Underwater Acoustic Conference and Exhibition Series.
- Gong Y., Lai C.-I., Chung Y.-A., Glass J. (2022), SSAST: self-supervised audio spectrogram transformer, [in:] Proceedings of the AAAI Conference on Artificial Intelligence, 36(10): 10699–10709, https://doi.org/10.1609/AAAI.V36I10.21315
- Gorovoy S. et al. (2014), A possibility to use respiratory noises for diver detection and monitoring physiologic status, The Journal of the Acoustical Society of America, 135 (4 Supplement): 2303, https://doi.org/10.1121/1.4877584
- Gorovoy S. et al. (2015), Detecting respiratory noises of diver equipped with rebreather in water, [in:] Proceedings of Meetings on Acoustics, 24(1): 070020, https://doi.org/10.1121/2.0000171
- Hari V.N., Chitre M., Too Y.M., Pallayil V. (2015), Robust passive diver detection in shallow ocean, [in:] OCEANS 2015 – Genova, https://doi.org/10.1109/OCEANS-Genova.2015.7271656
- Hewamalage H., Bergmeir C., Bandara K. (2021), Recurrent neural networks for time series forecasting: current status and future directions, International Journal of Forecasting, 37(1): 388–427, https://doi.org/10.1016/J.IJFORECAST.2020.06.008
- Hsu W.-N., Bolte B., Tsai Y.-H.H., Lakhotia K., Salakhutdinov R., Mohamed A. (2021), HuBERT: selfsupervised speech representation learning by masked prediction of hidden units, IEEE/ACM Transactions on Audio Speech and Language Processing, 29: 3451–3460, https://doi.org/10.1109/TASLP.2021.3122291
- Huang F., Zhang J., Zhou C., Wang Y., Huang J., Zhu L. (2020), A deep learning algorithm using a fully connected sparse autoencoder neural network for landslide susceptibility prediction, Landslides, 17(1): 217–229, https://doi.org/10.1007/s10346-019-01274-9
- Jin B., Xu G. (2020), A passive detection method of divers based on deep learning, [in:] 2020 IEEE 3rd International Conference on Electronics Technology (ICET), https://doi.org/10.1109/ICET49382.2020.9119556
- Jin G., Liu F., Wu H., Song Q. (2020), Deep learning-based framework for expansion, recognition and classification of underwater acoustic signal, Journal of Experimental and Theoretical Artificial Intelligence, 32(2): 205–218, https://doi.org/10.1080/0952813X.2019.1647560
- Johansson A.T., Lennartsson R.K., Nolander E., Petrović S. (2010), Improved passive acoustic detection of divers in harbor environments using pre-whitening, [in:] OCEANS 2010 MTS/IEEE SEATTLE, https://doi.org/10.1109/OCEANS.2010.5664549
- Karjalainen A.I., Mitchell R., Vazquez J. (2019), Training and validation of automatic target recognition systems using generative adversarial networks, [in:] 2019 Sensor Signal Processing for Defence Conference (SSPD 2019), https://doi.org/10.1109/SSPD.2019.8751666
- Korenbaum V., Kostiv A., Gorovoy S., Dorozhko V., Shiryaev A. (2020), Underwater noises of open-circuit scuba diver, Archives of Acoustics, 45(2): 349–357, https://doi.org/10.24425/aoa.2020.133155
- Lennartsson R.K., Dalberg E., Persson L., Petrović S. (2009), Passive acoustic detection and classification of divers in Harbor environments, [in:] OCEANS 2009, https://doi.org/10.23919/OCEANS.2009.5422407
- Li P., Wu J.,Wang Y., Lan Q., Xiao W. (2022), STM: spectrogram transformer model for underwater acoustic target recognition, Journal of Marine Science and Engineering, 10(10): 1428, https://doi.org/10.3390/JMSE10101428
- Lin J., Kilgour K., Roblek D., Sharifi M. (2020), Training keyword spotters with limited and synthesized speech data, [in:] ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing – Proceedings, https://doi.org/10.1109/ICASSP40776.2020.9053193
- Mahmoud S., Saleh L., Chouaib I. (2025), Experimental results of diver detection in Harbor environments using single acoustic vector sensor, Archives of Acoustics, 50(2): 173–185, https://doi.org/10.24425/aoa.2025.153663
- Phung S.L. et al. (2019), Mine-like object sensing in sonar imagery with a compact deep learning architecture for scarce data, [in:] 2019 Digital Image Computing: Techniques and Applications (DICTA 2019), https://doi.org/10.1109/DICTA47822.2019.8945982
- Radford C.A., Jeffs A.G., Tindle C.T., Cole R.G., Montgomery J.C., (2005), Bubbled waters: the noise generated by underwater breathing apparatus, Marine and Freshwater Behaviour and Physiology, 38(4): 259–267, https://doi.org/10.1080/10236240500333908
- Seo D., Oh H.-S., Jung Y. (2021), Wav2KWS: transfer learning from speech representations for keyword spotting, IEEE Access, 9: 80682–80691, https://doi.org/10.1109/ACCESS.2021.3078715
- Stolkin R. et al. (2006), Feature based passive acoustic detection of underwater threats, [in:] Proceedings SPIE 6204, Photonics for Port and Harbor Security II, https://doi.org/10.1117/12.663651
- Sun Y. et al. (2022), Multi-resonance flextensional hydrophone for open-circuit scuba diver detection, AIP Advances, 12(11): 115310, https://doi.org/10.1063/5.0101999
- Sun Y. et al. (2024), Feature extraction methods for underwater acoustic target recognition of divers, Sensors, 24(13): 4412, https://doi.org/10.3390/s24134412
- Sutin A., Salloum H., DeLorme M., Sedunov N., Sedunov A., Tsionskiy M. (2013), Stevens passive acoustic system for surface and underwater threat detection, [in:] 2013 IEEE International Conference on Technologies for Homeland Security (HST), https://doi.org/10.1109/THS.2013.6698999
- Tu Q., Yuan F., Yang W., Cheng E. (2020), An approach for diver passive detection based on the established model of breathing sound emission, Journal of Marine Science and Engineering, 8(1): 44, https://doi.org/10.3390/JMSE8010044
- Vincent P., Larochelle H., Lajoie I., Bengio Y., Manzagol P.-A. (2010), Stacked denoising autoencoders: learning useful representations in a deep network with a local denoising criterion, Journal of Machine Learning Research, 11: 3371–3408.
- VoiceBosch (2024), Sonar, Pixabay, https://pixabay.com/sound-effects/sonar-183436/ (access: 21.05.2025).
- Wang W., Huang Y., Wang Y., Wang L. (2014), Generalized autoencoder: a neural network framework for dimensionality reduction, [in:] IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, https://doi.org/10.1109/CVPRW.2014.79
- Wang X., Meng J., Liu Y., Zhan G., Tian Z. (2022), Self-supervised acoustic representation learning via acoustic-embedding memory unit modified space autoencoder for underwater target recognition, The Journal of the Acoustical Society of America, 152(5): 2905–2915, https://doi.org/10.1121/10.0015138
- Wang Y. et al. (2021), Underwater communication signal recognition using sequence convolutional network, IEEE Access, 9: 46886–46899, https://doi.org/10.1109/ACCESS.2021.3067070
- Zhang X., Lu Z., Kang C. (2003), Underwater acoustic targets classification using support vector machine, [in:] Proceedings of 2003 International Conference on Neural Networks and Signal Processing, https://doi.org/10.1109/ICNNSP.2003.1280753
- Zhao J. et al. (2023), Underwater target perception algorithm based on pressure sequence generative adversarial network, Ocean Engineering, 286(Part 1): 115547, https://doi.org/10.1016/J.OCEANENG.2023.115547
- Zhao W. et al. (2016), Passive acoustic detection of diver based on SVM, [in:] 2016 IEEE International Conference on Mechatronics and Automation, https://doi.org/10.1109/ICMA.2016.7558635

