Implementation of a COVAREP-Integrated Application for Voice Analysis Using LLMs

Downloads

Authors

  • Hubert Daniszewski Polish-Japanese Academy of Information Technology, Poland
  • Krzysztof Szklanny Polish-Japanese Academy of Information Technology, Poland ORCID ID 0000-0001-6540-1671
  • Piotr Wrzeciono Warsaw University of Life Sciences, Poland ORCID ID 0000-0003-3433-3906

Abstract

This paper presents VocalCheck, a mobile application designed to monitor voice conditions using recordings made with a smartphone. The initial phase involved a literature review and an evaluation of existing voice processing applications, which informed the project’s design requirements. Based on this analysis, an application was developed to record and preliminarily analyze voice samples using parameters provided by the Collaborative Voice Analysis Repository (COVAREP) for the Speech Technologies Toolkit. Additionally, an experimental study was conducted to assess the effectiveness of large language models (LLMs), specifically Microsoft Copilot and ChatGPT, in generating MATLAB scripts for computing voice quality parameters. Although these tools proved helpful for general programming tasks, their performance in speech signal processing was unsatisfactory.

Keywords:

voice quality, voice diagnostics, mobile app, speech processing, LLM

References


  1. Alhussein M., Muhammad G. (2019), Automatic voice pathology monitoring using parallel deep models for smart healthcare, IEEE Access, 7: 46474–46479, https://doi.org/10.1109/ACCESS.2019.2905597

  2. Alku P., Strik H., Vilkman E. (1997), Parabolic spectral parameter – a new method for quantification of the glottal flow, Speech Communication, 22(1): 67–79, https://doi.org/10.1016/s0167-6393(97)00020-4

  3. Alku P., Backstrom T., Vilkman E. (2002), Normalized amplitude quotient for parametrization of the glottal flow, The Journal of the Acoustical Society of America, 112(2): 701–710, https://doi.org/10.1121/1.1490365

  4. Airas M., Alku P. (2007), Comparison of multiple voice source parameters in different phonation types, [in:] Proceedings Interspeech 2007, https://doi.org/10.21437/interspeech.2007-28

  5. Angadi V., Chih M.-Y., Stemple J. (2023), Developing and testing a smartphone application to enhance adherence to voice therapy: a pilot study, International Journal of Environmental Research and Public Health, 20(3): 2436, https://doi.org/10.3390/ijerph20032436

  6. Askenfelt A.G., Hammarberg B. (1986), Speech waveform perturbation analysis: a perceptual-acoustical comparison of seven measures, Journal of Speech, Language, and Hearing Research, 29(1): 50–64, https://doi.org/10.1044/jshr.2901.50

  7. Asci F. et al. (2020), Machine-learning analysis of voice samples recorded through smartphones: the combined effect of ageing and gender, Sensors, 20(18): 5022, https://doi.org/10.3390/s20185022

  8. Bjorkner E., Sundberg J., Alku P. (2005), Subglottal pressure and NAQ variation in voice production of classically trained baritone singers, [in:] Proceedings Interspeech 2005, pp. 1057–1060, https://doi.org/10.21437/interspeech.2005-424

  9. Boersma P. (2001), Praat, a system for doing phonetics by computer, Glot International, 5(9/10): 341–345.

  10. Childers D.G., Lee C.K. (1991), Vocal quality factors: analysis, synthesis, and perception, The Journal of the Acoustical Society of America, 90(5): 2394–2410, https://doi.org/10.1121/1.402044

  11. Compton E.C. et al. (2022), Developing an Artificial Intelligence tool to predict vocal cord pathology in primary care settings, The Laryngoscope, 133(8): 1952–1960, https://doi.org/10.1002/lary.30432

  12. Cooper W.E., Sorensen J.M. (1981), Fundamental Frequency in Sentence Production, Springer Science & Business Media.

  13. Crowson M.G. et al. (2020), A contemporary review of machine learning in otolaryngology–head and neck surgery, The Laryngoscope, 130(1): 45–51, https://doi.org/10.1002/lary.27850

  14. Degottex G., Kane J., Drugman T., Raitio T., Scherer S. (2014), COVAREP – a collaborative voice analysis repository for speech technologies, [in:] 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 960–964, https://doi.org/10.1109/icassp.2014.6853739

  15. Hacki T. (1989), Classification of glottal dysfunctions on the basis of electroglottography [in German: Klassifizierung von glottiscysfunktionen mit hilfe der elektroglottographie], Folia Phoniatrica, 41(1): 43–48, https://doi.org/10.1159/000265931

  16. Hanson H.M. (1997), Glottal characteristics of female speakers: acoustic correlates, The Journal of the Acoustical Society of America, 101(1): 466–481, https://doi.org/10.1121/1.417991

  17. Hendriksz C.J. et al. (2014), Burden of disease in patients with Morquio A syndrome: results from an international patient-reported outcomes survey, Orphanet Journal of Rare Diseases, 9(1): 32, https://doi.org/10.1186/1750-1172-9-32

  18. Hillenbrand J., Houde R.A. (1996), Acoustic correlates of breathy vocal quality: dysphonic voices and continuous speech, Journal of Speech, Language, and Hearing Research, 39(2): 311–321, https://doi.org/10.1044/jshr.3902.311

  19. Ho T.K. (1995), Random decision forests, [in:] Proceedings of 3rd International Conference on Document Analysis and Recognition, 1: 278–282, https://doi.org/10.1109/ICDAR.1995.598994

  20. Holland J.H. (1975), Adaptation in Natural and Artificial Systems. An Introductory Analysis with Applications to Biology, Control, and Artificial Intelligence, The MIT Press.

  21. Howard D.M. (1995), Variation of electrolaryngographically derived closed quotient for trained and untrained adult female singers, Journal of Voice, 9(2): 163–172, https://doi.org/10.1016/s0892-1997(05)80250-4

  22. Kane J., Gobl C. (2011), Identifying regions of nonmodal phonation using features of the wavelet transform, [in:] Proceedings Interspeech 2011, pp. 177–180, https://doi.org/10.21437/interspeech.2011-76

  23. Kane J., Gobl C. (2013),Wavelet maxima dispersion for breathy to tense voice discrimination, [in:] IEEE Transactions on Audio, Speech, and Language Processing, 21(6): 1170–1179, https://doi.org/10.1109/tasl.2013.2245653

  24. Kasuya H., Shigeki O., Mashima K., Ebihara S. (1986) Normalized noise energy as an acoustic measure to evaluate pathologic voice, The Journal of the Acoustical Society of America, 80(5): 1329–1334, https://doi.org/10.1121/1.394384

  25. Kaur P., Stoltzfus, J.C. (2017), Bland–Altman plot: a brief overview, International Journal of Academic Medicine, 3(1): 110–111, https://doi.org/10.4103/IJAM.IJAM 54 17.

  26. Klingholz R. (1987), The measurement of the signal-to-noise ratio (SNR) in continuous speech, Speech Communication, 6(1): 15–26, https://doi.org/10.1016/0167-6393(87)90066-5

  27. Kosztyła-Hojna B., Moskal D., Kuryliszyn-Moskal A., Rutkowski R. (2014), Visual assessment of voice disorders in patients with occupational dysphonia, Annals of Agricultural and Environmental Medicine, 21(4): 898–902, https://doi.org/10.5604/12321966.1129955

  28. Lebacq J., Schoentgen J., Cantarella G., Bruss F.T., Manfredi C., DeJonckere P. (2017), Maximal ambient noise levels and type of voice material required for valid use of smartphones in clinical voice research, Journal of Voice, 31(5): 550–556, https://doi.org/10.1016/j.jvoice.2017.02.017

  29. Manfredi C. et al. (2017), Smartphones offer new opportunities in clinical voice research, Journal of Voice, 31(1): 111, https://doi.org/10.1016/j.jvoice.2015.12.020

  30. Nawka T., Anders L., Wendler J. (1994), The auditory assessment of hoarse voices according to the RBH system [in German], Sprache Stimme, Gehor, 18: 130–133.

  31. Parsa V., Jamieson D.G. (2001), Acoustic discrimination of pathological voice: sustained vowels versus continuous speech, Journal of Speech, Language, and Hearing Research, 44(2): 327–339, https://doi.org/10.1044/1092-4388(2001/027)

  32. Patel R.R. et al. (2018), Recommended protocols for instrumental assessment of voice: American Speech Language-Hearing Association expert panel to develop a protocol for instrumental assessment of vocal function, American Journal of Speech-Language Pathology, 27(3): 887–905, https://doi.org/10.1044/2018 ajslp-17-0009.

  33. Suvvari T.K. (2023), The role of Artificial Intelligence in diagnosis and management of laryngeal disorders, Ear, Nose & Throat Journal, 104(11), https://doi.org/10.1177/01455613231175053

  34. Szklanny K. (2019), Acoustic parameters in the evaluation of voice quality of choral singers. Prototype of mobile application for voice quality evaluation, Archives of Acoustics, 44(3): 439–446, https://doi.org/10.24425/aoa.2019.129257

  35. Szklanny K., Gubrynowicz R., Iwanicka-Pronicka K., Tylki-Szymańska A. (2016), Analysis of voice quality in patients with late-onset Pompe disease, Orphanet Journal of Rare Diseases, 11(1): 99, https://doi.org/10.1186/s13023-016-0480-5

  36. Szklanny K., Gubrynowicz R., Ratyńska J., Chojnacka-Wądołowska D. (2019), Electroglottographic and acoustic analysis of voice in children with vocal nodules, International Journal of Pediatric Otorhinolaryngology, 122: 82–88, https://doi.org/10.1016/j.ijporl.2019.03.030

  37. Szklanny K., Gubrynowicz R., Tylki-Szymańska A. (2018), Voice alterations in patients with Morquio A syndrome, Journal of Applied Genetics, 59(1): 73–80, https://doi.org/10.1007/s13353-017-0421-6

  38. Szklanny K., Tylki-Szymańska A. (2018), Follow-up analysis of voice quality in patients with late-onset Pompe disease, Orphanet Journal of Rare Diseases, 13(1): 189, https://doi.org/10.1186/s13023-018-0932-1

  39. Szklanny K., Wrzeciono P. (2019a), Relation of RBH auditory-perceptual scale to acoustic and electroglottographic voice analysis in children with vocal nodules, IEEE Access, 7: 41647–41658, https://doi.org/10.1109/ACCESS.2019.2907397

  40. Szklanny K., Wrzeciono P. (2019b), The application of a genetic algorithm in the noninvasive assessment of vocal nodules in children, IEEE Access, 7: 44966–44976, https://doi.org/10.1109/ACCESS.2019.2908313

  41. Tirronen S., Javanmardi F., Kodali M., Reddy Kadiri S., Alku P. (2023), Utilizing Wav2Vec in database-independent voice disorder detection, [in:] ICASSP 2023 – 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1–5, https://doi.org/10.1109/ICASSP49357.2023.10094798

  42. Uloza V. et al. (2015), Exploring the feasibility of smart phone microphone for measurement of acoustic voice parameters and voice pathology screening, European Archives of Oto-rhino-laryngology, 272(11): 3391–3399, https://doi.org/10.1007/s00405-015-3708-4

  43. Uloza V. et al. (2023), Accuracy of acoustic voice quality index captured with a smartphone – measurements with added ambient noise, Journal of Voice, 37(3): 465, https://doi.org/10.1016/j.jvoice.2021.01.025

  44. Verde L., De Petro G., Alrashoud M., Ghoneim A., Al-Mutib K.N., Sannino G. (2019), Leveraging Artificial Intelligence to improve voice disorder identification through the use of a reliable mobile app, IEEE Access, 7: 124048–124054, https://doi.org/10.1109/ACCESS.2019.2938265

  45. Verikas A., Gelzinis A., Bacauskiene M., Uloza V. (2006), Towards a computer-aided diagnosis system for vocal cord diseases, Artificial Intelligence in Medicine, 36(1): 71–84, https://doi.org/10.1016/j.artmed.2004.11.001

  46. Yumoto E., Gould W.J., Baer T. (1982), Harmonics-to-noise ratio as an index of the degree of hoarseness, The Journal of the Acoustical Society of America, 71(6): 1544, https://doi.org/10.1121/1.387808

  47. Zheng F., Zhang G., Song Z. (2001), Comparison of different implementations of MFCC, Journal of Computer science and Technology, 16(6): 582–589, https://doi.org/10.1007/BF02943243