Speech resources of adequate quality are crucial for the development of speech technologies and spoken language research, but they are still scarce in the field of spontaneously produced spoken Slovenian due to the complexity of collection. The production of speech corpora and databases is costly, which is why researchers are increasingly recognizing the potential of citizen science. By leveraging crowdsourcing and other methods, it enables the efficient remote collection of large amounts of speech data. In this paper, we discuss the key factors - technical, financial, legal, ethical and motivational - that need to be considered when designing a sustainable and scalable system for speech acquisition. Based on a literature review, analysis of existing methods and global initiatives for collecting speech resources, we provide recommendations suitable for implementation in the Slovenian context.
|