Draft:Low-resource languages
This article may incorporate text from a large language model, which is prohibited in Wikipedia articles. (May 2026) |
A low-resource language is a language for which limited digital or computational resources are available for use in natural language processing, computational linguistics, or language technology. The term is commonly used to refer to the availability of machine-readable text, annotated corpora, parallel corpora, speech recordings, linguistic databases, or software tools.[1] It does not refer to the linguistic complexity, cultural value, or number of speakers of a language.[2]
In research contexts, low-resource status is generally treated as relative and task-dependent. A language may have sufficient resources for one application, such as written text processing, while having few resources for another, such as automatic speech recognition or machine translation.[1][2]
Definition and scope
In natural language processing, a language is often described as low-resource when the amount or type of available data is insufficient for developing, training, adapting, or evaluating computational systems for a specific task.[1] Relevant resources may include monolingual corpora, annotated datasets, bilingual or parallel corpora, speech recordings with transcriptions, treebanks, lexicons, pronunciation dictionaries, or benchmark datasets.
The term does not have a single universal threshold. Natural language processing researchers found limited consensus on what qualifies as a "low-resource language" and argued that low-resourcedness depends on several factors, including the amount and type of available data, the task, the domain, and the language technology setting.[2]
Low-resource status is not equivalent to being an endangered language, minority language, Indigenous language, or demographically small language. Some languages with large speaker populations have limited computational resources, while some endangered languages may have documentation resources for specific scholarly purposes.[2]
Types of resources
Resources relevant to language technology include several types of data, tools, and evaluation materials.
- Monolingual corpora: collections of written or transcribed text in one language.
- Annotated corpora: text or speech labelled for linguistic or computational tasks, such as part-of-speech tagging, named-entity recognition, syntactic parsing, or sentiment analysis.
- Parallel corpora: aligned texts in two or more languages, commonly used in machine translation.[3]
- Speech corpora: audio recordings, transcriptions, speaker metadata, pronunciation data, or related speech resources.
- Lexical and grammatical resources: dictionaries, morphological analysers, treebanks, wordnets, terminology databases, or pronunciation dictionaries.
- Evaluation datasets: datasets used to compare the performance of language technology systems.
The availability of these resources can differ by language, region, domain, and application.[1][3]
Causes of resource scarcity
The scarcity of computational resources for a language may arise from several factors, including limited access to digital infrastructure, limited digitisation of written material, dispersed speaker communities, limited institutional support, limited commercial incentives, or the use of a language mainly in oral domains.[2]
Resource scarcity is generally not treated as an inherent property of a language. Rather, it describes the current availability of data, tools, research infrastructure, and technology support for particular computational purposes.[2]
Impact on language technology
Low-resource languages may be less supported by language technologies such as machine translation, automatic speech recognition, text-to-speech systems, spell-checkers, search tools, conversational systems, and information retrieval systems. Studies of linguistic diversity in natural language processing have found substantial disparities in the representation of the world's languages in language technology research and applications.[4]
The effects vary by language and by task. A language may be supported for one technology or language pair but not for another. For this reason, researchers often analyse low-resource conditions at the level of specific tasks, datasets, domains, or language pairs.[1][3]
In European policy and research, the related concept of digital language equality refers to the goal of ensuring adequate language-technology support across languages. The European Language Equality project examined language technology support for European languages and produced a strategic agenda and roadmap for digital language equality in Europe.[5]
Computational approaches
Several methods have been used in natural language processing to address limited data availability. These approaches are not specific to any single language and may be combined depending on the task and available resources.[1]
- Transfer learning: adapting models trained on one language, domain, or task to another.
- Cross-lingual learning: using data or representations from languages with more available resources to support languages with fewer resources.
- Multilingual modelling: training a single model on data from multiple languages.
- Data augmentation: generating or transforming training examples to increase the amount or variety of data.
- Weak, distant, or semi-supervised supervision: supplementing limited labelled data with automatically generated or indirectly obtained labels.
- Community-based data collection: involving speakers in the contribution, validation, or review of language data.
In machine translation, low-resource research often focuses on language pairs for which little translated training data is available.[3] For speech technology, crowdsourced and community-based data collection has been used to create multilingual speech corpora.[6]
Relation to other terms
The term low-resource language overlaps with several related terms that are not interchangeable.[2]
| Term | Main meaning |
|---|---|
| Low-resource language | A language with limited data, tools, or computational resources for one or more language technology tasks. |
| Under-resourced language | A near-synonym often used to emphasize lack of technological, institutional, or research support. |
| Endangered language | A language at risk of falling out of use by its speaker community. |
| Minority language | A language spoken by a minority population within a particular state, region, or political context. |
| Indigenous language | A language associated with Indigenous peoples; it may or may not be low-resource or endangered. |
These categories may overlap. For example, a language can be low-resource without being endangered, or endangered while having some documentation resources.[2]
Initiatives
Several research communities, public initiatives, and data platforms address language-resource scarcity or the digital representation of under-supported languages.
- Masakhane is a grassroots research organization focused on natural language processing for African languages.[7]
- AmericasNLP is a workshop series focused on natural language processing for Indigenous languages of the Americas, with proceedings published through the ACL Anthology.[8]
- Language Technologies for All is an initiative associated with UNESCO, the European Language Resources Association, and SIGUL that addresses language technologies, linguistic diversity, multilingualism, and under-resourced languages.[9]
- European Language Equality was a European initiative that produced reports, metrics, and a strategic agenda on digital language equality and language technology support in Europe.[5]
- Mozilla Common Voice is a crowdsourced multilingual speech corpus used in speech technology research and development.[6]
- UNESCO World Atlas of Languages provides information on spoken and signed languages, including language status and domains of use.[10]
- The European Language Grid provides access to language technology tools, services, datasets, corpora, models, and information about language technology organisations in Europe.[11]
See also
- Computational linguistics
- Natural language processing
- Machine translation
- Language documentation
- Endangered language
- Minority language
- Corpus linguistics
- Digital divide
- Language revitalization
Further reading
- Hedderich, Michael A.; Lange, Lukas; Adel, Heike; Strötgen, Jannik; Klakow, Dietrich (2021). "A Survey on Recent Approaches for Natural Language Processing in Low-Resource Scenarios". Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics. pp. 2545–2568. doi:10.18653/v1/2021.naacl-main.201.
- Haddow, Barry; Bawden, Rachel; Miceli Barone, Antonio Valerio; Helcl, Jindřich; Birch, Alexandra (2022). "Survey of Low-Resource Machine Translation". Computational Linguistics. 48 (3). MIT Press: 673–732. doi:10.1162/coli_a_00446.
- Nigatu, Hellina Hailu; Tonja, Atnafu Lambebo; Rosman, Benjamin; Solorio, Thamar; Choudhury, Monojit (2024). "The Zeno's Paradox of 'Low-Resource' Languages". Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics. pp. 17753–17774. doi:10.18653/v1/2024.emnlp-main.983.
- Poria, Sampoorna; Huang, Xiaolei (2025). "Bhaasha, Bhāṣā, Zaban: A Survey for Low-Resourced Languages in South Asia – Current Stage and Challenges". Findings of the Association for Computational Linguistics: EMNLP 2025. Association for Computational Linguistics. doi:10.18653/v1/2025.findings-emnlp.73.
- Rehm, Georg; Way, Andy, eds. (2023). European Language Equality: A Strategic Agenda for Digital Language Equality. Springer. doi:10.1007/978-3-031-28819-7.
External links
- ACL Anthology
- UNESCO World Atlas of Languages
- UNESCO International Decade of Indigenous Languages
- European Language Equality
- European Language Grid
- Masakhane
- AmericasNLP
- Mozilla Common Voice
- European Language Resources Association
Category:Natural language processing
Category:Computational linguistics
Category:Corpus linguistics
Category:Language documentation
- ^ a b c d e f Hedderich, Michael A.; Lange, Lukas; Adel, Heike; Strötgen, Jannik; Klakow, Dietrich (2021). "A Survey on Recent Approaches for Natural Language Processing in Low-Resource Scenarios". Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics. pp. 2545–2568. doi:10.18653/v1/2021.naacl-main.201. Retrieved 14 May 2026.
- ^ a b c d e f g h Nigatu, Hellina Hailu; Tonja, Atnafu Lambebo; Rosman, Benjamin; Solorio, Thamar; Choudhury, Monojit (2024). "The Zeno's Paradox of 'Low-Resource' Languages". Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics. pp. 17753–17774. doi:10.18653/v1/2024.emnlp-main.983. Retrieved 14 May 2026.
- ^ a b c d Haddow, Barry; Bawden, Rachel; Miceli Barone, Antonio Valerio; Helcl, Jindřich; Birch, Alexandra (2022). "Survey of Low-Resource Machine Translation". Computational Linguistics. 48 (3). MIT Press: 673–732. doi:10.1162/coli_a_00446. Retrieved 14 May 2026.
- ^ Joshi, Pratik; Santy, Sebastin; Budhiraja, Amar; Bali, Kalika; Choudhury, Monojit (2020). "The State and Fate of Linguistic Diversity and Inclusion in the NLP World". Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics. pp. 6282–6293. doi:10.18653/v1/2020.acl-main.560. Retrieved 14 May 2026.
- ^ a b Rehm, Georg; Way, Andy, eds. (2023). European Language Equality: A Strategic Agenda for Digital Language Equality. Springer. doi:10.1007/978-3-031-28819-7. Retrieved 14 May 2026.
- ^ a b Ardila, Rosana; Branson, Megan; Davis, Kelly; Kohler, Michael; Meyer, Josh; Henretty, Michael; Morais, Reuben; Saunders, Lindsay; Tyers, Francis M.; Weber, Gregor (2020). "Common Voice: A Massively-Multilingual Speech Corpus". Proceedings of the Twelfth Language Resources and Evaluation Conference. European Language Resources Association. pp. 4218–4222. Retrieved 14 May 2026.
- ^ "Masakhane". Masakhane. Retrieved 14 May 2026.
- ^ "Workshop on Natural Language Processing for Indigenous Languages of the Americas". ACL Anthology. Retrieved 14 May 2026.
- ^ "Language Technologies for All – LT4All 2025". UNESCO. Retrieved 14 May 2026.
- ^ "UNESCO launches World Atlas of Languages to celebrate and protect linguistic diversity". UNESCO. 24 November 2021. Retrieved 14 May 2026.
- ^ "European Language Grid". European Language Grid. Retrieved 14 May 2026.
Content Disclaimer
Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.
- The information displayed on this website is sourced in part or in whole from Wikipedia and has been adapted for the purpose of restating it. We strive to provide accurate and relevant information, however:
- There is no guarantee of absolute accuracy. Wikipedia is an open, collaborative project that can be edited by anyone, so information is subject to change.
- It is not intended to constitute professional advice. The content displayed is for informational and educational purposes only. For important decisions (e.g., medical, legal, or financial), please consult a professional.
- Content copyright. Wikipedia is licensed under the Creative Commons Attribution-ShareAlike License (CC BY-SA). This means that content may be reused with appropriate attribution and shared under a similar license.
- Responsible use. Any risk arising from the use of information from this website is entirely the responsibility of the user.