You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Community-controlled voice data collection for language preservation and AI development. Companion to 'AI Techniques for Indigenous Cultural Expression' in Envisioning Indigenous Methods in Digital Media and Ecologies.
Zolai AI Second Brain — NLP toolkit to digitize, standardize & preserve the Tedim Zolai language. Includes data pipelines, 152k-entry dictionary, Bible corpus, LLM fine-tuning (Qwen2.5-3B LoRA), and a Next.js learning app.
This project is rooted in the fascinating history of the Sakha language's written form. In the early 20th century, the language was transcribed using a Latin script developed by the eminent linguist Semyon Andreevich Novgorodov. This historical period holds significant cultural resonance, especially since the language transitioned to the Cyrillic.
A volunteer project to create the first modern, freely available Bible translation in Northern Hindko — now recruiting native Hindko speakers to translate, review, and proofread.
The first word processor built for English and Tibetan. Free, open source, 100% offline. Wylie and TCRC-Bodyig input, Tibetan-English-Tibetan dictionary,
Angika-LowResource-Translator is a neural machine translation model designed to translate between Angika and English. Built for low-resource language support, it leverages fine-tuned transformers to preserve cultural context and improve accessibility for underrepresented communities.
Open, reproducible toolkit for learning and preserving Abkhaz (аҧсуа бызшәа), a Northwest Caucasian language UNESCO lists as vulnerable — corpus frequency data (CC0), an A0–A2 course, and a bilingual site.