What It Solves
AI models struggle with Assamese because of a severe lack of structured, open-license regional datasets. Open Assam Data Hub aggregates, cleans, and publishes open datasets for Assamese speech, text, dictionary terms, and geographical statistics.
Key Features
- Assamese Text Corpus: Cleaned, open-license text datasets for fine-tuning LLMs.
- Audio & Speech Dataset: Crowdsourced voice recordings for Assamese Speech-to-Text models.
- Automated Data Cleaning: Dataform & Python pipelines for deduplication and formatting.
Status & Roadmap
🚧 Under Construction — Coming Soon! Active dataset collection and cleaning pipelines are under active development. Full documentation, download links, and Hugging Face dataset releases will be updated once ready and available.