What It Solves

AI models struggle with Assamese because of a severe lack of structured, open-license regional datasets. Open Assam Data Hub aggregates, cleans, and publishes open datasets for Assamese speech, text, dictionary terms, and geographical statistics.

Key Features

  • Assamese Text Corpus: Cleaned, open-license text datasets for fine-tuning LLMs.
  • Audio & Speech Dataset: Crowdsourced voice recordings for Assamese Speech-to-Text models.
  • Automated Data Cleaning: Dataform & Python pipelines for deduplication and formatting.

Status & Roadmap

🚧 Under Construction — Coming Soon! Active dataset collection and cleaning pipelines are under active development. Full documentation, download links, and Hugging Face dataset releases will be updated once ready and available.

Coming Soon Coming Soon!

This project is currently in concept/development phase. Live links and repositories will be available upon launch.

#Open-Data #Datasets #Assamese-NLP

Explore Other Projects

Check out our live tools and upcoming community proposals.