Back to datasets
Speech infrastructure

AxomiyaVoice

An open Assamese speech resource for preserving authentic voices, building verified transcriptions, and creating the foundation for Assamese speech technology.

Assamese speech dataset
Assamese speech and voice infrastructure represented through AxomiyaVoice
Open speech infrastructure
Why AxomiyaVoice

A language is more than written text. It lives in the voices of its people.

Assamese has a rich spoken tradition shaped by region, community, generation and everyday life. AxomiyaVoice aims to help preserve that living linguistic diversity as structured, documented and reusable speech data.

Authentic voices

Collect real Assamese speech from native speakers so future language technology can learn from natural pronunciation, rhythm and expression.

Verified transcriptions

Pair recordings with carefully reviewed Assamese transcriptions to create reliable speech-text resources for research and development.

Open & documented

Build speech resources transparently with clear documentation, contribution practices and responsible data stewardship.

Preserving Assamese voices for the AI era.
Language preservation

Preserving Assamese voices for the AI era.

Speech carries things that text alone cannot fully capture — pronunciation, rhythm, emphasis, accent and the natural way people communicate. Open speech resources can help ensure these characteristics remain represented as Assamese language technology evolves.

Language preservation & technology
Voice is part of cultural memory.
Assamese voice & culture

Voice is part of cultural memory.

Assamese culture has always carried stories, emotions and identity through music, conversation and oral tradition. Cultural voices can inspire how we think about preserving Assamese speech — not as a collection of isolated recordings, but as part of a living language.

Preserving Assamese voices is not only about creating data for AI — it is also about documenting the spoken heritage of the language for the future.

What it can enable

Building blocks for Assamese speech technology.

Speech recognition

Support future systems that can understand spoken Assamese more naturally.

Speech synthesis

Provide language resources that can contribute to more natural Assamese voice systems.

Voice applications

Create opportunities for accessibility, education, assistants and other voice-based applications.

Built with speakers

Every Assamese voice can help strengthen the dataset.

Native speakers, students, educators, researchers and developers can contribute recordings, transcriptions, reviews and documentation.

See how to contribute