Projects

Two programmes, one method

NepSwor works across the regional languages of Nepal's plains and hills. YetiVoices works with Himalayan and Buddhist heritage languages. Both follow the same community-first pipeline, from consent to a working model.

Programmes

Where each one stands

NepSwor

Regional languages of Nepal · 9 languages

Parallel sentence corpora, read and spontaneous speech, and shared benchmarks for speech recognition and translation across the major languages of the Terai and the mid-hills.

Languages in scope
9
Sentence target
4 lakh
Validated audio
1000+ hrs
Validation passes
2 per clip
NepaliMaithiliBhojpuriTharuTamangBajjikaAwadhiNepal Bhasha (Newari)Magar Dhut
Full methodology and progress

YetiVoices

Himalayan heritage languages · 7 languages

Documentation of high-altitude and Buddhist heritage languages: structured corpora alongside ritual recitation, song and oral history, archived with the monasteries and communities that hold them.

Languages in scope
7
Endangered focus
High
Cultural recordings
Song, ritual, story
Archival standard
FAIR
SherpaLimbuGurungTibetanTamangThakaliHumli
Full methodology and progress

Methodology

From a community conversation to a working voice model

  1. 01

    Community consent and design

    We begin with the speech community: mother-tongue organisations, elders and local schools agree what may be recorded, who owns it and how it may be used. Nothing enters a dataset without a signed community licence.

  2. 02

    Readiness audit

    Each language is scored on written resources, institutional support, digital presence and existing corpora, producing a preparedness index that tells us whether to start with an alphabet, a wordlist or a full speech corpus.

  3. 03

    Prompt and sentence design

    Native writers build balanced prompt sets covering everyday speech, numerals, place names, agriculture, ritual vocabulary and code-switching, so the corpus is phonetically and topically representative.

  4. 04

    Field recording

    Speakers of different ages, genders and dialect areas record in their own villages on calibrated equipment, with a supervisor checking noise floor, clipping and speaker metadata on the spot.

  5. 05

    Transcription and double validation

    Every clip is transcribed in the community's chosen script, then reviewed by a second native validator. Disagreements go to a dialect panel rather than being silently overwritten.

  6. 06

    Model training and benchmarks

    Validated audio and parallel text feed speech recognition, text-to-speech and translation baselines, published with held-out test splits so improvements are measurable.

  7. 07

    Return to the community

    Learning material, dictionaries, word cards and working voice tools go back to the speakers, alongside archival copies held to FAIR standards.

Bring your language into a programme.

If your community keeps a language that is not yet in NepSwor or YetiVoices, we can start with a readiness audit and a small pilot recording.

Start a conversation