Vernacular Pedagogy: Hindi-to-Santhali Edge AI Models

This repository hosts production-ready, lightweight offline models for real-time Hindi $\rightarrow$ Santhali translation and natural speech synthesis, designed for Foundational Literacy & Numeracy (FLN) Grade 1โ€“3 primary education.

Model Contents

  1. sat_piper_model.onnx (~60.6 MB): End-to-end VITS neural TTS model trained on authentic multi-speaker Santhali speech (AI4Bharat IndicVoices-R & XKaab).
  2. sat_piper_model.onnx.json: Model configuration, audio sample rate (16,000 Hz), and native Ol Chiki phoneme symbol mapping (U+1C50โ€“U+1C7F).
  3. indictrans2_sat_int8_ct2.tar.gz (~286.7 MB): CTranslate2 INT8 quantized neural machine translation model (hin_Deva $\rightarrow$ sat_Olck), optimized for ultra-fast mobile CPU inference (<120 ms).
  4. fln_lexicon.sqlite: Pre-indexed B-Tree SQLite cache of 368 verified classroom interactions (<0.1 ms retrieval).

Benchmark Performance (Standard CPU)

  • Speech Synthesis (Piper ONNX): Real-Time Factor (RTF) = 0.051 โ€“ 0.081 (>4x faster than real-time)
  • Synthesis Latency: 41 ms โ€“ 87 ms per classroom command
  • Translation Latency: <120 ms per sentence

License & Attribution

Trained and packaged as part of the Vernacular Pedagogy initiative.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support