Instructions to use FineEnvs/gemma-4-E4B-it-kannada-ocr-grpo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use FineEnvs/gemma-4-E4B-it-kannada-ocr-grpo with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-E4B-it") model = PeftModel.from_pretrained(base_model, "FineEnvs/gemma-4-E4B-it-kannada-ocr-grpo") - Notebooks
- Google Colab
- Kaggle
Gemma 4 E4B · Kannada document OCR (GRPO)
A LoRA adapter for google/gemma-4-E4B-it, trained
with GRPO against an OpenEnv document-OCR environment on
4,000 Kannada section crops from
NayanaOCR_Corpus_2025.
On the Kannada part of Sarvam Indic OCR Bench,
a benchmark it never trained on, character error falls 16% by Sarvam's own metric.
Part of the Multilingual Multimodal Envs collection. The environment is live at FineEnvs/nayana-ocr-env.
Results
Sarvam Indic OCR Bench, test split, all 300 Kannada crops, greedy decoding through vLLM. CER and WER
come from the benchmark's own metrics.py, vendored unmodified, so they are comparable with
Sarvam's published numbers. The change is paired crop by crop against the untuned base model.
| base | this adapter | change (95% CI) | |
|---|---|---|---|
| Sarvam CER | 0.4277 | 0.3601 | −0.068 (−0.083, −0.052) |
| Sarvam WER | 0.773 | 0.748 | −0.025 (−0.037, −0.014) |
| Sarvam word accuracy | 22.7% | 25.2% | |
| loop or catastrophic outputs | 16 | 10 | |
| environment reward | 0.455 | 0.513 | |
| crops containing another Indic script | 37 (12.3%) | 15 (5.0%) |
223 crops get better, 62 get worse and 15 are unchanged. No crop is read exactly, before or after: the remaining errors are a letter or two inside most words, which is why word error barely moves. Most of the gain arrives in the first 100 steps, and held-out CER is flat from step 300 to 500.
What it learned
It learned to stay in the script. One crop in eight from the untuned model has letters from another
Indic script in it; after training it is one in twenty. In this crop, at the 25th percentile of
improvement, the untuned model writes Devanagari (वायु) and Latin (nel) into a Kannada
passage, and the trained one does not.
| text | Sarvam CER | |
|---|---|---|
| reference | ಬಯ್ತ - ಅಡಗಿಸಿಟ್ಟ (ದಿವಿಜತತಿ ಬಯ್ತು ಕೈದುವಂ ಅವಸರದೊಳ್ / ಬೇಡೆ ಸುರಭಿಯ ನಕ್ಕಿಸೆ .. .. ಕಯ್ದುವಾದುವಾ / ಗೊರವನೆಲ್ವು: ಸಮಯಪ, ೧೦. ೧೪೨) | |
| base | ಬತ್ತು - ಅಡಗಿಸಿಟ್ಟಿ (ದವಿಚತೇ वायु ಕೃಡಮ ಅವಕಾಶದೊಳ / ಬೇದ ಸುರಭೆಯ ನಕ್ಷೆ ... ... ಕಮ್ಯುದದಾ / ಗೌರವnel: ಸಮುಮ, ೧೦.೧೭) | 0.455 |
| tuned | ಬತ್ತು - ಅಡಗಿಸಿಟಿ (ದಿವಸತೇ ಬತ್ತು ಕಡಮಂ ಅವಕಾಶದೊಳ / ಬೇಡ ಸುರಭಿಯ ನಕ್ಷೆ ... ... ಕಮ್ಯುವಾದುಹಾ / ಗಾರ್ದನಲ್ಯ: ಸಮಯವ, ೧೦.೧೭) | 0.347 |
Use it
With vLLM, serving the adapter over the base model:
vllm serve google/gemma-4-E4B-it --enable-lora --max-lora-rank 16 \
--limit-mm-per-prompt '{"image": 1}' \
--lora-modules kannada-ocr=FineEnvs/gemma-4-E4B-it-kannada-ocr-grpo
import base64, requests
crop = base64.b64encode(open("crop.png", "rb").read()).decode()
reply = requests.post("http://localhost:8000/v1/chat/completions", json={
"model": "kannada-ocr",
"temperature": 0,
"messages": [{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{crop}"}},
{"type": "text", "text": "Transcribe the text in this cropped document region (language: kn). "
"Return only the text, preserving its language and punctuation."},
]}],
}).json()
print(reply["choices"][0]["message"]["content"])
With transformers and PEFT:
import torch
from peft import PeftModel
from PIL import Image
from transformers import AutoProcessor, Gemma4ForConditionalGeneration
processor = AutoProcessor.from_pretrained("google/gemma-4-E4B-it")
model = Gemma4ForConditionalGeneration.from_pretrained(
"google/gemma-4-E4B-it", dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, "FineEnvs/gemma-4-E4B-it-kannada-ocr-grpo")
inputs = processor.apply_chat_template([{"role": "user", "content": [
{"type": "image", "image": Image.open("crop.png").convert("RGB")},
{"type": "text", "text": "Transcribe the text in this cropped document region (language: kn). "
"Return only the text, preserving its language and punctuation."},
]}], add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=768, do_sample=False)
print(processor.batch_decode(output[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])
Use the prompt above: it is the one the environment serves for every section-OCR task, and the one the adapter was trained on. To read a crop and score your own answer, try the playground.
How it was trained
| base model | google/gemma-4-E4B-it |
| method | GRPO (TRL 1.13), LoRA r=16, α=32 on the attention q/v projections of the language model and the vision tower (98 modules) |
| data | 4,000 distinct Kannada section crops streamed from the 164,366 in the Nayana train split (500 steps × 8) |
| rollouts | 8 transcriptions per crop, temperature 0.9, up to 512 tokens |
| reward | 0.8 × max(0, 1 − CER) + 0.2 × exact match, zero for overlong output, computed by the environment server |
| optimiser | lr 5e-5, 25 warmup steps, cosine decay, no KL penalty |
| hardware | one A100 80GB on HF Jobs, 10.5 hours |
| selection | step 300: best Sarvam CER of all 20 scored checkpoints; the final step 500 scores the same (0.3602) |
The benchmark is served by the same environment, as evaluation-only splits, with the same reward and Sarvam's CER/WER reported beside it. Training never sees it.
The trainer, the environment and both job launchers are in
06-multilingual-ocr/
on GitHub. Every checkpoint's held-out scores, the curves and the exact job scripts are in
FineEnvs/multilingual-multimodal-rl-runs,
and the training and evaluation runs are in the
Trackio dashboard.
Limits
One language. The crops cover printed books and periodicals; handwriting and scene text are out of scope. Held-out CER stopped improving after step 300 at this learning rate, so more steps on the same mix are unlikely to help.
The adapter is released under CC BY-NC 4.0 because its training data, NayanaOCR_Corpus_2025, is CC BY-NC 4.0. The base model is Apache-2.0.
Citation
@misc{fineenvs,
author = {Kolavi, Adithya S},
title = {FineEnvs: Open Source RL Environments for LLM Agents},
year = {2026},
url = {https://github.com/adithya-s-k/FineEnvs}
}
@misc{sarvam-indic-ocr-bench,
title = {Sarvam Indic OCR Bench},
author = {Sarvam AI},
year = {2026},
url = {https://huggingface.co/datasets/sarvamai/indic-ocr-bench}
}
- Downloads last month
- -