Sarvam AI introduces Vision 2.1, which improves document reading and recognizes Indic handwriting


Homegrown AI startup Sarvam AI has launched Sarvam Vision 2.1, an upgraded version of its vision-language model built to understand and process documents. The latest model brings enhanced capabilities for handling complicated tables, extracting information from structured forms and identifying handwritten text in Indian languages.

According to the company, Vision 2.1 addresses several limitations identified in its previous version. This time, the focus extends beyond simply reading documents to converting messy and unstructured scans into clean, organised data that businesses can directly integrate into production workflows.

What can Sarvam Vision 2.1 do

The latest model comes with stronger tabular processing capabilities. In practical terms, it can handle nested, complicated and multi-page tables that conventional vision models often find difficult to interpret accurately. It can also pull key-value information from forms, including names, dates and financial figures, while recognising handwritten notes and annotations in Indic languages.

Sarvam said it introduced specific training techniques to minimise hallucinations and spatial inconsistencies, which are common problems for AI systems processing dense or irregularly formatted document scans.

Benchmark performance

Alongside the launch, Sarvam introduced a new Sarvam Indic OCR benchmark covering all 22 official Indian languages. The evaluation set consists of 6,909 samples, including 6,609 examples from the 22 regional languages and 300 English baseline samples. The dataset includes newspapers, brochures, textbooks and historical archives spanning the period from 1800 to the present.

On the benchmark, Sarvam Vision 2.1 achieved an overall word accuracy score of 87.39 per cent. It also recorded a score of 87.30 on the global olmOCR-Bench, showing competitive results against international document-processing standards.

Training and deployment

Sarvam trained the new model using a combination of synthetic and real-world document scans, with particular emphasis on regional handwriting variations and extracting information from forms. The model was subsequently fine-tuned through supervised learning and then enhanced using reinforcement learning with verifiable rewards (RLVR). This process was intended to improve exact structural and textual accuracy rather than simply generating output that appears plausible.

The model is available through Sarvam’s document intelligence APIs, which offer developers tools to convert multi-page documents into structured text and automatically extract information from forms and tables.


 

buttons=(Accept !) days=(20)

Our website uses cookies to enhance your experience. Learn More
Accept !