Series · 12 parts
eMunshi Engineering Series
A technical series on the data foundation Inspire AI Lab built for eMunshi, a legal research and practice platform for Indian advocates: recovering text from tens of millions of court documents in multiple scripts, establishing case identity, verifying links between documents and cases, generating headnotes that cannot invent content, and maintaining a verified statute library. Each part records what was measured, what went wrong, and what was learned.

12 partspublishes Thursdaysfirst part 0 of 12 live
Next up Part 1: Working with Indian Court Records at Scale ·
Data foundation Parts 1–9
Recovering text from court documents in multiple scripts, establishing case identity, verifying links between documents and cases, headnotes that cannot invent content, and a verified statute library.
- Part 1Coming
- Part 2Coming
Pre-Unicode Legacy Fonts: Recovering Indic Text Encoded as ASCII
- Part 3Coming
OCR as Verification, Not Default
- Part 4Coming
Headnotes That Cannot Invent Content
- Part 5Coming
Party Extraction Across Inconsistent Formats
- Part 6Coming
A Verified Statute Library
- Part 7Coming
Case Identity: Why a Printed Case Number Is Not an Identifier
- Part 8Coming
Imperfect Source Data: Misfiled Documents, Missing Metadata, and Relinking
- Part 9Coming
Operating AI Coding Agents on a Large Document Pipeline
Product capabilities Parts 10–12
Citation verification, party identity in due diligence, and court-rules-aware drafting on the legal research and practice platform.
- Part 10Coming
Citation Existence Verification
- Part 11Coming
Resolving Party Identity in Due Diligence
- Part 12Coming
Court-Rules-Aware Drafting