
FiMI Banking Uses Verifiable Rewards to Beat a 12B Baseline on Its Own Tasks
NPCI's new FiMI Banking study makes a practical claim about financial AI: a 4.5-billion-effective-parameter model, trained inside a replayable bank environment, raised held-out reward from 0.610 to 0.697—just above a 12-billion-parameter baseline's 0.690—while generating 29% fewer tokens per dialog. The result does not mean small models generally beat large ones. It holds on 1,000 synthetic Indian retail-banking tasks with fixed tools, database states and rewards; even after training, 21.8% of those tasks failed in both trials. That boundary is the point. Exact tool order and final account state carried most of the useful signal, while a separate judged evaluation relied on one language-model judge. A concurrent preregistered study found severe instability in black-box LLM observers, strengthening the case for code-verifiable gates. GitHub's HydraFusion preview and Gimlet's $300 million financing show the same commercial pressure from different directions: spend computation selectively, measure the whole workflow, and do not confuse a cheaper path with a proven production system.









