Skip to main content
Back to timeline
arXivSource publication:

IndicBankBench tests 799 Indian retail-banking cases and finds eleven models reach only 43.7%–58.2% strict three-run reliability while at-least-once success runs 60%–74%

Synopsis

The authors release IndicBankBench, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes, graded in four stages—safety, action and tool use, response adequacy, and advisory quality—with deterministic tool-use and most safety checks, an LLM judge only for semantic response adequacy, and a narrow resolver for ambiguous confirmation-before-write; running every case three times across eleven models yields strict pass3 of 43.7%–58.2% versus at-least-once success of 60%–74%, a gap of 10.8–21.4 percentage points.

AI-generated editorial illustration: IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

Interpretation

The benchmark splits a banking-assistant interaction into four ordered stages—safety, action and tool use, response adequacy, and advisory quality—and records the stage at which an interaction first failed. Existing tool-calling benchmarks commonly summarize performance with aggregate accuracy or success metrics, whereas this one localizes failure to a stage so that fabricating an answer and asking for information already available can be told apart. 799 cases, 32 tools, 20 primary axes; safety and action checks are deterministic, the response stage is judged by GLM-5.2 at temperature 0 with seed 42, ambiguous confirmation-before-write cases go to a separate narrow resolver returning a yes/no signal, and the grading code rather than either model makes the final pass/fail decision.

With three repeated trials per case, strict pass3 and at-least-once success diverge markedly. Prior work used repeated samples to measure at-least-once success; here repeated trials distinguish occasional task completion from reliable completion and report Inconsistent and Failed-all case counts. 2,397 trajectories per model across eleven models; strict pass3 ranges from 43.7% (349/799, Gemma 4 E4B IT) to 58.2% (465/799, DeepSeek-V4-Pro-0813), at-least-once success from 60.1% to 74.3%, a gap of 10.8–21.4 percentage points; GLM-5.3-Flash has the highest at-least-once rate (594/799, 74.3%) but a strict rate of 56.8%.

Case-level diagnostics show that systems fail in different ways, and no reliable difference is detected among the top five. Aggregate scores cannot separate over-clarification from acting without reconciling context; per-axis results and gate counts make those differences visible. The top five span 56.8%–58.2% strict pass3, a difference of 11 cases; across all ten pairs every bootstrap interval includes zero and every exact McNemar test is non-significant both unadjusted and Holm-adjusted; the wrong-information axis is the lowest-scoring task axis for seven of eleven models, with no model passing more than 40.1% of its 142 cases; among 26,291 trajectories where both decisions could be determined, 3,390 (12.9%) passed every safety and action gate but failed R, while 1,497 (5.7%) passed R but failed at least one safety or action gate.

The authors provide a descriptive nine-pattern failure taxonomy and release the cases, mock environment, and evaluation harness. The taxonomy groups recurring failures into trusting the surface, knowing versus doing, and miscalibrated action, supporting diagnosis rather than ranking alone. The taxonomy was developed from case-level gate records and transcripts of a separate five-model screening run at temperature 0 with one run per case; the authors state it describes how failures occur and is not a frequency estimate for the eleven-model evaluation; code and case data are public at github.com/npci/IndicBankBench.

Perspective

The benchmark is aimed at evaluators and developers who want to diagnose where a banking assistant fails, in an English-language Indian retail-banking setting with synthetic customer state and deterministic mock tools; the authors release the cases, mock environment, and evaluation harness and note it can compare changes to an assistant or prompt under the same cases and evaluation settings, with a new version and rerun required if a case or grading rule changes.

Response adequacy depends on an LLM judge and ambiguous confirmation-before-write cases on a narrow resolver, while the blinded audit covers only 40 transcripts from the original eight evaluated models, so its agreement rate may not hold across the full benchmark; three repeated trials per case give only a limited estimate of run-to-run reliability, and results may depend on provider endpoints, serving behavior, and reasoning controls; cases use synthetic customer state and mock tools and do not cover every production policy, operational failure, or customer preference, nor do they establish performance across banks, jurisdictions, Indian languages, code-switched interactions, accessibility needs, or production authorization and fraud-monitoring systems; the nine failure patterns come from a separate five-model screening run and are a descriptive taxonomy rather than frequency estimates; this reading is full text, but the appendix prompt and grading artifacts appear as placeholders, so their exact instruction content cannot be checked here.

Sources