Ringg's AI agents handle 7M calls a month and resolve up to 65% of customer requests
Synopsis
Ringg built an enterprise agent platform centered on GPT‑5.6 family models spanning voice, chat, WhatsApp, and the web, using an orchestration layer, knowledge retrieval, and multi-model routing to complete multi-step tasks across CRMs, ticketing, payment, and scheduling systems; it now handles more than 7 million connected calls each month, resolves up to 65% of routine customer inquiries without human involvement, reports an average customer CSAT of 4.8, and cut model costs by roughly 90% after migrating suitable real-time workloads from GPT‑4.1 to GPT‑5.6.
Interpretation
The platform runs on a division of labor across models: GPT‑4.1 handles most real-time voice and chat traffic, GPT‑5.6 Luna is used when its performance, latency, or price-performance profile fits better, GPT‑5.6 Terra handles post-call analysis including summaries and sentiment classification, and GPT‑5.6 Sol supports evaluation, prompt improvement, and model-as-judge workflows. Rather than having one model carry every task, real-time interaction, post-call analysis, and evaluation improvement are split across models and routed by task need. The text describes each model's role and the routing logic at the platform-architecture level, without per-model comparative experiment data.
In an evaluation of the post-call analysis workflow, GPT‑5.6 Terra came out on top against alternatives such as Gemini 2.5 Flash, and Ringg moved summaries and sentiment classification to Terra, maintaining high accuracy while materially improving unit economics. This is the one model evaluation in the text with a named comparison target, tying model choice directly to cost structure. The text says evaluation used historical conversations and simulated customer flows and reports up to 97% accuracy on common regional languages, but does not disclose sample sizes or statistical detail.
On the customer side, agents resolve up to 65% of routine inquiries without a human agent; Policybazaar connects more than 57,000 customer requests with 67% of calls handled without human intervention and average response time falling from 8–12 minutes to under 60 seconds; Practo reached an 85% first-call resolution rate with response times below three seconds, operating costs down 70%, and more than 1,000 appointment bookings per day; Groww resolves 72% of inbound IPO, futures, and options queries through self-service with an average handling time of two minutes. These are post-deployment operational metrics from named customers, grounding platform capability in measurable business outcomes. All are customer-side operational figures reported in the text by the deploying parties, without independent audit or controlled-experiment description.
On the engineering side, the system creates a structured summary when context approaches approximately 80,000 tokens to continue long conversations; models are tested offline on historical conversations and simulated flows before entering a small share of production traffic; and the production router monitors latency and endpoint health across regions, shifting traffic when an endpoint becomes unavailable or crosses a latency threshold. Context management, staged rollout, and traffic shifting are presented together as parts of platform reliability. This is a description of system design and operational process; the text does not provide failure rates or latency distributions for these mechanisms.
Perspective
This material speaks to teams evaluating enterprise voice and chat agent deployment, especially high-volume consumer businesses in markets such as India that need multilingual and cross-channel service. It shows that, given supporting engineering such as an orchestration layer, knowledge retrieval, staged rollout, and latency routing, assigning real-time interaction, post-call analysis, and evaluation improvement to different models can balance quality, latency, and cost. The browser agents and cross-channel context layer mentioned remain under development, and their applicable scope awaits further disclosure.
The text does not disclose sample sizes, statistical methods, or control settings for the model evaluations, nor the measurement definitions and time windows behind the 65% resolution rate and 4.8 CSAT; customer-side metrics are reported by the deploying parties without independent verification. In addition, the browser agents and cross-channel context layer are still in development, and their actual performance and applicable conditions await further information.
