Enterprise Strategy

A 99% Accurate Agent Shouldn't Move Money Alone. Banks Draw The Line Where Mistakes Can't Be Undone.

September 29, 2026

Wipro's Alok Kumar explained why an agent's confidence never equals authority in banking, and where a human still has to make the final call.

A 99% Accurate Agent Shouldn't Move Money Alone. Banks Draw The Line Where Mistakes Can't Be Undone.
Credit:
powered by

Make State of AI one of your go-to sources on Google

Google Icon
Add thestateofai.com on Google
Quote Icon
Even if our AI becomes 97, 98, 99% accurate, these cases still require a human in the middle to validate before giving the final authority to AI to take decisions. If something wrong happens, it becomes very difficult to roll back.

Alok Kumar

Chief Enterprise AI Architect
Wipro

An AI agent can decide a payment in seconds. If that payment reaches settlement, there may be no practical way to take it back. An IMF note on agentic payments this spring drew a line between the agent's decision and the systems that actually authorize and settle the payment. The FSB's June consultation similarly called for guardrails around newer forms of AI, including agents. For banks, the question is becoming where an agent gets to act alone, and where a person has to take over.

Alok Kumar is Chief Enterprise AI Architect in the R&D AI vertical at Wipro, the global technology services and consulting firm. Kumar has architected fraud detection and anti-money laundering systems for banking clients, including work supporting First Abu Dhabi Bank, and now designs multi-agent orchestration for regulated customer onboarding. For Kumar, the architecture question starts with deciding where people belong in the workflow.

"We cannot put a human in the middle for all stages. The stages with critical decision-making, financial transactions, or some critical action after the response are the places where we require a human in the middle," Kumar says. His onboarding systems hand agents the document verification, eKYC checks, credit bureau pulls, and fraud scoring that once moved by hand. That work can be substantially cheaper to automate, but the design challenge is deciding which outputs can trigger an action on their own.

Probability isn't permission

Vector search returns a number that can look more authoritative than it is. Kumar's objection is to how often teams let that number decide what happens next. "We should not only depend on probabilistic search. It's very fast, but it gives the probability, like 0.9 or 0.8. We take the highest probability as the final result, but it isn't correct all the time. Sometimes it fails," he says.

His architecture stacks retrieval methods so each one checks another. Semantic search runs alongside BM25 keyword matching and relational or graph queries. A dedicated validation agent then reviews the combined output, a version of the agent-on-agent monitoring the FSB's consultation says some deployments may require. An applicant whose uploaded ID lists a different address than the application is the kind of conflict a single similarity score resolves confidently and wrongly.

"If we have multiple gates and one gate misses the validation, the second or third will at least catch it before we deliver the final response," Kumar explains. Each gate raises confidence in the output, but confidence alone doesn't give an agent permission to act. Authority comes from the rules around the decision.

Fraud scoring already sorts decisions by consequence

The clearest example runs inside card fraud systems, where Kumar walks through the process score by score. Each transaction gets rated on signals like location and amount, so a customer who normally transacts in India and suddenly moves a large sum from the UK draws a higher score. Low scores clear automatically, and mid-range scores trigger a step-up check by OTP or email. Anything scoring six to eight gets blocked and routed to a person. In the example Kumar gives, around 80% of the transactions that reach a person turn out to be legitimate. He argues the tier earns its cost in the remaining fraction it stops.

The same approach can apply beyond fraud: let the system handle routine decisions, add checks when the risk rises, and bring in a person when the consequences are harder to reverse. Routine decisions stay fully automated and ambiguous ones get an automated second check. Consequential ones stop and wait. Mastercard takes a similar approach with agentic payments, using preset spending limits to define how far an agent can go on its own.

A human checkpoint only works when the human has a job

In Kumar's fraud example, the reviewer has information and powers the model lacks. They can check the customer's history, call to verify the transaction, and hold the funds until the question is settled. "The human will check the history of that person, because AI is not that intelligent to check everything. They call the person and verify the transaction, and until then the money is blocked," he says.

For Kumar, that makes the reviewer part of the decision itself, rather than a final approval step. "If I'm the human in the middle and I get a request, I should make sure things are really correct before I approve. It's also a question on the person doing the review. It's not only a stamp. It's a required and mandatory step."

Accuracy and economics both draw the line

The boundary carries a price, too. Banks have long paid vendors to run static fraud rules, so replacing those systems with in-house models can look like an obvious way to cut costs. But the bill doesn't disappear. It moves into model usage, retrieval and orchestration, where token costs can be hard to predict as workflows grow more complex. Kumar says he has seen companies cut staff for AI, run into token costs they hadn't planned for, and start hiring again. He traces those failures to architecture and points to caching at multiple layers, so repeated requests never reach the vector database or the model.

That changes how architects have to think about the tradeoff. A workflow with several agents, layered retrieval, validation, and human review still has to justify itself against the legacy system on cost as well as accuracy. Planning for token costs upfront lets teams keep human checkpoints where the risk warrants them without pricing the system out of production.

Falling error rates won't necessarily change the architecture. Kumar's custom models already run around 95% accuracy, and he expects that number to climb. For banks, a model's error rate matters most when one wrong answer can be difficult to reverse.

"Even if our AI becomes 97, 98, 99% accurate, these cases still require a human in the middle to validate before giving the final authority to AI to take decisions. If something wrong happens, it becomes very difficult to roll back," Kumar says.

Outlever Logo

If this caught your attention, that’s not accidental.


Text Decoration Line

The best editorial systems don’t happen by accident. Outlever builds them.

Decorative Circular LinesDecorative Circular LinesDecorative Circular Lines Mobile

Get the latest AI insights first.

Sign up for updates, interviews, and fresh analysis on how AI is reshaping business, brands, and technology.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.