HMDeveloper
All articles
agentes-ia
agentes-iaragatendimento

The five decisions that define the cost of a RAG-powered customer support agent

When a small business owner asks how much an AI agent costs, the answer lies in five technical decisions that start well before the first line of code.

When a small business owner asks "how much does an AI support agent cost?", the honest answer is not a number. It is a list of five decisions nobody shows you in the demo.

Each of these decisions multiplies or divides the final cost. And the worst part: several are ignored until the system breaks in production.

1. Channel and volume

An agent that answers via website chat has one cost profile. An agent that answers via WhatsApp, with voice, during business hours, has another.

Volume matters more than channel. If you get 30 questions a day, a small local model handles them with acceptable quality at near-zero cost. If you get 3,000, the local model may not keep up with latency, and every question that escalates to a cloud model costs tokens.

The first decision is measuring: how many conversations per day, on which channels, at which times. Without that number, any quote is fiction.

2. What goes into the knowledge base and who maintains it

RAG is only as good as the content it consults. If the base has an outdated manual, a generic FAQ, and a crooked scanned PDF, the agent answers confidently and gets it wrong with conviction.

Someone needs to decide what goes in, keep it updated, and test it. That is ongoing human work — not model training, content curation. In systems I have built, document ingestion is automated, but quality validation still goes through a person. If you have no one for that role, the real cost is the cost of the wrong answers the agent will give.

3. Per-task model and local → cloud routing

Not every question needs the most expensive model. "What are your business hours?" can be answered by a small local model. "Does product X cover case Y under regulation Z?" may need a larger cloud model.

The pattern I apply is routing: the local model tries first. If the answer confidence falls below a threshold, or if the question demands reasoning the local model demonstrably cannot deliver, it escalates to cloud.

This turns the cost from "every question pays full price" to "only the hard questions pay full price". At high volume, the difference is the bill fitting the budget or not.

4. When the agent says "I don't know" and escalates to a human

RAG always returns a top result — cosine similarity does not judge relevance, it only orders vectors. As I wrote in another article, a good system knows when the retrieved chunk does not help.

Translated to cost: an agent that never says "I don't know" produces wrong answers with sources. An agent that says "I don't know" too often frustrates the customer. Between these two extremes sits a product decision: at what similarity threshold does the agent escalate to a human?

That decision determines how many human agents you still need to keep. And human agents are the most expensive line on the spreadsheet.

5. Observability and cost per conversation

Without logs, you do not know what each conversation costs. Without knowing the cost, you do not know whether the system pays for itself.

Minimum metrics: tokens consumed per conversation, share of conversations resolved without human escalation, p95 latency, and cost in your currency per conversation. Week 1 is low. Week 12, if nobody looked, it may have tripled because someone uploaded a 400-page PDF the agent reads in full on every question.

Where it breaks

A support agent is not a project with an end. It is a product with maintenance. If you treat it as a project — shipped, tested, closed — it works until the knowledge base goes stale, volume doubles, or the cloud model provider changes pricing. And that will happen.

If you want to explore this for your business, tell me about your context.

I work with teams adopting AI agents in their day-to-day engineering. If that is your situation, let us talk.

See consulting
The five decisions that define the cost of a RAG-powered customer support agent | Hugo Minari Diniz