AI agent vs. chatbot: what's actually different
Scripted bots follow a flowchart. AI agents understand intent, pull from your knowledge base, and know when to hand off to a human.
“Chatbot” and “AI agent” get used as if one were a newer word for the other. They are different things, and the difference is not sophistication — it is who decides what happens next.
A chatbot follows a flowchart you drew. An agent works out what the customer wants and chooses from the actions you allowed. That distinction produces five practical differences, and it is worth knowing which of them you actually need before paying for either.
1. Who handles the question you did not anticipate
This is the whole difference, and everything below follows from it.
A flowchart can only answer questions on the flowchart. Every real inbox is full of questions nobody anticipated — not exotic ones, just phrased sideways. “Are you open Sunday” and “can I come this weekend” are the same question and a different branch.
A chatbot meets the second one with “Sorry, I didn’t understand that”, offers a menu, and the customer leaves. An agent recognises the intent and answers from the opening hours you gave it.
If your inbound questions genuinely are a short, closed list, a chatbot is cheaper and more predictable. Almost nobody’s are.
2. Where the answer comes from
A chatbot’s answers are written into the flow, so they are as current as the last time someone edited the flow. In practice that means prices drift, hours drift, and the bot confidently states things that stopped being true months ago.
An agent should answer from a source you maintain anyway — your catalogue, your price list, your policies, your calendar. Update the source, and every answer updates.
The question to ask any vendor: where does an answer come from, and what happens when the source and the model disagree? If the answer is “the model fills the gap”, you are buying a confident guesser. The behaviour you want is: answers only from what it was given, and an honest escalation for everything else.
3. What it does when it should stop
A chatbot ends. An agent hands over — and the quality of a deployment is mostly the quality of that handover.
Good escalation has three parts: it triggers on the right things (complaints, anything outside policy, an explicit request for a human), it routes to the right person rather than a shared pile, and it arrives with the conversation attached so nobody asks the customer to start again.
Most disappointment with “AI support” is not the AI answering badly. It is the AI not knowing when to stop.
4. Whether it can act, or only talk
A chatbot tells you the booking link. An agent can take the booking, in the conversation, against the real calendar.
This is where the value concentrates, and also where the risk does — an agent that can act can act wrongly. Scope it deliberately: which actions, on whose behalf, with what limits. “It can read everything and do nothing” is a reasonable place to start, and a reasonable place to stay for a while.
5. Whether you can tell if it is working
A chatbot reports completions: how many people reached the end of a flow. That number can be high while the experience is poor.
What you want to know is the share of conversations resolved without a person, how often it escalated, and what people actually asked for — including the things it could not answer, which is your list of what to teach it next.
If a tool cannot show you the questions it failed, it cannot be improved and neither can you.
Which one do you need?
A chatbot is enough if your questions are a genuinely closed set, you only need routing rather than answers, and the cost of a wrong answer is high enough that you want every sentence written by hand.
You want an agent if the same questions arrive phrased fifty ways, your answers live in information you already maintain, messages arrive outside working hours, or you want the conversation to end in a booking rather than a link.
How to evaluate one in an afternoon
Skip the demo script. Take the last fifty real conversations out of your inbox and run them through. Then count three things:
- How many it answered correctly from your own information
- How many it escalated — and whether those were the right ones
- How many it answered confidently and wrongly
The third number is the one that matters. It should be zero, and a tool that cannot get it to zero is not ready to talk to your customers, however well it performs on the first.
Fellix’s agent is briefed on your business, answers only from what you gave it, and escalates with the whole conversation attached. See how you brief it.