A chatbot conducts a conversation and answers questions. An AI agent completes a task: it reasons about a goal, calls tools, reads the results and acts until the request is resolved or escalated. The practical difference is tool use and write access, not conversational quality - and it is a spectrum, not a binary.
Key takeaways
- The dividing line is action, not intelligence. A chatbot retrieves and replies. An agent selects a tool, calls it, reads the result and decides what to do next.
- It is a five-level spectrum. Scripted (L0), understanding (L1), grounded answering (L2), read-only tool use (L3), acting with write access (L4). Most products marketed as "AI agents" today are Level 2 or Level 3.
- Level 3 is the best value jump. Read-only tool access - order status, availability, account lookup - delivers most of the perceived intelligence at a fraction of the risk of write access.
- Capability trades against cost and latency. A scripted reply is effectively instant and nearly free; a grounded answer is one model call; an agent is several model calls plus tool round trips.
- The failure modes differ in kind. Chatbots fail visibly with dead ends. Agents fail quietly - a confident wrong answer, a wrong tool argument, a loop, or an action that should have required approval.
- You do not have to choose up front. Start at Level 2, add read-only tools where they pay off, and gate write actions behind approval.
Definitions, Without the Marketing
What a chatbot is
A chatbot is software that holds a conversation in order to answer a question or collect information. In its simplest form it is a decision tree: buttons, keywords, and a path drawn in advance by a person. In its modern form it uses natural language understanding to interpret whatever the visitor typed, and a large language model to phrase the answer using your own documented content. Either way, the interaction ends with a reply. The full definition, with history and types, is on what is a chatbot.
What an AI agent is
An AI agent is software that pursues a goal. Given a request, it reasons about what needs to happen, chooses from a set of tools it has been given - API calls, lookups, write operations - executes one, reads the result, and repeats until the goal is met or it decides to hand over to a human. The conversation is the interface, not the point. The longer definition and the classical agent taxonomy are on what is an AI agent, and agentic AI covers the umbrella term.
The sentence that actually separates them
A support chatbot can tell a customer your return policy. A support agent can look up the order, check eligibility against it, and start the return. Both begin with a conversation. Only one finishes the job. Everything else in this page is detail on that one difference.
The Terminology Mess, Sorted
Half the confusion in this comparison is vocabulary. Six terms are used interchangeably by vendors who mean six different things, and some of them contradict each other - a few platforms treat conversational AI as a subset of agents, others treat agents as an evolution of chatbots. Here is the working set we use throughout this page.
| Term | What it means here |
|---|---|
| Rule-based chatbot | A scripted decision tree. Buttons, keywords, and a fixed dialogue graph. No language model decides anything. |
| AI chatbot / conversational AI | Understands free-text phrasing using natural language understanding, then answers - usually from a knowledge base or a generated response. Still answers rather than acts. |
| Copilot / assistant | An AI that works alongside a human and suggests. A person reviews and approves before anything happens in a real system. |
| AI agent | An AI that works toward a goal: it reasons about the request, chooses and calls tools, reads the results, and continues until the task is done or it escalates. |
| Agentic AI | The umbrella term for systems built this way. Usually a property of the architecture, not a product category. |
| Multi-agent / orchestration | Several specialised agents coordinated by a router or supervisor. Relevant at scale; usually premature for a first deployment. |
Related definitions: conversational AI, AI agent, AI agent orchestration, intent recognition and slot filling.
The Capability Matrix
Two-column tables force a false binary and hide where almost every real product actually sits. Three columns are more honest: the scripted bot, the grounded AI answerer that most people mean when they say "AI chatbot", and the tool-using agent.
| Rule-based chatbot | AI chatbot (grounded) | AI agent (tool-using) | |
|---|---|---|---|
| Core job | Answer a question | Answer a question in the customer's own words | Complete a task end to end |
| What decides the next step | A fixed dialogue graph | The model, one turn at a time | The model, reasoning over a goal across turns |
| Unseen phrasing | Falls back or misroutes | Handled - intent is inferred, not matched | Handled, and used to pick a tool |
| Tool and API use | Only hardcoded calls wired per branch | None, or a single lookup | Selects from an available tool set at runtime |
| Multi-step work | Requires an explicit branch per path | Cannot chain actions | Decomposes the task itself |
| Memory | Session variables and slot filling | The conversation so far, within the context window | Conversation, working state, and retrieved context |
| Knowledge | Exact matches and keywords | Grounded in your documented content | Grounded content plus live system state |
| Write access | None, or one pre-wired form submission | None | Yes - scoped, permissioned, and audited |
| Typical latency | Effectively instant | One model call | Several model calls plus tool round trips |
| Cost per interaction | Near zero | One inference | Inference per reasoning step, plus tool execution |
| Main failure mode | Dead ends and silent misrouting | Confident wrong answers | Wrong tool, wrong argument, or a loop |
| Maintenance work | Add branches and intents | Improve the source content | Expand and constrain the tool set |
Latency and cost rows are directional, not benchmarked: they depend on your model, the number of reasoning steps and your own tool latency. Measure yours.
Start as a grounded chatbot, add tool actions later. Free forever plan, no card.
The Autonomy Spectrum: Five Levels
"Is this an agent?" is the wrong question, because the honest answer for most products is "partly". A more useful question is how much of the decision-making you have handed over. These five levels are a working framing rather than an industry standard, but they map cleanly onto what is actually being sold, and they let you say where your own project sits today and where you want it next quarter.
Scripted
What it is: Buttons and a decision tree. Every path was drawn by a person.
What the model decides: Nothing. The designer decided in advance.
Where it fits: Menus, routing, opening hours, simple qualification. Regulated wording that must be pre-approved.
Understanding
What it is: Free-text input is mapped to an intent, then a scripted response runs.
What the model decides: Which branch to enter.
Where it fits: FAQ deflection where the answers are stable and the phrasing is not.
Grounded answering
What it is: A language model answers from your documented content instead of a fixed script.
What the model decides: What the answer says, within the bounds of your content.
Where it fits: Support content that changes faster than anyone can rewrite a flow. This is where most 'AI chatbots' actually sit.
Reading tools
What it is: The model can call read-only tools - order status, availability, account lookup - and use the result in its answer.
What the model decides: Which tool to call and how to interpret what comes back.
Where it fits: 'Where is my order', 'am I still covered', 'what slots are free'. The largest single jump in usefulness, at the smallest jump in risk.
Acting
What it is: The model can also write: book, cancel, refund, update, escalate - within permissions and approval gates.
What the model decides: What to do, in what order, and when to stop and hand to a human.
Where it fits: Returns, rescheduling, plan changes. Requires guardrails, audit logging, and a defined blast radius.
Two observations worth more than the table itself. First, the largest usefulness gain per unit of risk is the step from Level 2 to Level 3: letting the bot read live state turns generic answers into specific ones without giving it the ability to break anything. Second, most of the discomfort people feel about agents is really discomfort about Level 4, which is a permissions and governance problem rather than an AI problem. Our agentic AI chatbots explainer and the agentic AI in customer service guide go deeper on both.
When a Chatbot Is the Right Choice
A chatbot at Level 1 or 2 is the pragmatic pick when your inbound volume is dominated by repeatable questions and simple capture: hours, pricing, policy, order status explained rather than looked up, booking a call, collecting a lead. It launches in an afternoon, it is easy to reason about, it answers instantly, and grounded in your own documentation it resolves a large share of routine conversations. Most teams should start here even if they intend to end up somewhere else.
There are also situations where a scripted Level 0 flow is not a compromise but the correct answer: wording that has to be pre-approved for regulatory reasons, extremely high volume where per-inference cost matters, and any path where a deterministic outcome is worth more than a fluent one. Our customer support chatbot and lead generation chatbot guides cover these jobs in depth, and the template library has flows you can start from rather than a blank canvas. If you are picking a platform, the chatbot builder comparison covers twelve of them with verified pricing.
When You Need an AI Agent
Reach for agentic behaviour when resolving a request means doing something across your systems rather than explaining it: processing a refund, rescheduling against a live calendar, changing a plan, checking coverage against a policy record, or qualifying a lead through branching questions and routing it to the right owner. These are precisely the conversations a chatbot escalates - which means they are also the conversations where the human cost sits.
A useful diagnostic: look at your escalation transcripts for a fortnight and sort them into "the bot did not know" and "the bot could not do". The first pile is a content problem and is fixed at Level 2 - see training a chatbot on your knowledge base. The second pile is the business case for Level 3 or 4, and it is usually far smaller and far more valuable than people expect. Put a number on it with the chatbot ROI method before you build anything.
What an Agent Actually Needs
The gap between an agent demo and an agent deployment is four things. Skip any of them and you have something that works in a screenshot.
1. Tools it is allowed to call
An agent is only as capable as the tools it has, and only as safe as the permissions on them. In practice this means authenticated API calls with a clear description of when each should be used, sensible defaults, and read and write scopes kept separate. The mechanism the model uses to invoke them is function calling; the emerging standard for exposing them consistently is Model Context Protocol, covered in what MCP is and the MCP server guide. On the platform side this is ordinary integration work: the chat API, API integration, the developer reference, webhooks, and the integrations directory or Zapier if you would rather not write it.
2. Grounding in your own content
An agent that cannot check what is true will confidently invent it. Grounding means the model answers from your documented content rather than from what it absorbed during training - see knowledge-base training and, for the theory, retrieval-augmented generation and RAG vs fine-tuning. Grounding is also the cheapest fix for the most common complaint about AI chat, which is covered in why your AI chatbot gives wrong answers and preventing hallucinations.
3. Memory, and knowing which kind you need
Three different things get called memory. Conversation history is what has been said so far, bounded by the context window. Working state is what the agent has learned during this task - the order ID it looked up, the eligibility answer it got back. Cross-session memory is what it remembers about a customer next time, which is the one with real privacy implications and the one most deployments should defer. Getting this distinction wrong is why agents either forget the order number three turns later or remember things they should not.
4. Guardrails, escalation and an audit trail
The non-negotiable layer. Scope permissions so the blast radius of a wrong decision is small. Put an approval gate in front of anything irreversible - refunds, cancellations, anything touching money or a legal record. Define a human handoff path with a real person on the other end, and design the handover rather than bolting it on: handoff best practices and live chat cover the mechanics. Log every tool call with its arguments and result, because "why did it do that" is unanswerable without it. Read AI guardrails, human in the loop, chatbot security risks and, if you operate in the EU, EU AI Act compliance for chatbots.
Visual flows plus webhooks and API access from the $19 Starter plan.
How Each One Fails
Comparisons of this kind usually skip failure, which is odd, because failure mode is the difference you will actually live with. The two categories fail in different ways and need different monitoring.
Chatbots fail loudly and cheaply. A dead end, a fallback message, an obviously wrong branch, a loop back to the main menu. Customers notice immediately and either rephrase or ask for a human. It is annoying, it is visible in your fallback rate, and it is straightforward to fix.
Agents fail quietly and expensively. Five patterns to watch for specifically:
- Confident wrong answers. The generic risk of any generative system - see hallucination - but worse in an agent because the wrong answer may already have triggered an action.
- Right tool, wrong argument. The agent correctly decides to issue a refund and refunds the wrong order. Validation belongs in the tool, not in the prompt.
- Loops. Reasoning steps that keep retrying without progressing, which burns tokens and time. Cap the number of steps and fail to a human.
- Prompt injection through untrusted content. The moment an agent reads a ticket body, an email, or a web page, that text can contain instructions. Prompt injection and defending against it are mandatory reading before you give an agent write access.
- Partial completion. Step three of five succeeded and step four did not, leaving a half-finished state that nobody owns. Decide in advance what rolls back and what escalates.
On the operational side, both categories inherit the ordinary problems of running on other people's platforms - documented error codes and rate limits per channel are what turn a two-day outage into a ten-minute fix, and performance monitoring is what tells you it happened at all.
Cost and Latency, Honestly
Capability runs in the opposite direction to speed and cost, and very few comparisons admit it. A scripted reply is a lookup: effectively instant, effectively free. A grounded AI answer is one model call, so it arrives in about the time it takes to read the question. An agent runs a loop - decide, call a tool, wait for your API, read the result, decide again - so its latency is the sum of several model calls and several network round trips, and its cost is per reasoning step rather than per conversation.
We are deliberately not quoting benchmark numbers here, because they would be meaningless: your latency is dominated by your slowest internal API, and your cost is dominated by your model choice and how many steps your prompts encourage. What is worth doing is measuring both from day one, and modelling the unit that matters. That unit is not cost per conversation but cost per resolved request against the human alternative. An agent that costs several times more per conversation than a scripted bot can still be dramatically cheaper than the ticket it replaced.
Two practical mitigations. Route by complexity: send the ninety percent of traffic that is a repeat question down the fast grounded path and reserve the agent loop for requests that actually need it. And stream partial output, so a three-second answer feels like a one-second answer. For the arithmetic, use the ROI framework alongside the numbers in reducing call centre volume; our own plan pricing is on the pricing page.
What You Measure Changes Too
This is the part most teams discover late. Chatbot programmes are measured on whether the conversation stayed with the bot: containment rate, deflection rate, fallback rate, first response time. Those metrics quietly assume that not escalating is the same as succeeding, which is fine for FAQ deflection and actively misleading for an agent.
Agent programmes need a different set: resolution rate and first contact resolution for whether the task actually finished; action accuracy, tracked separately, for what proportion of tool calls did the right thing; escalation rate with reasons rather than escalation rate alone; and cost per resolved request. Track action accuracy on its own line, because an agent can be articulate, well-reviewed and still acting wrongly - a high CSAT will not surface that. Benchmarks and definitions are in the KPI guide and containment rate benchmarks, and analytics is where they land in the product.
Getting From One to the Other
There is a genuine disagreement in this space about whether a chatbot can become an agent. One camp says the assets transfer directly; the other says true agents need an architectural rewrite. Both are right about different things, and the distinction is worth being precise about.
What transfers: your knowledge content, your channel deployment, your integrations and credentials, your escalation routing, and everything you learned about how customers actually phrase things. What does not: the flow logic itself. A decision tree encodes the sequence a designer chose, and an agent is specifically the thing that chooses its own sequence - so scripted branches become tools and constraints rather than being ported. On a platform that supports both, this is configuration work. If you built directly on a model API, it is a rewrite of your orchestration layer.
The sequence that works in practice:
- Ground first (Level 2). Point the bot at your real content and fix the answers before adding any capability. Most perceived "agent" wins are actually grounding wins.
- Add one read-only tool (Level 3). Pick the single lookup that appears most in your escalation transcripts - usually order or booking status. Ship that alone and measure it.
- Add one write action behind a gate (Level 4). Choose something reversible, require confirmation, log everything, and cap how many it can do before a human reviews.
- Widen slowly, on evidence. Every new tool is new surface area for the failure modes above. Add them when escalation data says so, not when a roadmap does.
Conferbot is built for exactly this progression: a visual flow builder for the scripted path, AI answering grounded in your own content, and webhooks and API calls for the actions - all on one bot, deployed to your website widget plus WhatsApp, Messenger, Instagram, Telegram, Slack, Microsoft Teams, Discord and LINE. See AI agents, the chatbot builder, the no-code editor, omnichannel deployment and worked use cases. If you are still choosing a platform, the comparison of twelve chatbot builders covers which ones support which levels, and ChatGPT vs a chatbot platform covers the build-versus-buy version of the question.
AI Agent vs Chatbot FAQ
What is the main difference between an AI agent and a chatbot?
A chatbot conducts a conversation; an AI agent completes a task. A chatbot retrieves an answer and stops. An AI agent reasons about a goal, calls tools, reads the results and keeps going until the task is done or it escalates - for example, looking up an order and starting a return rather than explaining the return policy. The difference is the ability to act, not the quality of the writing.
Is an AI agent better than a chatbot?
Neither is universally better; they fit different jobs, and cost and latency run in the opposite direction to capability. For frequent questions, routing and simple capture, a well-built chatbot answers in milliseconds at near-zero marginal cost. For requests that need action across systems - refunds, bookings, account changes - an agent resolves what a chatbot would escalate, at several model calls and several seconds per turn. Most teams end up running both, with the scripted path handling the common cases.
Is ChatGPT an AI agent or a chatbot?
Out of the box, a raw chat interface is a conversational assistant - closer to an AI chatbot. It becomes an AI agent when it is given tools, memory and permission to act toward a goal with limited supervision. The label depends on capability and configuration, not the brand name. In our five-level framing, a plain chat interface is Level 2; the same model with tool access is Level 3 or 4.
What are the levels of AI agent autonomy?
We use five practical levels. Level 0 is scripted: a decision tree with no model deciding anything. Level 1 adds natural language understanding, so free text is mapped to an intent. Level 2 grounds answers in your own content rather than a script. Level 3 adds read-only tool calls, so the bot can look up an order or check availability. Level 4 adds write actions - booking, cancelling, refunding - inside permissions and approval gates. This is a working framing rather than an industry standard, but it maps cleanly onto what vendors are actually selling.
What does an AI agent need that a chatbot does not?
Four things. Tools: authenticated API calls it is allowed to make, with clear descriptions of when to use each. Grounding: access to your documented content so it is not inventing answers. Memory: conversation state plus whatever working context the task needs. Guardrails: scoped permissions, approval gates for irreversible actions, an escalation path to a human, and logging of every action it took and why. Skip the fourth and you have a demo, not a deployment.
How much slower and more expensive is an AI agent?
Directionally: a scripted chatbot replies effectively instantly at near-zero marginal cost; a grounded AI answer is one model call; an agent is several model calls plus tool round trips, so it is measurably slower and costs more per conversation. Exact figures depend on the model, the number of reasoning steps and your tool latency, so measure your own rather than trusting a benchmark. The economic question is not cost per conversation but cost per resolved request against the human alternative.
How do AI agents fail, and how is that different from chatbot failure?
A chatbot fails visibly and cheaply: a dead end, a fallback message, an obviously wrong branch. An agent fails less visibly and more expensively: a confident wrong answer, the right tool called with the wrong argument, a loop that burns tokens without progressing, or an action taken that should have needed approval. Agents also inherit prompt injection risk once they read untrusted text such as ticket bodies or web pages. The mitigations are scoped permissions, approval gates for anything irreversible, and logging every tool call.
Do I need an AI agent or just a chatbot for my website?
Start from the job. If your goal is deflecting repeat questions and capturing contact details, a chatbot grounded in your documentation is the pragmatic choice and will be live in an afternoon. If customers routinely need something done - reschedule an appointment, check an order, change a plan - an agent resolves what a chatbot hands back. On a no-code platform you can begin at Level 2 and add tool access later, so it is not a permanent fork.
Can one platform do both a chatbot and an AI agent?
Yes, and that is increasingly the norm. Conferbot lets you build a scripted conversational flow, add AI answering grounded in your own content, and connect actions through webhooks and API calls on the same bot, deployed across your website and eight messaging channels. You are not choosing a category up front; you are choosing how much autonomy each flow gets.
How do I measure an AI agent differently from a chatbot?
The metric has to move from 'did it reply' to 'did it finish'. Chatbot programmes are usually measured on containment and deflection rate. Agent programmes should be measured on task success rate, resolution rate, cost per resolved request, escalation rate with reasons, and action accuracy - what proportion of tool calls did the right thing. Track action accuracy separately from answer quality, because an agent can be articulate and still act wrongly.
Start at Level 2. Move Up When the Data Says So.
Build a grounded chatbot today and add tool actions to individual flows as they prove out - on the same bot, across nine surfaces.
Build your first bot or agent free600 chats a month, free forever, no credit card required.