AI customer support that answers from your own documents
Customer support is still the easiest place to prove AI's value: the questions repeat, the volume is high and the results are measurable. The catch is that a general chatbot does not know your return policy. Here is how support agents are grounded in your own documents, how to choose the right retrieval approach, and what to track.
By Monolith

Picture an outdoor gear shop with a busy website. The same questions arrive all day: can I return boots I have worn once, when will my tent ship, does this jacket's warranty cover a torn zipper? A general AI chatbot will answer every one of them fluently, and some of those answers will be wrong, because it has never read the shop's return policy or warranty terms. The fix is not a smarter model. It is giving the model your documents and making it answer from them.
That approach has a name, retrieval-augmented generation or RAG, and it sits behind most useful support AI. This guide explains it in plain terms, shows the published results from companies that measured it, and gives a ladder for deciding how much retrieval machinery a small business actually needs.
Why support is still the best first AI project
Support has everything a first AI project needs: high volume, repeating patterns and numbers you already track, such as time to resolution, repeat contacts and satisfaction. The best-known early result came from Klarna. In its first month, February 2024, its AI assistant handled 2.3 million conversations, two-thirds of Klarna's customer service chats, doing work Klarna put at the equivalent of 700 full-time agents across 23 markets and more than 35 languages. 1
Two details in Klarna's announcement matter as much as the headline. Customer satisfaction was on par with human agents, not higher. And these were launch-month figures from a company with millions of customers. A small business should treat them as proof the pattern works, not as a forecast, and should keep an easy path to a person for anything the assistant cannot settle. 1
What an AI support agent actually does
A useful support agent does more than chat. It sorts incoming questions, looks up the relevant policy or order, resolves what it safely can, and hands the rest to a person with the context already gathered. The division of labor below is the practical version for a small team.
| Job | What the AI agent does | What your team keeps |
|---|---|---|
| Triage | Reads the message, tags the topic and urgency, and routes it | Rules for what counts as urgent |
| Policy lookup | Finds the relevant passage in your docs and cites it | Keeping the documents current |
| Order and account questions | Checks status in connected systems and replies | Access permissions for those systems |
| Resolution | Handles returns, rebookings or FAQs within set limits | Refunds, exceptions and goodwill gestures |
| Escalation | Hands off with a summary, the customer's history and what was tried | Every conversation that needs judgment |

RAG in plain English
Retrieval-augmented generation works like an open-book exam. When a question comes in, the system first searches your own material, such as the return policy, warranty terms, help articles and past tickets, and pulls out the few passages that look relevant. It then hands those passages to the AI model with an instruction: answer using this. Because the answer comes from your documents, the agent can cite the passage it used, and a person can check it.
RAG is not always needed. Anthropic points out that if your whole knowledge base is under about 200,000 tokens, roughly 500 pages, you can simply include all of it in the prompt and skip the search step. Many small businesses fit comfortably under that line: a return policy, a warranty page, shipping rules and fifty help articles. That guidance dates from 2024, and newer models accept far more, so the threshold has only grown. 3
When plain search is not enough
Basic RAG usually searches by meaning, using what are called embeddings: a question about a torn zipper finds the warranty passage about damaged fasteners even though the words differ. Meaning-based search has a blind spot, though. It can miss exact identifiers. Anthropic's example is an error code: a search by meaning might miss the exact match for TS-999, where a plain keyword search would find it at once. 3
Anthropic's fix combines three ideas. Hybrid search runs keyword and meaning-based search together. Contextual retrieval adds a short description of where each passage sits in its document before indexing it, so a lone paragraph about refunds still knows it belongs to the boots policy. Reranking then re-scores the top results before they reach the model. In its tests, the combination cut failed retrievals by 67%. 3
When a knowledge graph is worth it
Some questions need more than one passage. Microsoft Research found that baseline RAG struggles to connect the dots when an answer depends on relationships spread across many documents, or on understanding a whole collection. Its GraphRAG approach first builds a map of the entities in your data and how they relate, then answers from that map. On simple lookups, Microsoft notes, both approaches perform well; the graph pays off on the multi-step questions. 4
LinkedIn published the clearest support example. Its customer service team treated past tickets as a knowledge graph, linking each issue to related issues, rather than as loose text. In LinkedIn's tests the method beat a plain-text baseline by 77.6% on a standard retrieval score, and after about six months in use by its support team, it reduced the median time to resolve an issue by 28.6%. 2
- Put it all in the promptSmall knowledge base, roughly under 500 pages. No search step at all.Move up whenYour documents outgrow the model's context or cost too much per question.
- Plain RAGSearch by meaning, pass the top passages to the model, cite them.Move up whenAnswers miss exact codes, SKUs or product names.
- Hybrid plus rerankingKeyword and meaning search together, with context added and results re-scored.Move up whenGood answers need several linked records, such as related past tickets.
- Knowledge graphMap entities and relationships first, then retrieve along the connections.Move up whenTop of the ladder: invest here only for multi-step questions at volume.
What to measure
Track the same numbers the published cases report, so you can compare honestly. Resolution time and repeat contacts show whether customers actually got their answer; Klarna reported both. Median resolution time for your own team shows whether agent assist is helping staff; that is LinkedIn's metric. Satisfaction should hold steady or rise, never fall. Add one AI-specific check: how often the agent's cited passage actually supports its answer, sampled weekly by a person. 12
| Metric | What it tells you | Check |
|---|---|---|
| Time to resolution | Customers get answers faster | Weekly, against a pre-launch baseline |
| Repeat contacts on the same issue | The first answer actually worked | Weekly |
| Share handled without a person | How much volume the agent absorbs | Weekly, alongside satisfaction |
| Customer satisfaction | Speed did not cost quality | Monthly; investigate any drop |
| Citation accuracy | Answers are grounded in your documents | Sample 20 conversations a week |
Start this week
- Export last month's support messages and group them into the ten most common questions.
- Collect the documents that answer them and fix anything outdated first; the agent will repeat your mistakes.
- Measure your baseline: average time to resolve and how many customers contact you twice about the same thing.
- Start with the smallest rung on the ladder that fits, often the whole knowledge base in the prompt.
- Require the agent to cite the passage it used and to hand off with a summary whenever it is unsure.
- Keep refunds, exceptions and complaints with a person, and review a sample of conversations every week.
Monolith's take: support AI works when it answers from your documents and hands off gracefully when it cannot. Most small businesses should start with their whole policy set in the prompt, add hybrid search only when exact codes get missed, and treat knowledge graphs as a later investment for complex, high-volume questions.
What is RAG in customer support?
Do I need RAG for a small business chatbot?
When is GraphRAG worth it?
Read for this feature. The numbers match the markers in the text.
- Klarna: AI assistant handles two-thirds of customer service chats in its first month (February 2024)klarna.com
- LinkedIn (Xu et al.): Retrieval-Augmented Generation with Knowledge Graphs for Customer Service Question Answering (SIGIR 2024)arxiv.org
- Anthropic: Introducing Contextual Retrieval (September 2024)anthropic.com
- Microsoft Research: GraphRAG, unlocking LLM discovery on narrative private data (February 2024)microsoft.com
Want a support agent that answers from your policies and hands off cleanly to your team? We build and tune them.
Want to apply this to a project? Here is the related service.
25 AI instructions to try on everyday work
Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.