InsightsPublished 7 min read

AI customer support that answers from your own documents

Customer support is still the easiest place to prove AI's value: the questions repeat, the volume is high and the results are measurable. The catch is that a general chatbot does not know your return policy. Here is how support agents are grounded in your own documents, how to choose the right retrieval approach, and what to track.

By Monolith

A support rep wearing a headset smiling as she helps a customer at the service desk of an outdoor gear shop
A support rep at an outdoor gear shop helping a customer at the service desk.

Picture an outdoor gear shop with a busy website. The same questions arrive all day: can I return boots I have worn once, when will my tent ship, does this jacket's warranty cover a torn zipper? A general AI chatbot will answer every one of them fluently, and some of those answers will be wrong, because it has never read the shop's return policy or warranty terms. The fix is not a smarter model. It is giving the model your documents and making it answer from them.

That approach has a name, retrieval-augmented generation or RAG, and it sits behind most useful support AI. This guide explains it in plain terms, shows the published results from companies that measured it, and gives a ladder for deciding how much retrieval machinery a small business actually needs.

Why support is still the best first AI project

Support has everything a first AI project needs: high volume, repeating patterns and numbers you already track, such as time to resolution, repeat contacts and satisfaction. The best-known early result came from Klarna. In its first month, February 2024, its AI assistant handled 2.3 million conversations, two-thirds of Klarna's customer service chats, doing work Klarna put at the equivalent of 700 full-time agents across 23 markets and more than 35 languages. 1

Two details in Klarna's announcement matter as much as the headline. Customer satisfaction was on par with human agents, not higher. And these were launch-month figures from a company with millions of customers. A small business should treat them as proof the pattern works, not as a forecast, and should keep an easy path to a person for anything the assistant cannot settle. 1

What an AI support agent actually does

A useful support agent does more than chat. It sorts incoming questions, looks up the relevant policy or order, resolves what it safely can, and hands the rest to a person with the context already gathered. The division of labor below is the practical version for a small team.

JobWhat the AI agent doesWhat your team keeps
TriageReads the message, tags the topic and urgency, and routes itRules for what counts as urgent
Policy lookupFinds the relevant passage in your docs and cites itKeeping the documents current
Order and account questionsChecks status in connected systems and repliesAccess permissions for those systems
ResolutionHandles returns, rebookings or FAQs within set limitsRefunds, exceptions and goodwill gestures
EscalationHands off with a summary, the customer's history and what was triedEvery conversation that needs judgment
Monolith's suggested split of support work between an AI agent and your team. Planning guidance, not results from a specific deployment.
A bike shop mechanic reading a message on his phone at the counter while a customer waits holding a helmet
A bike shop mechanic reading a handed-off support summary while the customer waits.

RAG in plain English

Retrieval-augmented generation works like an open-book exam. When a question comes in, the system first searches your own material, such as the return policy, warranty terms, help articles and past tickets, and pulls out the few passages that look relevant. It then hands those passages to the AI model with an instruction: answer using this. Because the answer comes from your documents, the agent can cite the passage it used, and a person can check it.

RAG is not always needed. Anthropic points out that if your whole knowledge base is under about 200,000 tokens, roughly 500 pages, you can simply include all of it in the prompt and skip the search step. Many small businesses fit comfortably under that line: a return policy, a warranty page, shipping rules and fifty help articles. That guidance dates from 2024, and newer models accept far more, so the threshold has only grown. 3

When plain search is not enough

Basic RAG usually searches by meaning, using what are called embeddings: a question about a torn zipper finds the warranty passage about damaged fasteners even though the words differ. Meaning-based search has a blind spot, though. It can miss exact identifiers. Anthropic's example is an error code: a search by meaning might miss the exact match for TS-999, where a plain keyword search would find it at once. 3

Anthropic's fix combines three ideas. Hybrid search runs keyword and meaning-based search together. Contextual retrieval adds a short description of where each passage sits in its document before indexing it, so a lone paragraph about refunds still knows it belongs to the boots policy. Reranking then re-scores the top results before they reach the model. In its tests, the combination cut failed retrievals by 67%. 3

When a knowledge graph is worth it

Some questions need more than one passage. Microsoft Research found that baseline RAG struggles to connect the dots when an answer depends on relationships spread across many documents, or on understanding a whole collection. Its GraphRAG approach first builds a map of the entities in your data and how they relate, then answers from that map. On simple lookups, Microsoft notes, both approaches perform well; the graph pays off on the multi-step questions. 4

LinkedIn published the clearest support example. Its customer service team treated past tickets as a knowledge graph, linking each issue to related issues, rather than as loose text. In LinkedIn's tests the method beat a plain-text baseline by 77.6% on a standard retrieval score, and after about six months in use by its support team, it reduced the median time to resolve an issue by 28.6%. 2

The retrieval ladder: climb only when you must
  1. Put it all in the promptSmall knowledge base, roughly under 500 pages. No search step at all.Move up whenYour documents outgrow the model's context or cost too much per question.
  2. Plain RAGSearch by meaning, pass the top passages to the model, cite them.Move up whenAnswers miss exact codes, SKUs or product names.
  3. Hybrid plus rerankingKeyword and meaning search together, with context added and results re-scored.Move up whenGood answers need several linked records, such as related past tickets.
  4. Knowledge graphMap entities and relationships first, then retrieve along the connections.Move up whenTop of the ladder: invest here only for multi-step questions at volume.
Monolith's ladder for choosing a retrieval approach, based on Anthropic's contextual retrieval guidance, LinkedIn's knowledge-graph results and Microsoft Research's GraphRAG findings. [2][3][4]

What to measure

Track the same numbers the published cases report, so you can compare honestly. Resolution time and repeat contacts show whether customers actually got their answer; Klarna reported both. Median resolution time for your own team shows whether agent assist is helping staff; that is LinkedIn's metric. Satisfaction should hold steady or rise, never fall. Add one AI-specific check: how often the agent's cited passage actually supports its answer, sampled weekly by a person. 12

MetricWhat it tells youCheck
Time to resolutionCustomers get answers fasterWeekly, against a pre-launch baseline
Repeat contacts on the same issueThe first answer actually workedWeekly
Share handled without a personHow much volume the agent absorbsWeekly, alongside satisfaction
Customer satisfactionSpeed did not cost qualityMonthly; investigate any drop
Citation accuracyAnswers are grounded in your documentsSample 20 conversations a week
A starter scorecard. Metric choices follow those reported by Klarna and LinkedIn; the weekly citation check is Monolith's recommendation. 12

Start this week

  • Export last month's support messages and group them into the ten most common questions.
  • Collect the documents that answer them and fix anything outdated first; the agent will repeat your mistakes.
  • Measure your baseline: average time to resolve and how many customers contact you twice about the same thing.
  • Start with the smallest rung on the ladder that fits, often the whole knowledge base in the prompt.
  • Require the agent to cite the passage it used and to hand off with a summary whenever it is unsure.
  • Keep refunds, exceptions and complaints with a person, and review a sample of conversations every week.
The short version

Monolith's take: support AI works when it answers from your documents and hands off gracefully when it cannot. Most small businesses should start with their whole policy set in the prompt, add hybrid search only when exact codes get missed, and treat knowledge graphs as a later investment for complex, high-volume questions.

Questions, answered
What is RAG in customer support?
Retrieval-augmented generation means the AI first searches your own policies, help articles and past tickets, then answers using the passages it found, so it can cite its source instead of guessing.
Do I need RAG for a small business chatbot?
Often not at first. Anthropic notes that knowledge bases under about 500 pages can be placed directly in the prompt; add search when your documents outgrow that or exact codes start being missed.
When is GraphRAG worth it?
When answers depend on connecting several related records, such as linked past tickets. LinkedIn's knowledge-graph approach cut its support team's median resolution time by 28.6% over about six months.
Sources

Read for this feature. The numbers match the markers in the text.

  1. Klarna: AI assistant handles two-thirds of customer service chats in its first month (February 2024)klarna.com
  2. LinkedIn (Xu et al.): Retrieval-Augmented Generation with Knowledge Graphs for Customer Service Question Answering (SIGIR 2024)arxiv.org
  3. Anthropic: Introducing Contextual Retrieval (September 2024)anthropic.com
  4. Microsoft Research: GraphRAG, unlocking LLM discovery on narrative private data (February 2024)microsoft.com

Want a support agent that answers from your policies and hands off cleanly to your team? We build and tune them.

Where this leads

Want to apply this to a project? Here is the related service.

Free download

25 AI instructions to try on everyday work

Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.

Enter your email to access the PDF. This form does not sign you up for a marketing sequence.