What is RAG development and when does your company need it?
RAG development is the engineering work of connecting a large language model to your own knowledge so that answers come from your documents, not from what the model memorized during training. A company needs RAG when people keep asking questions whose answers already exist somewhere internal: support macros, product manuals, contracts, policies, past tickets, a wiki nobody can search. Instead of retraining a model, a RAG system retrieves the few passages that matter for each question and hands them to the model with instructions to answer from them and cite them.
Typical RAG implementations we are asked about:
- An internal assistant that answers staff questions from policies, SOPs and the wiki.
- A customer support assistant grounded in help center articles and resolved tickets (see AI chatbot development).
- Search and question answering over contracts, specs or technical documentation.
- A "memory" layer for an AI product, so the assistant remembers facts about a specific user or account.
How does a production RAG pipeline work?
A production RAG pipeline has two halves: an offline ingestion side that prepares your knowledge, and an online query side that answers questions. Most accuracy problems come from the ingestion side, which is why we spend real engineering time there instead of treating it as a one-off script.
| Stage | What happens | Where it usually goes wrong |
|---|---|---|
| 1. Ingestion | Connectors pull content from files, Confluence, Google Drive, Notion, databases, ticketing tools or APIs | Missing sources, no incremental sync, stale copies |
| 2. Cleaning and parsing | Text, tables and headings are extracted; duplicates and boilerplate removed; scanned PDFs sent through OCR | Tables flattened into noise, headers and footers repeated in every chunk |
| 3. Chunking | Documents split into passages with metadata (source, section, date, access level) | Chunks cut mid-sentence or too large to be specific |
| 4. Embeddings | Each chunk turned into a vector with an embedding model | Model mismatched to language or domain |
| 5. Storage | Vectors and metadata stored in a vector database | No filtering by permission or date |
| 6. Retrieval | Hybrid search (vector plus keyword), then reranking of the top results | Right answer exists but is ranked 15th |
| 7. Generation | The LLM answers from the retrieved passages, with citations | Model ignores the context or blends in outside knowledge |
| 8. Evaluation and monitoring | A test set and live feedback measure accuracy over time | Nobody notices quality dropping after a data change |
How should documents be chunked for RAG?
Chunking should follow the structure of your documents, not a fixed character count. A good default is to split by headings and paragraphs into passages of a few hundred tokens, keep a small overlap, and attach metadata to every chunk: document title, section path, date, language and who is allowed to see it. The heading path matters because a chunk that says "the limit is 30 days" is useless without knowing it came from "Refunds > EU customers".
Specific content types need specific handling. Tables are kept whole or converted into row-level statements. FAQs are split one question per chunk. Long contracts are chunked by clause with the clause number preserved. Chat logs and tickets are grouped by conversation, not by line. We test two or three chunking strategies against the same evaluation set and keep the one that retrieves the right passage most often, instead of guessing.
Which embeddings and vector database should you use?
The embedding model and vector database should be chosen for your data volume, languages and hosting rules, and both should be replaceable later. We keep them behind a small interface in the code so that switching providers is a re-index job, not a rewrite.
| Option | Good fit | Trade-offs |
|---|---|---|
| pgvector (PostgreSQL) | Teams already on PostgreSQL, up to a few million chunks, permissions stored in the same database | Tuning needed at larger scale; fewer built-in search features |
| Qdrant | Larger collections, rich metadata filtering, self-hosting in your own cloud | One more service to operate |
| Weaviate or Milvus | Very large collections, hybrid search built in | Heavier operations footprint |
| Pinecone (managed) | Teams that want zero database operations | Vendor dependency, data leaves your infrastructure |
| OpenSearch or Elasticsearch with vectors | Companies that already run it for keyword search | Vector features less mature than dedicated engines |
For embeddings, hosted models from OpenAI and similar providers are the fastest start. Open embedding models run on your own servers when data cannot leave your environment. Multilingual content (for example English plus French or Ukrainian) needs a model tested on those languages, which we check during evaluation rather than assume.
How do you measure RAG accuracy?
RAG accuracy is measured with a fixed evaluation set: 50 to 200 real questions from your users, each with the expected answer and the source it should come from. Without that set, every discussion about quality is opinion. With it, every change to chunking, prompts, models or data gets a score you can compare.
We track four numbers:
- Retrieval hit rate: did the correct passage appear in the top results?
- Faithfulness: does the answer only state what the retrieved passages say?
- Answer correctness: does the answer match the expected one, judged by a reviewer or an LLM grader checked against human samples?
- Refusal quality: when the answer is not in the knowledge base, does the system say so instead of inventing one?
In production we add user feedback (thumbs up or down with a reason), logging of retrieved sources per answer, cost per request and latency. The evaluation set runs automatically before each release, the same way unit tests do.
What about data privacy, permissions and model choice?
A RAG system must never show a user a passage they could not open in the original source. We store access rules as chunk metadata and filter at retrieval time, so permissions are enforced before the model sees anything. Sensitive data can be encrypted at ingest; in our AI Grief Companion project every message is encrypted with per-message AES-256-GCM under a KMS envelope key, and a user can delete everything in one click.
For the generation step we work with the OpenAI and Anthropic Claude APIs and with open models served on your own infrastructure (the same project runs Qwen2.5-7B adapters through vLLM). The choice depends on data rules, languages, latency and cost per answer. Prompts live in the database with version history where it helps, so tone and instructions can be tuned without a code release.
What does RAG development cost?
RAG development cost depends mostly on how many sources you connect, how messy they are and how strict the accuracy and permission requirements are. A proof of concept on one source with an evaluation set typically takes 3 to 6 weeks; a production assistant with several sources, permissions and a UI takes 2 to 4 months; platform-level RAG across products with custom models takes longer.
Running costs are separate: model tokens per answer, embedding costs on each re-index, and hosting for the vector database. We estimate cost per 1,000 questions during the proof of concept so there are no surprises, and give a fixed quote for the agreed milestones after a short scoping call. A RAG assistant over your knowledge starts from $5,000 with us; our AI chatbot development cost guide covers larger scopes.
Where have we built RAG and LLM systems?
Our most complete retrieval work is the AI Grief Companion for a US startup: a Python ingestion gateway parses ten chat-export formats into one normalized message table, and an ML pipeline builds a personality profile, a fact graph in Neo4j, RAG memory and LoRA adapters. The chat always uses whichever stage is ready, so users can talk to the system before training finishes.
On the product side, AI Resume Master, our own resume builder, uses LLMs to generate, rewrite and improve resume content and cover letters, and reached 50,000 monthly active users. For the broader picture of adding LLM features to an existing product, see AI integration services and our RAG implementation guide.
How we work on RAG projects
We start with a short scoping call and discovery: we look at your sources, collect real questions, and build the evaluation set together with your team. Then we deliver a proof of concept on one source with measured accuracy, and only after that scale to more sources, permissions and a production interface. Senior engineers own the architecture; internally we use AI coding agents to move faster on the plumbing.
If you have a knowledge base that people keep asking the same questions about, book a scoping call and bring five example questions. We will tell you honestly whether RAG is the right tool.