What would your project cost with us? Describe it in a few lines and see our range in two minutes. Get an estimate

What is RAG development and when does your company need it?

RAG development is the engineering work of connecting a large language model to your own knowledge so that answers come from your documents, not from what the model memorized during training. A company needs RAG when people keep asking questions whose answers already exist somewhere internal: support macros, product manuals, contracts, policies, past tickets, a wiki nobody can search. Instead of retraining a model, a RAG system retrieves the few passages that matter for each question and hands them to the model with instructions to answer from them and cite them.

Typical RAG implementations we are asked about:

  • An internal assistant that answers staff questions from policies, SOPs and the wiki.
  • A customer support assistant grounded in help center articles and resolved tickets (see AI chatbot development).
  • Search and question answering over contracts, specs or technical documentation.
  • A "memory" layer for an AI product, so the assistant remembers facts about a specific user or account.

How does a production RAG pipeline work?

A production RAG pipeline has two halves: an offline ingestion side that prepares your knowledge, and an online query side that answers questions. Most accuracy problems come from the ingestion side, which is why we spend real engineering time there instead of treating it as a one-off script.

Stage What happens Where it usually goes wrong
1. Ingestion Connectors pull content from files, Confluence, Google Drive, Notion, databases, ticketing tools or APIs Missing sources, no incremental sync, stale copies
2. Cleaning and parsing Text, tables and headings are extracted; duplicates and boilerplate removed; scanned PDFs sent through OCR Tables flattened into noise, headers and footers repeated in every chunk
3. Chunking Documents split into passages with metadata (source, section, date, access level) Chunks cut mid-sentence or too large to be specific
4. Embeddings Each chunk turned into a vector with an embedding model Model mismatched to language or domain
5. Storage Vectors and metadata stored in a vector database No filtering by permission or date
6. Retrieval Hybrid search (vector plus keyword), then reranking of the top results Right answer exists but is ranked 15th
7. Generation The LLM answers from the retrieved passages, with citations Model ignores the context or blends in outside knowledge
8. Evaluation and monitoring A test set and live feedback measure accuracy over time Nobody notices quality dropping after a data change

How should documents be chunked for RAG?

Chunking should follow the structure of your documents, not a fixed character count. A good default is to split by headings and paragraphs into passages of a few hundred tokens, keep a small overlap, and attach metadata to every chunk: document title, section path, date, language and who is allowed to see it. The heading path matters because a chunk that says "the limit is 30 days" is useless without knowing it came from "Refunds > EU customers".

Specific content types need specific handling. Tables are kept whole or converted into row-level statements. FAQs are split one question per chunk. Long contracts are chunked by clause with the clause number preserved. Chat logs and tickets are grouped by conversation, not by line. We test two or three chunking strategies against the same evaluation set and keep the one that retrieves the right passage most often, instead of guessing.

Which embeddings and vector database should you use?

The embedding model and vector database should be chosen for your data volume, languages and hosting rules, and both should be replaceable later. We keep them behind a small interface in the code so that switching providers is a re-index job, not a rewrite.

Option Good fit Trade-offs
pgvector (PostgreSQL) Teams already on PostgreSQL, up to a few million chunks, permissions stored in the same database Tuning needed at larger scale; fewer built-in search features
Qdrant Larger collections, rich metadata filtering, self-hosting in your own cloud One more service to operate
Weaviate or Milvus Very large collections, hybrid search built in Heavier operations footprint
Pinecone (managed) Teams that want zero database operations Vendor dependency, data leaves your infrastructure
OpenSearch or Elasticsearch with vectors Companies that already run it for keyword search Vector features less mature than dedicated engines

For embeddings, hosted models from OpenAI and similar providers are the fastest start. Open embedding models run on your own servers when data cannot leave your environment. Multilingual content (for example English plus French or Ukrainian) needs a model tested on those languages, which we check during evaluation rather than assume.

How do you measure RAG accuracy?

RAG accuracy is measured with a fixed evaluation set: 50 to 200 real questions from your users, each with the expected answer and the source it should come from. Without that set, every discussion about quality is opinion. With it, every change to chunking, prompts, models or data gets a score you can compare.

We track four numbers:

  1. Retrieval hit rate: did the correct passage appear in the top results?
  2. Faithfulness: does the answer only state what the retrieved passages say?
  3. Answer correctness: does the answer match the expected one, judged by a reviewer or an LLM grader checked against human samples?
  4. Refusal quality: when the answer is not in the knowledge base, does the system say so instead of inventing one?

In production we add user feedback (thumbs up or down with a reason), logging of retrieved sources per answer, cost per request and latency. The evaluation set runs automatically before each release, the same way unit tests do.

What about data privacy, permissions and model choice?

A RAG system must never show a user a passage they could not open in the original source. We store access rules as chunk metadata and filter at retrieval time, so permissions are enforced before the model sees anything. Sensitive data can be encrypted at ingest; in our AI Grief Companion project every message is encrypted with per-message AES-256-GCM under a KMS envelope key, and a user can delete everything in one click.

For the generation step we work with the OpenAI and Anthropic Claude APIs and with open models served on your own infrastructure (the same project runs Qwen2.5-7B adapters through vLLM). The choice depends on data rules, languages, latency and cost per answer. Prompts live in the database with version history where it helps, so tone and instructions can be tuned without a code release.

What does RAG development cost?

RAG development cost depends mostly on how many sources you connect, how messy they are and how strict the accuracy and permission requirements are. A proof of concept on one source with an evaluation set typically takes 3 to 6 weeks; a production assistant with several sources, permissions and a UI takes 2 to 4 months; platform-level RAG across products with custom models takes longer.

Running costs are separate: model tokens per answer, embedding costs on each re-index, and hosting for the vector database. We estimate cost per 1,000 questions during the proof of concept so there are no surprises, and give a fixed quote for the agreed milestones after a short scoping call. A RAG assistant over your knowledge starts from $5,000 with us; our AI chatbot development cost guide covers larger scopes.

Where have we built RAG and LLM systems?

Our most complete retrieval work is the AI Grief Companion for a US startup: a Python ingestion gateway parses ten chat-export formats into one normalized message table, and an ML pipeline builds a personality profile, a fact graph in Neo4j, RAG memory and LoRA adapters. The chat always uses whichever stage is ready, so users can talk to the system before training finishes.

On the product side, AI Resume Master, our own resume builder, uses LLMs to generate, rewrite and improve resume content and cover letters, and reached 50,000 monthly active users. For the broader picture of adding LLM features to an existing product, see AI integration services and our RAG implementation guide.

How we work on RAG projects

We start with a short scoping call and discovery: we look at your sources, collect real questions, and build the evaluation set together with your team. Then we deliver a proof of concept on one source with measured accuracy, and only after that scale to more sources, permissions and a production interface. Senior engineers own the architecture; internally we use AI coding agents to move faster on the plumbing.

If you have a knowledge base that people keep asking the same questions about, book a scoping call and bring five example questions. We will tell you honestly whether RAG is the right tool.

Case studies

Frequently asked questions

RAG (retrieval-augmented generation) searches your own content at question time and gives the most relevant passages to the model, which then answers from them and cites them. A general chatbot does not know your internal documents, cannot respect who is allowed to see what, and cannot show where an answer came from. RAG adds those three things while keeping your data in a store you control.

Use RAG when the model needs facts that change: policies, product docs, tickets, contracts. Use fine-tuning when the model needs a style, a format or a narrow behavior it keeps getting wrong. Most business assistants need RAG first. Fine-tuning comes later, if at all, and the two combine well: in our AI Grief Companion project a LoRA adapter carries the voice while retrieval carries the facts.

A focused proof of concept on one knowledge source with an evaluation set typically takes 3 to 6 weeks. A production system with several sources, permissions, incremental sync, monitoring and a user interface typically takes 2 to 4 months. The biggest variable is not the model but the state of the source data: scanned PDFs, duplicates and outdated pages add cleanup work.

If you already run PostgreSQL and have up to a few million chunks, pgvector is usually the simplest choice because it keeps vectors next to your relational data and permissions. Qdrant or Weaviate make sense for larger collections, heavy filtering or dedicated scaling. Managed services like Pinecone reduce operations work but add a vendor. We pick after looking at data volume, filters and hosting rules.

Hallucinations cannot be removed completely, but they can be measured and reduced. We instruct the model to answer only from retrieved passages and to say when it does not know, show citations for every answer, add reranking so the right passages reach the model, and run a fixed evaluation set on every change. Answers below a confidence threshold go to a human or return a safe fallback.

Let’s start your project
Book a call