What is a RAG implementation and how does it work?
A RAG implementation connects a large language model to your own knowledge so it answers from your documents instead of from memory. At request time the system searches your indexed content, selects the most relevant passages, inserts them into the prompt and asks the model to answer with citations.
A production RAG architecture has two pipelines:
- Ingestion pipeline: connect sources, parse files, clean text, split it into chunks, create embeddings, store chunks with metadata and access rules, and keep everything in sync.
- Query pipeline: understand the question, run hybrid search, rerank results, assemble the prompt, generate the answer, attach citations and log everything for evaluation.
Most RAG failures come from the first pipeline, not the model. If a table is parsed as noise or a policy is split in half, no model can recover the answer. This guide walks through each stage the way we approach it at Lytvynov Production. For the delivery side, see our RAG development services.
How should you ingest and parse documents?
Ingestion should turn every source into clean text with structure and metadata: title, headings, source URL, author, dates, document type, language and who is allowed to read it. Spend real time here, because parsing quality sets the ceiling for answer quality.
Common sources and what they need:
- PDFs and Office files: layout-aware parsing that keeps headings, lists and tables; OCR for scanned pages.
- Wikis and help centers (Confluence, Notion, Zendesk, Intercom): API connectors that also pull permissions and update timestamps.
- Tickets, chats and emails: thread reconstruction, deduplication of quoted replies and removal of signatures.
- Databases and product catalogs: convert records into short readable text, or query them directly with tools instead of embedding them.
Normalizing many messy formats into one clean table is often the hardest part. In the AI Grief Companion project, we wrote parsers for ten chat-export formats (WhatsApp, Messenger, Instagram, Discord, iMessage from an iPhone backup, Android SMS and others) and normalized all of them into one message table before any AI step ran. The same lesson applies to business data: invest in a single normalized model first. Our sports calendar platform shows the same discipline outside AI, with a parsing engine that unifies data from hundreds of external calendars.
What chunking strategy should you use?
Chunk by document structure first and by size second. Split on headings, sections, list items or message threads, then cap chunks at a size that holds one complete idea, typically 200-800 tokens with a small overlap.
| Strategy | How it works | Good for | Watch out for |
|---|---|---|---|
| Fixed size with overlap | Split every N tokens, overlap 10-20% | Quick baseline, uniform text | Cuts sentences, tables and lists in half |
| Structure-aware | Split on headings, sections, paragraphs | Documentation, policies, contracts | Very long sections need a second split |
| Semantic | Split where topic similarity drops | Long narrative text, transcripts | More compute at ingestion, harder to debug |
| Parent and child | Search small chunks, return the larger parent section | Precise search with enough context | More storage, more prompt tokens |
| Record-based | One chunk per ticket, product, FAQ entry or message thread | Structured or semi-structured data | Very short records may lack context |
Add context to each chunk: prepend the document title and section path ("Refund policy > EU customers > Digital goods") so the chunk makes sense on its own. This one step often improves retrieval more than switching embedding models.
How do you choose an embedding model?
Pick an embedding model that handles your languages and domain vocabulary, then test two or three candidates on your own golden set. Hosted embedding APIs from OpenAI and others are a sensible default; open-weight embedding models are a good choice when data must stay in your infrastructure.
Record the embedding model and version with every stored vector. Changing the model later means re-embedding the whole corpus, so plan that as a normal background job, not an emergency. For multilingual content, check that a question in one language retrieves documents written in another if your users need that.
Which vector database should you use for RAG?
Use the database your team can operate well. For most business RAG systems below tens of millions of chunks, retrieval quality depends far more on chunking, hybrid search and reranking than on the choice of vector store.
| Option | Type | Strengths | Trade-offs | Good fit |
|---|---|---|---|---|
| pgvector (PostgreSQL) | Extension to your existing database | One database for data, vectors and permissions; transactions; simple ops | Needs tuning at large scale; keyword search is basic without extra work | SaaS products already on PostgreSQL, up to several million chunks |
| Qdrant | Open-source vector database, self-hosted or cloud | Fast filtering on metadata, hybrid search support, efficient memory use | Another service to run and back up | Large corpora, heavy metadata filtering |
| Weaviate | Open-source vector database, self-hosted or cloud | Built-in hybrid search, modules for vectorization | More concepts to learn, heavier to operate | Teams wanting search features out of the box |
| Pinecone | Fully managed service | No infrastructure work, scales easily | Vendor lock-in, data leaves your cloud, usage-based cost | Teams without ops capacity |
| Elasticsearch / OpenSearch | Search engine with vector support | Mature keyword search (BM25), aggregations, existing ops knowledge | Resource-hungry; vector features vary by version and license | Companies already running them for search |
Our default for SaaS clients on PostgreSQL is pgvector, because permissions and tenant IDs live next to the vectors and there is one less system to secure. We move to a dedicated engine when scale or filtering needs justify it.
Why use hybrid search and reranking?
Hybrid search combines keyword search (BM25) with vector search, because each finds things the other misses. Vector search understands paraphrases; keyword search catches exact product codes, error messages, names and acronyms. A reranker then reorders the combined candidates by true relevance to the question.
A typical query flow: rewrite the user question into a standalone search query (resolving "it" and "that" from the conversation), run keyword and vector search in parallel with permission filters, merge the results (reciprocal rank fusion is a simple method), rerank the top 30-50 candidates with a cross-encoder or reranking API, and keep the top 5-10 for the prompt. Reranking usually gives one of the biggest quality gains per hour of engineering work in a RAG project.
How do you assemble the prompt and add citations?
Assemble the prompt with clear sections: system rules, retrieved passages each labeled with a source ID, the conversation so far and the question. Instruct the model to answer only from the passages, cite source IDs for each claim and say plainly when the sources do not contain the answer.
After generation, validate the citations in code: every cited ID must exist in the retrieved set, and answers without citations for factual claims can be flagged or regenerated. Show citations in the UI as links to the original document and section. Users trust answers they can check, and support teams can fix the source document instead of arguing with the AI. For the wider integration picture (streaming, guardrails, cost), see our guide on how to integrate ChatGPT or Claude into your product.
How do you handle permissions and access control in RAG?
Enforce permissions at retrieval time, before any passage reaches the model. Store access metadata (tenant ID, team, role, document ACL) with every chunk and apply it as a filter in the search query for the current user.
Never rely on the prompt to hide content ("do not reveal HR documents") because models can be talked out of instructions. Sync permission changes from source systems quickly, and remove chunks when documents are deleted. In multi-tenant SaaS, a tenant filter on every query is mandatory, and we add automated tests that try to retrieve another tenant's data. Sensitive content may also need encryption at rest; the AI Grief Companion platform encrypts every message at ingest with per-message AES-256-GCM and supports one-click deletion of everything.
How do you evaluate a RAG system?
Evaluate retrieval and generation separately, using a golden set of real questions. Without a golden set, every change is a guess, and teams end up tuning prompts against whichever example someone complained about last.
- Golden set: 100-300 real questions with reference answers and the source documents that should be found. Include questions with no answer in the corpus.
- Retrieval metrics: recall at k (did the correct passage appear in the top k?) and ranking quality.
- Generation metrics: faithfulness (is every claim supported by retrieved passages?), answer relevance, completeness and citation accuracy. A second model can grade these with a rubric, with humans reviewing a sample.
- Production signals: thumbs down, follow-up rephrasing, escalations to a human and unanswered questions.
Run the set on every change to parsing, chunking, embeddings, search settings, prompts or models, and block releases that lower scores.
How do you keep a RAG index fresh?
Keep the index fresh with incremental sync: detect new, changed and deleted documents in each source and update only the affected chunks. Store a content hash and an updated-at timestamp per document so unchanged files are skipped.
Webhooks from source systems give near real-time updates; scheduled jobs (hourly or nightly) are fine for slower content. Add freshness metadata to search so newer versions of a policy outrank older ones, and archive superseded documents instead of letting them compete. Monitor sync jobs like any other production pipeline, because a silent sync failure makes the assistant confidently wrong about last month's changes.
What are the most common RAG failure modes?
The most common failures are retrieval failures that look like model failures. When an answer is wrong, first check whether the right passage was retrieved at all.
| Symptom | Likely cause | Fix |
|---|---|---|
| Answer is vague or says "I don't know" when the answer exists | Bad parsing or chunking, missing keyword search | Structure-aware chunking, hybrid search, chunk context headers |
| Answer mixes old and new rules | Stale or duplicate documents | Incremental sync, versioning, freshness boost |
| Confident answer not in the sources | Weak grounding instructions, no citation check | Answer-only-from-context rule, citation validation, faithfulness evaluation |
| User sees content they should not | Permissions applied in the prompt, not in search | ACL filters at query time, tenant isolation tests |
| Good demo, poor production | Test questions written by the team, not users | Golden set from real logs, weekly review of failures |
| Costs grow with usage | Too many or too long passages per call | Reranking, fewer passages, prompt caching |
RAG vs fine-tuning vs long context: which should you choose?
Choose RAG for knowledge, fine-tuning for behavior and long context for small, per-request material. They are complementary, not competing, and many production systems use two of them together.
| Approach | Best for | Updates | Cost profile | Limits |
|---|---|---|---|---|
| RAG | Large, changing, permissioned knowledge; answers with citations | Instant, re-index a document | Indexing plus retrieval and prompt tokens per request | Quality bound by parsing and retrieval |
| Fine-tuning (e.g. LoRA) | Consistent style, tone, format or narrow task behavior | Requires retraining | Training runs plus hosting the tuned model | Poor at storing facts that change; no citations |
| Long context | One contract, a codebase slice or a report per request | Nothing to update | High token cost per request unless cached | Slower, expensive at scale, attention can drift in very long inputs |
The AI Grief Companion platform is a real example of combining approaches. It builds a personality profile, a fact graph in Neo4j and RAG preparation from the user's chat history, and trains LoRA adapters (SFT and CPT on Qwen2.5-7B with 4-bit QLoRA) for the voice of one specific person, with vLLM loading the right adapter at request time. The chat picks up whichever stage is ready, so a conversation is possible before training finishes. Fine-tuning handles how the person writes; retrieval handles what they said.
How we work on RAG projects
At Lytvynov Production we start RAG projects with a short scoping call: we look at a sample of your real documents and the questions users ask, then give you a fixed quote for the build. The first milestone builds a golden set and tests parsing and retrieval on it. We then deliver in milestones with evaluation scores reported at each one. If you are planning a knowledge assistant or AI chatbot over company data, get in touch.