Skip to content
Back to Articles
Machine Learning

How to Build RAG with Python and Vector Database

Learn step by step how to build a Retrieval-Augmented Generation (RAG) system using Python and vector databases in 2026, complete with architecture, code, and best practices.

September 30, 2026
How to Build RAG with Python and Vector Database

The enterprise generative AI solutions market is projected to surpass USD 220 billion by the end of 2026, and more than 65% of corporate AI implementations now rely on Retrieval-Augmented Generation (RAG) architecture to meet the needs for accuracy, contextuality, and internal data security. If at the beginning of this decade RAG was still considered a laboratory experiment, by 2026 it has become the standard foundation for internal company chatbots, legal research assistants, technical recommendation systems, and cross-document semantic search engines. This surge is driven by the growing demand for AI that is not only fluent in language but also capable of grounding itself in verified and auditable data sources. RAG is an architectural pattern that combines the power of large language models (LLMs) with a semantic search engine based on vector databases, so that every generated answer is always grounded in real documents and the latest context.

What is RAG? Combining the LLM Brain with a Smart Library

Imagine you have a highly intelligent personal assistant, capable of writing, summarizing, and explaining anything. However, that assistant was born without memory of your company's internal documents, the latest financial reports, or the 2026 edition of product manuals. Whenever asked something specific, it can only fabricate answers that sound convincing but are potentially completely wrong. RAG solves this problem by giving that assistant a digital library card: before answering, it must search for relevant books (documents) in the library (vector database), read a few of the most important pages, then compose an answer based on what it actually read, complete with source footnotes. So, the LLM remains the brain that assembles language, but it no longer works from an empty memory but from real evidence found in your knowledge base.

Technically, RAG consists of three main components that work sequentially: indexing, retrieval, and generation. Indexing converts raw documents into numerical vector representations stored in a vector database. Retrieval searches for vectors most similar to the user's query using distance metrics such as cosine similarity. Generation hands the query plus the retrieval results to the LLM to produce a final answer that is contextual and minimizes hallucination. In practice, modern RAG systems in 2026 also often add layers of reranking, metadata filtering, and retrieval quality evaluation before the answer is sent to the user.

Common types of RAG implemented in industry today include:

  • Naive RAG: the simplest flow, going directly from query to retrieval to LLM, suitable for prototypes and small to medium document bases.

  • Advanced RAG: adds optimizations at the indexing stage such as smarter chunking, domain-specific embeddings, as well as query rewriting and query expansion before retrieval.

  • Modular RAG: the most flexible architecture, allowing independent component replacement, for example using a separate reranker, hybrid search that combines vector and keyword search, or a memory module for multi-turn conversations.

  • Graph RAG: a cutting-edge 2026 approach that combines vector search with knowledge graphs to answer questions requiring relational reasoning, such as "which clients have filed type X insurance claims in region Y in the last quarter?".

  • Agentic RAG: RAG integrated with autonomous AI agents, where the agent can decide when to search, when to ask for clarification, and when to explore external sources independently.

Why RAG Matters: The Foundation of Trustworthy Enterprise AI

1. Reducing Hallucination and Improving Factual Accuracy

Large language models, even the latest ones in 2026, still have a tendency to produce information that sounds convincing but is wrong when faced with data they have never seen during training. RAG fundamentally changes this dynamic by limiting the LLM's answer space to the provided context. When the system retrieves three relevant paragraphs from your internal HR policy, the LLM no longer needs to guess the contents of that policy; it simply summarizes and rearranges what is in front of it. As a result, the hallucination rate on specific documents can drop dramatically, and for many enterprise use cases, factual accuracy becomes more important than mere language fluency.

Case Study – Financial Services Company: A leading digital bank in Southeast Asia implemented RAG for its customer service chatbot in early 2026 and recorded a 41% reduction in escalations to human agents within the first six months, primarily because chatbot answers now always refer to the latest product terms and FAQs, rather than generic answers that were often wrong. The system also logs every source document used for every answer, making compliance audits easy.

2. Delivering Real-Time Knowledge Without Expensive Retraining

Retraining a foundation model to understand the latest company policies, this quarter's product prices, or newly issued government regulations is no small task: it requires large amounts of data, expensive GPU infrastructure, and weeks of time. RAG offers an elegant shortcut: simply add new documents to the vector database, and within seconds to minutes, your AI system can already answer questions based on that latest information. This enables companies to respond to market changes, product updates, or financial report releases at a speed impossible to achieve with traditional fine-tuning approaches.

By 2026, the knowledge update cycle in many companies has shifted from monthly to daily, even real-time for news and social media data sources. RAG architecture with vector databases that support streaming ingestion such as Qdrant, Weaviate, Pinecone, or Milvus allows new data to be searchable just seconds after entering the pipeline. This is crucial for industries such as e-commerce, media, and financial services where the value of information declines over time.

3. Supporting Compliance, Security, and Data Sovereignty

Companies in banking, healthcare, government, and legal sectors cannot simply send all their internal documents to public LLM APIs. Regulations such as GDPR, HIPAA, and various national data protection laws demand strict control over where data is stored, who can access it, and how it is used. RAG enables a clear separation between the knowledge base (which can be stored in on-premise or private cloud infrastructure) and the language model (which may run in a separate environment). Thus, sensitive documents never leave the company's security perimeter, while the LLM only receives selected text snippets temporarily.

Case Study – Healthcare Institution: A private hospital network in Indonesia built an internal medical question-answering system based on RAG with a vector database hosted in a private cloud. The system allows medical staff to search internal medical literature, patient treatment protocols, and anonymized medical records without sending raw data to external LLM providers. Initial results show that clinical information search time dropped from an average of 12 minutes to less than 2 minutes per query.

4. Increasing User Trust with Citations and Transparency

One of the biggest weaknesses of generative AI in the eyes of corporate users is its inability to explain where answers come from. RAG inherently solves this problem: because answers are generated from documents actually retrieved from the vector database, the system can easily display a list of sources, page numbers, or direct links to original documents for every answer. This citation feature is a major differentiator between chatbots that merely "sound right" and AI systems that are truly auditable and trusted for critical decision-making.

In the 2026 enterprise market, this auditability feature is no longer just an added value but a mandatory requirement in many tenders and requirement documents. Compliance teams can inspect retrieval logs, see which documents were retrieved for each question, and verify that the answers provided are indeed based on legitimate sources, not LLM fabrications.

RAG and Vector Database Adoption in Indonesia

Key Players: The global vector database ecosystem widely used in Indonesia in 2026 includes Qdrant (open-source, popular for on-premise deployment), Pinecone (cloud-based managed service), Weaviate (with strong hybrid search support), Milvus (for large scale and billions of vectors), and pgvector as a PostgreSQL extension that allows teams to start RAG without adding new infrastructure. On the embedding side, multilingual models such as sentence-transformers, embedding models from OpenAI, Cohere, as well as open-source models from BAAI (BGE-M3) and Jina AI are widely used for Indonesian-language documents. Meanwhile, local technology companies such as Kata.ai, Bahasa.ai, and Prosa.ai are beginning to offer ready-to-use RAG solutions optimized for the nuances of the Indonesian language.

Local Success Stories:

  • A large e-commerce company in Indonesia reported a 35% increase in product search relevance after replacing keyword-based search with vector-based semantic search integrated with an LLM to answer customer questions directly.

  • A legal technology startup in Jakarta built a RAG-based contract analysis platform capable of reviewing and flagging risky clauses in Indonesian and English contracts, cutting review time from hours to less than 10 minutes per contract.

  • A higher education institution implemented a RAG-based academic assistant that answers student questions about curriculum, schedules, and campus procedures with over 90% accuracy in internal trials, while also providing citations to official campus documents.

  • A state-owned enterprise in the energy sector uses RAG for an internal technical document search system containing thousands of inspection reports, enabling engineers to find repair recommendations from old reports in seconds.

Challenges & How to Overcome Them

1. Poor Chunking Quality

Chunking is the process of breaking long documents into smaller pieces before embedding and storing them in a vector database. If chunks are too large, retrieval results become less precise because too much mixed information is included; if too small, important context can be cut off and answers become partial. This challenge is even more complex for Indonesian-language documents that often have long paragraph structures, mixed languages, or table formats.

The way to overcome this is to implement structure-aware chunking: use separation based on headings, paragraphs, or even semantics with the help of segmentation models. Many teams in 2026 are shifting to semantic chunking, where chunk boundaries are determined by changes in meaning, not just character count. Experiment with chunk sizes between 500 and 1,500 tokens for narrative documents, and add 10-20% overlap between chunks to prevent context from being cut off at chunk boundaries. For tables and structured data, consider storing separate metadata or using table-specific embedding models.

2. Irrelevant or Too Generic Retrieval

A classic problem in RAG is the system retrieving documents that are semantically similar but do not directly answer the user's question. For example, the question "how much does international shipping cost?" might return chunks about international shipping policies in general without mentioning the actual cost figure. This happens because embedding based on semantic similarity does not always capture the user's specific intent.

The solution involves multiple optimization layers: first, perform query rewriting using an LLM to expand or clarify the question before retrieval; second, apply hybrid search that combines vector search with BM25-based keyword search to capture specific terms that might be missed by embeddings; third, add a cross-encoder reranker model after initial retrieval to filter and reorder results based on finer relevance. Periodic evaluation with metrics such as recall@k and nDCG is also important for monitoring retrieval quality as the document base grows.

3. Operational Costs and Scalability

Vector databases can grow quickly, and storage and computation costs for embedding millions of documents can balloon. Additionally, vector search at the scale of billions requires the right infrastructure to keep latency low, especially for real-time applications like chatbots that must respond within seconds.

The way to overcome this is to choose a vector database that supports quantization (such as product quantization or scalar quantization) to reduce memory usage without sacrificing too much accuracy. Apply strict metadata filtering before vector search to narrow the search space. Use tiered storage for old documents that are rarely accessed, and consider multi-tenant architecture to separate data between departments. For teams just starting out, pgvector with PostgreSQL can be an efficient choice because it leverages existing database infrastructure, while for large-scale needs, Milvus or Qdrant with Kubernetes deployment offer good horizontal scalability.

4. Inconsistent Answer Quality Evaluation

Without a good evaluation system, development teams find it difficult to know whether changes to chunking, embedding models, or prompt templates actually improve answer quality or actually degrade it. Manual evaluation does not scale for systems with thousands of queries per day, while LLM-based automatic evaluation still has limitations.

The way to overcome this is to build a continuous evaluation pipeline using a combination of retrieval metrics (recall@k, precision@k, MRR) and generation metrics (faithfulness, answer relevance, context relevance). Use an evaluation dataset representative of real user queries, and run automatic evaluation every time there is a change to the RAG components. Frameworks such as RAGAS and TruLens are widely adopted in 2026 for this purpose, and many teams also build internal dashboards to monitor evaluation scores over time. A/B testing on a subset of users also helps ensure that changes made truly have a positive impact.

The Future of RAG

  • Autonomous AI agents based on RAG will become increasingly common, where agents not only answer questions but also proactively search for information, compile reports, and take actions based on their knowledge base.

  • Graph RAG will see wider adoption due to its ability to answer questions requiring multi-hop and relational reasoning, especially in the financial, legal, and business intelligence sectors.

  • Multimodal RAG will develop rapidly, enabling systems to retrieve and reference not only text but also images, charts, tables, and even video, making the enterprise knowledge base truly comprehensive.

  • RAG will become increasingly standardized with the emergence of open protocols and APIs for interoperability between vector databases, embedding models, and LLMs, making it easier for companies to swap components without rebuilding the entire system from scratch.

Conclusion: RAG as a Strategic Imperative in the 2026 AI Era

Building a RAG system with Python and vector databases in 2026 is no longer a complex experimental project, but a core competency for engineering teams that want to deliver AI that is accurate, up-to-date, and trustworthy. With the right architecture, a vector database choice that fits the needs, and attention to chunking, retrieval, and evaluation quality, companies can transform piles of internal documents into intelligent assistants that truly deliver business value. Competitive advantage is no longer determined by who has the largest model, but by who most effectively manages and leverages the knowledge they already have. RAG is the bridge between a company's rich data assets and the increasingly mature generative capabilities of LLMs, and mastery of this architectural pattern will be the differentiator between companies that merely follow the AI trend and those that truly reap its benefits.

References

Tags

RAG
Python
Vector Database
LLM
Enterprise AI
Share this article