Frameworks & tools
You can build RAG from scratch with an embedding API, a vector DB client, and an LLM call — and for a simple pipeline, that's often the clearest choice. Frameworks add ready-made components, integrations, and patterns that save time on complex systems, at the cost of a layer of abstraction to learn and debug.
The main options
| Framework | Strengths | Good when |
|---|---|---|
| LangChain / LangGraph | Huge integration ecosystem; LangGraph for stateful, cyclic agent flows; LangSmith for tracing/eval | Complex, multi-step or agentic pipelines with many moving parts |
| LlamaIndex | Purpose-built for RAG: loaders, indices, query engines, advanced retrieval patterns out of the box | Data-centric RAG over many sources; you want retrieval patterns ready-made |
| Haystack | Production-oriented, modular pipeline design, strong on search | Search-heavy, production RAG with a clear pipeline structure |
When to go framework-free
For a straightforward retrieve-and-generate pipeline, direct API calls give you full control, fewer dependencies, easier debugging, and no abstraction tax. Many teams prototype with a framework to move fast, then drop to raw calls for the hot path once requirements are clear. There's no prize for using — or avoiding — a framework; pick by how much complexity you actually have.
Supporting tools you'll want
- Vector DBs: Pinecone, Weaviate, Qdrant, Milvus, Chroma, pgvector.
- Embeddings: OpenAI, Cohere, Voyage, Gemini; open models via sentence-transformers, BGE, E5.
- Rerankers: Cohere Rerank, Voyage, bge-reranker.
- Evaluation: RAGAS, TruLens, DeepEval, Phoenix.
- Observability: LangSmith, Phoenix, and general tracing tools.
- Parsing: layout-aware document parsers and OCR for messy PDFs.
This landscape changes fast — verify current options and capabilities when you build.