Files
SRAG-Page/README.md
T
2026-08-15 15:20:12 +02:00

4.2 KiB

SRag — Efficient Retrieval-Augmented Generation

A modern, lightweight RAG implementation designed for efficient document parsing and contextual AI queries. Runs entirely on local models via Ollama — no cloud dependency.

Live: srag.jrodriiguezg.link
GitHub: jrodriiguezg/SRag


Overview

SRag is a clean, maintainable Retrieval-Augmented Generation (RAG) pipeline that connects document store embeddings to a language model. It is designed to be containerized, dependency-minimal, and completely local — no external API calls required.

The system ingests documents, generates vector embeddings, and answers queries by retrieving relevant context before passing it to an LLM, drastically reducing hallucinations.

How It Works

Documents (text, PDF, Markdown)
         │
         ▼
  ┌──────────────────┐
  │  Ingestion Engine │  ← Parses and chunks documents
  └──────────────────┘
         │
         ▼
  ┌──────────────────┐
  │   Vectorization   │  ← Ollama Embedding (nomic-embed-text)
  └──────────────────┘
         │
         ▼
  ┌──────────────────┐
  │    Vector Store   │  ← In-memory / persistent index
  └──────────────────┘
         │
  User Query ──────► Query Router ──► Retrieves top-k context
                                            │
                                            ▼
                                    LLM (Ollama / qwen2.5)
                                            │
                                            ▼
                                       Answer

Core Architecture

Component Role
Ingestion Engine Parses text, PDFs, and Markdown into manageable chunks
Vectorization Generates dense embeddings optimized for semantic search
Query Router Intercepts user prompts, retrieves context, augments LLM input
Vector DB Lightweight in-memory or persistent storage for fast retrieval

Stack

Layer Technology
Framework LlamaIndex
LLM (local) Ollamaqwen2.5:7b
Embeddings Ollama — nomic-embed-text
Vector Store LlamaIndex VectorStoreIndex (in-memory / persistent)
Containerization Docker Compose

Quick Start

import os
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader, Settings
from llama_index.llms.ollama import Ollama
from llama_index.embeddings.ollama import OllamaEmbedding

# Setup local models via Ollama
Settings.llm = Ollama(model="qwen2.5:7b", request_timeout=120.0)
Settings.embed_model = OllamaEmbedding(model_name="nomic-embed-text")

# Load documents and generate vector index
documents = SimpleDirectoryReader("docs").load_data()
index = VectorStoreIndex.from_documents(documents)

# Initialize strict context-based chat engine
chat_engine = index.as_chat_engine(
    chat_mode="context",
    similarity_top_k=20,
    verbose=False
)

Key Design Decisions

  • LLM Agnostic: Compatible with Ollama for local inference or any standard API backend (OpenAI-compatible).
  • Embeddings Flexibility: Supports local models via Hugging Face or external APIs like OpenAI.
  • Context-strict mode: chat_mode="context" ensures the LLM answers only from retrieved documents, minimizing hallucinations.
  • similarity_top_k=20: Retrieves a generous context window to improve answer coverage.

Prerequisites

  • Ollama installed and running
  • Models pulled: ollama pull qwen2.5:7b and ollama pull nomic-embed-text
  • Python 3.10+ with llama-index and related packages
  • (Optional) Docker for containerized deployment

Support

Support on Ko-fi

Author

Juan Raul Rodriguez Gil — Built with passion for Linux & AI
jrodriiguezg.link | Instagram


© 2026 Juan Raul Rodriguez Gil