RAG Document Analyzer

An intelligent document analysis tool using Retrieval-Augmented Generation. Ingests PDFs, chunks them into vector embeddings, and allows users to query complex information with cited sources.
The Problem
Standard RAG tutorials work fine for simple text snippets but fall apart in real-world scenarios. I needed a system that could handle large, complex PDFs without crashing the server, losing markdown table structures, or hitting strict API rate limits during the embedding and retrieval process.
Approach
Built a dual-engine FastAPI ingestion pipeline: a 'Fast Mode' for standard text extraction and a 'Deep Scan' layout-aware mode using PyMuPDF4LLM to perfectly preserve tables and hierarchical formatting.
Implemented a global chunk buffer that batches vector chunks in memory, reducing network calls to Hugging Face and Supabase from 30+ per document to just a few, completely bypassing strict free-tier API rate limits.
Used Supabase and pgvector for hybrid search, ensuring only high-confidence context reaches Groq's high-speed gpt-oss models.
Added a segmented control UI in React to allow dynamic search scoping, letting users seamlessly switch between querying a single active session or their entire global knowledge base.
Integrated a lightweight health-check endpoint to keep the free-tier Render backend awake via external pings, eliminating 60-second cold start delays for end users.
Key Features
- PDF ingestion and semantic chunking.
- Vector embedding storage in a local database.
- Context-aware query answering with source citations.
- Streamlit/FastAPI interface for interactive querying.
Challenges
Network bottlenecks and vector database query limits were the biggest hurdles. Initially, the backend made a network request for every single PDF page, causing massive slowdowns and rate-limit errors; I solved this by engineering a 20-chunk global memory buffer to batch uploads invisibly. Later, files stopped appearing in the knowledge base UI because the Supabase fetch query was hitting a hidden 1,000-row limit on vector chunks. I had to restructure the database schema to include a dedicated summaries table and rewrite the endpoint with a 15,000-row limit and exact JSONB metadata filtering to ensure reliable document retrieval.
Outcome
A highly optimized, production-ready RAG workspace that handles complex document layouts without memory spikes. The React UI features real-time token streaming, explicit markdown formatting for bold text and tables, and mandatory API key modal flows, resulting in a smooth, professional user experience that stays robust within free-tier infrastructure constraints.