AI / ML2026completed

PDF RAG Assistant

A retrieval-augmented generation system for asking questions over uploaded PDFs. Documents are chunked, deduplicated, and indexed in a persistent Chroma vector store. Queries go through optional LLM expansion, hybrid retrieval (semantic MMR + BM25), cross-encoder reranking, and a context-only chat model so answers stay grounded — with explicit “I don’t know” behavior when evidence is thin.

The backend is a FastAPI service for ingest and query APIs (streaming supported); the frontend is a Streamlit chat UI with source citations (file + page). OpenAI embeddings and chat models power the generation path; Hugging Face supplies the reranker. The codebase is modular (ingestion, retriever, reranker, LLM) with config and structured logging aimed at production-minded iteration.

PythonFastAPIRAGChromaDBOpenAIStreamlit
2018

Year

0

Features

0

Technologies

Problem
  • Balancing retrieval breadth vs precision across diverse PDFs
  • Keeping answers strictly grounded without sounding useless
  • First-load latency from embedding and reranker models
Solution
  • Hybrid fusion plus reranking before the generator sees context
  • Strict system prompts and explicit insufficient-context replies
  • Modular stages so ingest, retrieve, rerank, and generate can be tuned independently

Key Features

✓

PDF upload, chunking, deduplication, and persistent Chroma indexing

✓

LLM query expansion before retrieval

✓

Hybrid fusion of vector search and BM25 with document filters

✓

Hugging Face cross-encoder reranking for top-k context

✓

Streaming answers with file/page source citations

✓

Anti-hallucination prompts and insufficient-context fallbacks

✓

FastAPI REST backend + Streamlit chat frontend

Technologies

PythonFastAPIStreamlitChromaDBOpenAI Embeddings + ChatBM25Hugging Face (reranker)PydanticLangChain-style pipelines

Learnings

  • Designed a full RAG path from PDF ingest through streamed, cited responses

  • Combined vector (MMR) and BM25 retrieval to improve recall on mixed documents

  • Used cross-encoder reranking to tighten context before generation

  • Practiced API design, config management, and modular pipeline structure

Highlights

Hybrid Retrieval

Cross-Encoder Rerank

Cited Answers

Next Project

MicroBanking System

Team full-stack banking platform: PostgreSQL/PLpgSQL transaction core and Node API, with a React TypeScript client for accounts, transfers, and live balances.