AI-powered research assistant that automates paper discovery, PDF analysis, RAG-based question answering, and research-topic monitoring using FastAPI, ChromaDB, MongoDB, and LLMs.
Researchers spend significant time searching for relevant papers, reading lengthy PDFs to find specific information, and repeatedly checking for new research in their areas of interest. This makes research discovery and analysis time-consuming and inefficient. The challenge was to create a centralized system that could simplify paper discovery, enable quick question answering from research documents, and automate monitoring of new papers.
Built an AI Research Automation platform with React and FastAPI. Users can discover papers through arXiv, upload PDFs, and ask questions using a RAG pipeline. Uploaded documents are extracted with PyMuPDF, split into chunks, converted into embeddings, and stored in ChromaDB for semantic retrieval, while MongoDB stores users, paper metadata, topics, and chat history. Retrieved document context is passed to a Groq-hosted LLM to generate answers. An APScheduler-based automation periodically monitors arXiv for user-selected topics, identifies new papers, generates AI summaries, and stores them for later access.
Developed a functional AI-powered research assistant combining paper discovery, PDF analysis, RAG-based question answering, and automated research monitoring. Enabled users to retrieve relevant information from uploaded research papers without manually searching through entire documents. Automated periodic arXiv monitoring for user-selected research topics and AI-generated summaries of newly discovered papers. Implemented secure user authentication with JWT and bcrypt-based password hashing, along with chat-history storage for research conversations. Created a modular backend architecture integrating FastAPI, MongoDB, ChromaDB, Groq, arXiv, and APScheduler into a single research workflow.
Automated research paper monitoring with a scheduled 6-hour interval. The RAG pipeline retrieves the top 4 relevant document chunks for user queries, reducing the need for manual searching through entire uploaded research papers. The platform centralizes paper discovery, document analysis, and research tracking in one workflow.