Description
We're looking for an AI engineer to build an internal knowledge base chatbot over about 200 of our PDFs: client deliverables, market research, SOPs, case studies. The team asks a question, the bot answers from our documents only, and cites the source PDF and page number. No hallucination from general knowledge. If the documents don't cover it, it says so. Required experience: Production RAG (retrieval-augmented generation), not prototypes Python and FastAPI LangChain for the retrieval pipeline (LlamaIndex also relevant) Pinecone or a comparable vector database (pgvector, Weaviate, Qdrant) OpenAI API, GPT-4o or similar Document ingestion: PDF parsing, chunking strategy, embeddings, metadata for source attribution Semantic search and vector search across a multi-document corpus Supabase for auth and per-user chat history Next.js, React, TypeScript for a clean internal chat UI Prior document Q&A or knowledge base chatbot work Core pipeline within a week, then tuning once we see performance on our real documents. Internal use only. Likely phase 2 for Slack integration and role-based access per team.