Back to jobs

RAG Expert AI ML

Search - AI Chatbot · local_filter_skipped · UID ~022081045698426255533

Open Job

Job Details

Budget $10.00 - $12.00/hr
ExperienceExpert
DurationUnknown
Weekly hoursLess than 30 hrs/week
Client countryAbout the client
Proposals15 to 20
Interviewing0
Invites sent0
First seenSat, Jul 25, 2026 6:37 PM
Last seenSun, Jul 26, 2026 12:55 AM

Description

Summary Project Overview We are building a domain-specific RAG (Retrieval-Augmented Generation) chatbot designed to answer complex, high-accuracy queries related to GST rates and legal compliance. Our knowledge base is derived from official CBIC PDFs containing narrative legal text and highly structured rate tables. The primary challenge is maintaining "ground truth" across frequently updated legal documents and ensuring perfect data extraction. Key Responsibilities Advanced Table Extraction: Develop robust, automated pipelines to extract and normalize complex rate tables from multi-page PDFs, ensuring structural integrity where headers and layouts vary. Retrieval Optimization: Refine hybrid search strategies (dense/sparse) to bridge the gap between user terminology ("chappals") and legal taxonomy ("footwear of sale value..."). Version Control for LLMs: Architect a solution for handling "updated info" (amendments, corrigenda, and superseding notifications) to ensure the model always pulls the current version of the truth, rather than superseded data. Deterministic Retrieval: Implement precision-first retrieval mechanisms to support exact-match requirements (e.g., specific HSN codes or tax rates), reducing reliance on approximate vector similarity where absolute accuracy is mandatory. Technical Requirements Deep Experience with RAG: Proven track record in building RAG systems for domain-specific, high-accuracy environments. PDF Parsing Proficiency: Strong experience with tools like Docling, Unstructured, or custom parsing logic to handle complex, multi-year PDF layouts. Semantic Search & Reranking: Expertise in hybrid search (BGE-M3, Qdrant) and implementing reranking pipelines to improve retrieval precision. Stack: Backend: FastAPI. Search/Database: Qdrant (Hybrid search), PostgreSQL. Models: Experience working with local SLMs (e.g., Qwen 3.5) for both reasoning and ingestion-time enrichment. Data Engineering: Understanding of document lifecycle management (linking amendments to original sources).

Skills

Artificial Intelligence Python

Notification History

ChannelTypeStatusSentError
No notifications.

User Actions

ActionActed at
No actions.