Description
Summary We are looking for an experienced AI/RAG engineer to help us design and implement a robust search chunk storage and retrieval mechanism for an AI agent/chatbot system. We have already built and deployed AI agents using AWS Bedrock, Claude, Lambda, API Gateway, EventBridge, Step Functions, S3, and DocumentDB. The current system powers an AI Business Buddy WhatsApp chatbot that performs business intelligence using RAG-based responses. Now, we want to improve the way documents and business knowledge are processed, chunked, stored, embedded, searched, and retrieved. What we need: We need someone who can help us build or improve the following: Document ingestion pipeline Text extraction and cleanup Smart chunking strategy for documents, web content, and business data Metadata structure for chunks Embedding generation using AWS Bedrock or similar models Storage of chunks and embeddings Efficient semantic search and retrieval Filtering based on metadata such as tenant, source, document type, date, etc. Retrieval logic for chatbot/agent responses Re-ranking or relevance improvement if required Handling fresh document updates and re-ingestion Clean APIs for search and retrieval Scalable architecture that can support multiple business tenants Existing system context: The system currently includes: AWS Bedrock / Claude for AI responses Lambda-based REST APIs API Gateway for exposing endpoints EventBridge for scheduled ingestion tasks S3 for storing raw documents/data DocumentDB for business data and indexed content Secrets Manager for credential management IAM-based least privilege access WhatsApp chatbot webhook integration Slack/Jira integration for exception triaging Ideal candidate should have experience with: .RAG architecture .Vector search and embeddings .AWS Bedrock .Claude or other LLMs .Python .Lambda .API Gateway .S3 .DocumentDB, MongoDB, .OpenSearch, Pinecone, Weaviate, .Qdrant, or similar .Chunking strategies for AI retrieval .Metadata-based filtering .Multi-tenant AI applications .REST API integrations .Secure AWS architecture .Expected deliverables: .Recommended architecture for .chunk storage and retrieval .Chunking and metadata strategy .Embedding and vector storage approach .Search/retrieval API design .Implementation of ingestion and .retrieval flow .Documentation of the approach .Suggestions for improving retrieval accuracy and scalability Nice to have: .Experience with MCP-like servers or agent tool integrations .Experience with AWS Step .Functions and EventBridge .Experience building production-grade AI agents .Experience with WhatsApp chatbot systems .Experience with business intelligence or document-based AI assistants Goal: We want to make our chatbot responses more accurate, grounded, and context-aware by building a reliable retrieval layer that can fetch the most relevant chunks from business documents and external knowledge sources. Please apply with examples of similar RAG, vector search, or AI retrieval systems you have built. Also mention what vector database or retrieval architecture you would recommend for this use case and why.