Description
Summary Looking for an engineer who can both build and rigorously validate AI infrastructure components. Roughly 60% hands-on technical work, 40% structured testing and written reporting. You will be writing small services, wiring up cloud and provider integrations, driving real traffic through them, and then proving with evidence what works and what does not. We need someone equally comfortable in a terminal and in a written report. LOCATION Strong preference for candidates based in Gurgaon or the wider Delhi NCR area. The work is remote day to day, but being in the same city makes occasional in-person working sessions possible and keeps coordination simple. Elsewhere in India will be considered for the right person. We are not looking for agencies or teams — this is for an individual freelancer. WHAT THE WORK INVOLVES - Writing small purpose-built services in Node.js or Python: MCP servers, webhook endpoints, mock upstreams, agent harnesses - Configuring LLM provider integrations and exercising them with real traffic: Gemini, OpenRouter, AWS Bedrock, and local/open-weight models via Ollama - Setting up and exercising guardrail and content-safety services: AWS Bedrock Guardrails, Google Model Armor, Azure Content Safety, Presidio for PII/PHI detection - Standing up identity providers and testing auth end to end: JWT, OAuth 2.1 with PKCE, OIDC, JWKS validation, token introspection, expiry and revocation - Working with Model Context Protocol: transports, tool schemas, tool discovery, and tool-level authorisation - Deploying and debugging on AWS EC2 with Docker Compose — reading container logs, rendered configs and REST APIs to find root causes rather than guessing - Writing test cases, capturing evidence, reproducing failures, classifying severity, and writing defect reports precise enough to survive being challenged - Node.js and Python — comfortable writing small services in either - Docker and Docker Compose, Linux, SSH - AWS: EC2, IAM, ideally Bedrock - Auth: JWT, OAuth 2.1 / PKCE, OIDC, JWKS. Hands-on with at least one IdP (Keycloak, Auth0, Okta, Entra) - HTTP at a low level: curl, headers, status codes, SSE and streaming responses, JSON-RPC - Git and GitHub, including issue and PR discipline - Model Context Protocol experience, or the ability to pick it up quickly STRONG PLUS - LLM APIs across more than one provider, and OpenAI-compatible / Anthropic-compatible endpoint shapes - Agent frameworks: PydanticAI, LangChain, LangGraph - Observability: OpenTelemetry / OTLP, Prometheus, Grafana, ClickHouse - Load testing: k6, wrk - Experience where correctness had real consequences — PII/PHI handling, payments, access control HOW WE WORK - Async first. Short written update at the end of each working day, responsive during agreed overlap hours - Hourly, time tracked - Conventions and structure already in place — you will be working inside an existing setup, not starting from scratch - Long-term engagement if it goes well WHAT WE WEIGH AS HEAVILY AS SKILLS - Clear written English - Evidence over assertion. "It worked" is not a result - Honesty about uncertainty — "I have not confirmed this yet" is worth more than a confident guess - Consistent daily communication, every working day TO APPLY Answer these briefly in your proposal: 1. Where are you based? 2. Describe a bug or misconfiguration you found where the consequence was serious rather than cosmetic. How did you find it, and how did you prove it? 3. Have you worked with MCP? If yes, what did you build. If no, how would you stand up an MCP server over HTTP and verify it works? 4. Which of these have you configured yourself, hands on: AWS Bedrock, Bedrock Guardrails, Google Model Armor, Azure Content Safety, Presidio, or an IdP? 5. Your hours of reliable overlap with IST. Proposals that do not answer these will not be reviewed.