Back to jobs

LLM Engineer to cut our AI costs (FastAPI + LangChain + RAG)

Search - AI Chatbot · local_filter_skipped · UID ~022079200201450235403

Open Job

Job Details

Budget Unknown
ExperienceExpert
DurationUnknown
Weekly hoursUnknown
Client countryUnknown
ProposalsUnknown
InterviewingUnknown
Invites sentUnknown
First seenMon, Jul 20, 2026 1:53 PM
Last seenMon, Jul 20, 2026 7:10 PM

Description

We run an AI chat assistant with some actions on a FastAPI backend. It uses LangChain, a RAG pipeline over a Milvus vector DB, and Claude models. It works, but our LLM cost per request is too high, mostly because some of our prompts are very large (one is around 12,000 input tokens per call). We need someone to bring that cost down without changing how the system actually behaves. That last part is the whole point — we don't want "shorter prompts that give worse answers." We want proof the answers stay the same. What we need done: 1. Go through our 4 main prompts and reduce token size where it's genuinely safe. We've already done an analysis that found real duplication (repeated examples, dead formatting instructions, etc.), so there's a starting point. 2. Build an automated test suite BEFORE making changes. Run our real queries through the current system, record what it does, then re-run after changes to prove decisions didn't change. English and Japanese both matter to us. 3. Evaluate 2–3 cheaper/alternative models against that same test suite, so we can see if switching models saves money while keeping quality. Deliver everything as reviewable pull requests with before/after numbers — token counts, cost, and test results. Our stack: Python, FastAPI, LangChain, Milvus, AWS Bedrock (Claude), some Node.js in front. A few things we care about: - You've worked on real production RAG/LLM systems, not just demos - You're comfortable using AI coding tools like Claude Code or Curso

Skills

Artificial Intelligence Machine Learning Python Deep Learning

Notification History

ChannelTypeStatusSentError
No notifications.

User Actions

ActionActed at
No actions.