Back to jobs

AI/LLM Engineer: Evaluation, Prompt & Schema Optimization for Production Audit Platform

Search - AI Chatbot · ai_analyzed · UID ~022076703533369455031

Open Job

Job Details

Budget $? - $?/hr
ExperienceExpert
DurationUnknown
Weekly hoursLess than 30 hrs/week
Client countryAbout the client
Proposals10 to 15
Interviewing0
Invites sent0
First seenMon, Jul 13, 2026 4:29 PM
Last seenMon, Jul 13, 2026 11:54 PM

Description

Summary We operate a production AI/LLM auditing platform used by a large enterprise client. This platform transcribes recordings of in-store visits completed by field representatives, then runs an LLM-driven audit schema against those transcripts to score whether specific criteria were met. The platform is live and delivering value. We're now investing in a dedicated AI specialist to push accuracy higher, expand the feature set, and bring deeper expertise to how we evaluate and iterate on our LLM pipeline. What you'll do - Build out our evaluation loop. We currently test our audit schema weekly. We want to move to more frequent, more granular testing, the full schema or individual audit questions, on demand. - Improve scoring accuracy. Our benchmark is 90%+ accuracy on most audit questions. You'll drive us there, identify the questions where 90% isn't realistic and explain why, and push a defined subset of high-priority questions well beyond 90%. - Assess the approach itself. Determine whether we need new audit schemas and questions, or whether the current approach should be refactored for better accuracy and capability. We want your judgment here, not just execution. - Ship new capabilities. -- Keyword search over transcript text, with Boolean operators, phrase search, and grouping. Returns matching snippets plus a count of transcripts containing the term. Transcripts are ASR output and imperfect, so the design needs to account for that. -- Sentiment analysis: merchant sentiment and agent sentiment across a transcript (negative / neutral / positive), plus detection of sentiment shift over the course of a conversation (e.g., a merchant who started negative and ended positive). - Evaluate emerging tooling and make recommendations we can act on. - Potentially build a schema management UI so our non-developer team can modify audit questions without engineering involvement. - Communicate frequently. We want regular written progress recaps, not silence between milestones. What we're looking for - Strong hands-on experience building and shipping LLM-powered systems in production, not just prototypes - Real depth in LLM evaluation: building eval sets, measuring accuracy against ground truth, regression testing prompts and schemas, and reasoning about where models fail and why - Prompt engineering and structured-output experience (classification, scoring, extraction) - Experience working with ASR/transcript data and its failure modes is a strong plus - Comfort with search implementation (Boolean query parsing, full-text search) is a plus - Sentiment analysis experience beyond off-the-shelf APIs - Ability to form an opinion, defend it, and communicate it clearly in writing to both technical and business stakeholders - Self-directed — you'll be trusted to identify the right problems, not just close tickets Nice to have - Experience with RAG, fine-tuning, or model selection tradeoffs at scale - Frontend capability sufficient to build an internal admin/config UI - Background in compliance, QA, or audit-adjacent domains To apply Please include: - A brief description of an LLM system you took from prototype to production, and specifically how you measured and improved its accuracy. - Your approach to building an evaluation harness for a scoring/classification task where ground truth is human-labeled. - Your availability and hourly rate. Applications that appear to be generic or AI-generated without substance will not be considered. Please reference the word "transcript" in your first sentence so we know you read this.

Skills

Azure OpenAI Service AI Model Integration Integration Testing Python JavaScript

Notification History

ChannelTypeStatusSentError
telegram pre_ai_job_alert sent Mon, Jul 13, 2026 4:32 PM -

User Actions

ActionActed at
No actions.