Back to jobs

Braintrust AI Evals Setup

Search - AI Chatbot · local_filter_skipped · UID ~022078585667317512955

Open Job

Job Details

Budget Unknown
ExperienceIntermediate
DurationUnknown
Weekly hoursUnknown
Client countryUnknown
ProposalsUnknown
InterviewingUnknown
Invites sentUnknown
First seenSat, Jul 18, 2026 9:14 PM
Last seenSun, Jul 19, 2026 1:17 PM

Description

We’re looking for a Braintrust (or similar AI Evals) specialist to help us build and scale automated evaluations for Punter, our conversational AI sports betting product. Punter combines operator betting/odds feeds, sports data APIs, user profile data and web search to answer user queries and select the relevant betting markets with our operator partners. Our primary goal is deterministic ground-truth testing of betting market-selection accuracy. Given a user prompt, we need to objectively verify that Punter has selected the correct underlying betting market and selection. This needs to be extremely reliable - if Punter is unsure of the user's intent, it should clarify rather than select the wrong market. We also want to use LLM-as-a-judge for more subjective evaluation, such as response quality, relevance and reasoning. We currently test manually in Excel and want help setting up Braintrust to run thousands of automated test cases/traces, evaluate results, identify failure modes and track performance over time.

Skills

AI Agent Development

Notification History

ChannelTypeStatusSentError
No notifications.

User Actions

ActionActed at
No actions.