Description
Summary We've built a voice agent that triages repair faults over the phone for UK social housing tenants. A tenant rings in, no app, the agent works out what's actually wrong with the boiler or the damp, and an engineer receives a diagnostic brief. It's live and taking calls. It now needs to be production-grade before it speaks to a real tenant, and that's the work. STACK Pipecat 0.0.108, self-hosted LiveKit, Speechmatics (EU1), Mistral (Paris), Sinch, Python 3.12, Hetzner + Coolify. Everything EU-hosted. No US-owned managed services anywhere in the pipeline. That's not a preference — it's a procurement requirement for our customers and it's most of why the product exists. WHAT NEEDS BUILDING — An eval harness for the conversation layer — A tree-execution engine: a state machine where the LLM does edge selection only, and never invents a node or a diagnosis — A deterministic safety gate that runs alongside the LLM and can veto it — Async photo evidence: a vision model extracts into a closed vocabulary and feeds the tree as answers, never as a diagnosis — Sinch telephony finished to production quality FIRST ENGAGEMENT Two weeks, paid, building the eval harness. Real deliverable, separable, and it's the thing everything else depends on. If it goes well: 3 days a week for six months, then a conversation about something permanent. MUST HAVE — Built a voice agent that took real phone calls. Not a demo. Pipecat, LiveKit Agents, Vapi, Retell or equivalent. — Strong Python async — Experience making an LLM produce reliable structured output under constraint — Have built an eval harness for an LLM system before, not just used one — Comfortable running self-hosted infrastructure USEFUL — SIP / telephony (Sinch, Twilio, Telnyx) — Vision models extracting structured facts from images — Multilingual ASR in production — Have worked somewhere that safety mattered NOT NEEDED Housing knowledge. Frontend. UK public sector experience. We handle all of that. TIMEZONE EU or UK hours. Remote.