Back to jobs

AI Voice Agent Platform Developer — Build the Application Layer on Our Voice AI Infrastructure

Search - AI Chatbot · local_filter_skipped · UID ~022082876805925551667

Open Job

Job Details

Budget $? - $?/hr
ExperienceExpert
DurationUnknown
Weekly hoursLess than 30 hrs/week
Client countryAbout the client
Proposals20 to 50
Interviewing2
Invites sent12
First seenThu, Jul 30, 2026 8:28 PM
Last seenFri, Jul 31, 2026 12:41 AM

Description

Summary Project Overview We provide a multilingual Voice AI infrastructure with enterprise APIs for automatic speech recognition, text-to-speech, speech translation, voice cloning and real-time streaming. Our speech infrastructure already exists. We're looking for an experienced AI Voice Agent developer to build the application layer that sits on top of our APIs. This is not a machine learning or speech model development project. Your responsibility is to build a scalable Voice AI application that consumes our APIs. Important You must use our APIs for all speech capabilities. Do not build or substitute speech recognition, text-to-speech, translation or voice cloning, and do not replace our APIs with ElevenLabs, OpenAI Speech, Azure Speech, Google Speech, Retell AI or Vapi. The objective is to make our infrastructure the underlying speech engine for every voice application you build. What You'll Build This is a single project covering the full scope below, not phased or billed by calendar week. We'll agree milestone checkpoints and payments with the freelancer we hire, based on their proposed delivery plan. A reusable, configurable Voice AI platform supporting multiple enterprise use cases, including AI customer support agents, AI receptionists, banking and telecom voice assistants, healthcare voice agents, internal enterprise assistants, appointment booking, lead qualification, call routing and FAQ agents. The platform should be configurable so new voice agents can be created without rebuilding the application. Voice Agent Framework Multi-turn conversations with conversation memory and session management, interruption handling (barge-in), voice activity detection and end-of-turn detection, low-latency streaming, and configurable prompts and agent personalities. AI Agent Capabilities Function calling and tool execution, retrieval-augmented generation, external API integration, and MCP-based tool integration so agents can consume external MCP servers. This includes business workflow automation and multi-agent orchestration where appropriate. Enterprise Integrations Modular, configurable connectors for CRM platforms, calendar systems, email providers, REST APIs, webhooks and internal business systems. MCP Integrations We expose our voice infrastructure through an MCP (Model Context Protocol) server. You'll extend and productionise this so AI assistants and AI-powered IDEs can use our speech capabilities directly, targeting Claude (claude.ai connectors, Claude Desktop, Claude Code), Cursor, Windsurf, ChatGPT and the OpenAI ecosystem, GitHub Copilot and other MCP-compatible clients. This means building a remote MCP server over Streamable HTTP with OAuth 2.1 authentication mapped to our API credentials, MCP tools covering ASR, TTS, translation, voice cloning, voice selection and account usage, one-click install links and configuration files for each target client, per-client setup guides, and usage metering, rate limiting and audit logging consistent with the main API, with MCP usage reported in the analytics dashboard. As with the rest of this project, all MCP tools must route through our APIs. Telephony Twilio Programmable Voice, LiveKit and SIP, with inbound and outbound calls, DTMF and WebRTC support, and telephony-grade (8 kHz) audio. The telephony layer should stay independent of the conversation engine. Multilingual Voice Automatic language detection, dynamic language switching and code-switching, with translation and voice cloning routed through our APIs. Administration Portal The ability to create and configure voice agents, prompts, voices and languages, manage API credentials, integrations, users and permissions, and view audit logs and configure rate limits. Analytics Dashboard Reporting on active and completed calls, call duration and response latency, speech recognition confidence where available, transcripts and call outcomes, language usage, API consumption, MCP tool usage, error rates, and searchable conversation history. Technical Requirements Preferred technologies are Python, FastAPI, TypeScript, React or Next.js, WebSockets, MCP and Docker. The architecture should clearly separate the transport layer (browser, Twilio, LiveKit, SIP), speech APIs, conversation engine, business logic, integrations, MCP server and analytics, so the same voice agent can run across web, mobile and telephony channels without changes to the core application logic. Deliverables Complete source code, technical documentation, a deployment guide and Docker configuration, API integration documentation, MCP client setup guides, an installation guide and architecture diagrams. In Your Proposal, Please Include Examples of Voice AI applications you've built Experience integrating custom ASR and TTS APIs, not only managed platforms Any MCP servers or MCP client integrations you've built Your proposed architecture, delivery plan and milestones Estimated timeline and project cost Any recommendations or technical considerations based on your experience Successful delivery may lead to a long-term engagement building enterprise Voice AI products on top of our platform.

Skills

AI Agent Development Artificial Intelligence Natural Language Processing Python Machine Learning

Notification History

ChannelTypeStatusSentError
No notifications.

User Actions

ActionActed at
No actions.