Back to jobs

Full-Stack AI Developer: Internal Web App for Google Gemini 3.1 Flash TTS (Custom Voice Profiles)

Search - AI Chatbot · local_filter_skipped · UID ~022072700349300828595

Open Job

Job Details

Budget Unknown
ExperienceExpert
DurationUnknown
Weekly hoursUnknown
Client countryAbout the client
ProposalsLess than 5
Interviewing0
Invites sent0
First seenThu, Jul 2, 2026 3:21 PM
Last seenFri, Jul 3, 2026 1:11 AM

Description

Summary We are looking for an experienced full-stack developer to build a lightweight, highly stable internal web application for our content editing team (2–5 concurrent users). The application will interface directly with Google’s **Gemini 3.1 Flash TTS API** (`gemini-3.1-flash-tts-preview`) to generate high-fidelity text-to-speech outputs, with a primary creative focus on creating and maintaining a highly consistent, repeatable **Lebanese Arabic accent**. While the default Google AI Studio platform has strong audio quality, it frequently cuts off long scripts mid-generation and suffers from severe "voice drift" (the voice randomly shifting tone, age, or performance style between consecutive generations). We need a tool that eliminates these platform bugs, allows us to create and reuse locked voice profiles, and maintains an archive of past generations. # Key Features & Scope of Work ### 1. Simple, Secure Team Access * A single, clean web page protected by a basic, shared group password (no complex user sign-up or registration flows required). * A secure backend server to safely mask and hold our Google Cloud / AI Studio API key so it is never exposed to the frontend browser. ### 2. Custom Voice Profile Builder & Library (Crucial Consistency Feature) We need a dashboard panel where editors can create, name, edit, and save custom "Voice Profiles" to a shared team library. * **The Inputs:** When creating a profile, the user must be able to select one of Gemini’s 30 base voices (e.g., *Charon, Kore, Puck, Zephyr*) and customize its permanent identity using specific configuration fields/sliders: * **Character Name:** (Crucial for grounding the AI's consistency). * **Age Description:** (e.g., Young Adult, Mid-30s, Mature/Elderly). * **Baseline Speed/Pacing:** (e.g., Slow, Balanced, Fast/Energetic). * **Default Emotional Persona:** (e.g., Warm, Corporate, Enthusiastic, Serious). * **Custom Director's Notes:** (Where we hardcode natural language instructions for the *Lebanese Arabic accent* and regional pacing). * **The Library:** Once saved, these profiles must be accessible via a dropdown menu on the main text-to-speech page. Selecting a profile must force the backend to send the exact same structural payload to the API every single time to ensure zero voice drift across separate script generations. * **Temperature Lock:** The backend must force a low temperature setting (0 or 0.1) on these profiles to eliminate unintended model randomness. ### 3. Instant Voice Previews * Next to the voice creation menu, include a play button that instantly plays a pre-recorded, 3-second static audio sample of each available Gemini base voice so editors can audit the baseline tones without triggering live API charges. ### 4. Auto-Chunking & Audio Stitching Engine (Fix for Text Cutoff) * To prevent Gemini’s native truncation bug on long text inputs, the backend must automatically split scripts into smaller text blocks/paragraphs, process the audio chunks asynchronously, and seamlessly stitch them into a single, cohesive high-fidelity output file (using a library like **FFmpeg**). ### 5. Block-Based Script UI (Line-by-Line Editing) * Build a block-style script interface where text is handled paragraph-by-paragraph. * Include separate, line-level "Regenerate" buttons so editors can re-render a single specific paragraph if the accent performance slips, without needing to re-generate the entire multi-minute project. ### 6. Generation History & Download Archive (New Feature) * **History Log:** Create a simple table or sidebar log tracking previous generations. * **Metadata Storage:** The log should display the timestamp, the user-defined script title or a short text preview, and the specific Voice Profile that was used. * **Archive Downloads:** Editors must be able to listen to previous generations directly from the log and instantly click a button to download the historical MP3 file without rerunning the script or spending API credits. ### 7. Script Formatting & Native Tag Support * The text input container must perfectly support Arabic characters and register Gemini 3.1's native inline audio tags (such as `[excited]`, `[whispers]`, `[laughs]`, `[short pause]`) so editors can make mid-sentence changes. * Provide a simple dropdown selector to change languages globally if needed (supporting Arabic, English, and French). ### 8. Media Export & Session Management * An instantaneous audio player to preview the generated script with a dedicated **Download MP3** button. * Basic backend file handling using unique, timestamped filenames to prevent concurrent editors from overwriting one another's audio data on the server. # Technical Requirements * **Backend:** Node.js (Express) or Python (FastAPI/Flask) * **Frontend:** Clean, modern, and highly responsive UI (React, Vue, or Tailwind CSS) * **APIs & Utilities:** Google GenAI SDK (Interactions API / `gemini-3.1-flash-tts-preview`), FFmpeg for audio processing, and a lightweight database (e.g., SQLite, PostgreSQL, or simple local storage) for voice profiles and history management. # How to Apply Please provide: 1. A brief overview of your experience working with AI Text-to-Speech models or Google GenAI APIs. 2. A rough timeline of how long this specific build will take you. 3. Your approach to saving custom persona profiles and logging audio history in the backend to ensure absolute voice consistency and efficiency. We are looking for someone who is ready to walk us step by step with the process, work fast and make sure that the final result meets our requirements. This is our general overview of how we want the app to work, but after implementation, slight edits might be required to properly fit our needs. Thank you

Skills

AI App Development API

Notification History

ChannelTypeStatusSentError
No notifications.

User Actions

ActionActed at
No actions.