Description
Summary We're building an autonomous AI "influencer" / explainer-video pipeline on top of our OpenClaw agent orchestration stack. The goal is to have an agent that, from a single high-level instruction, can generate a consistent AI character and a long-form talking explainer video for a specific website or process we define. The workflow should be similar to reference tutorials we'll share: building AI influencers using Google Flow + Nano Banana (Imagen) + Veo 3.1 Lite / Google Vids, and long-form AI presenters using Google Flow Agent and Gemini video APIs. What you'll build: - An OpenClaw orchestrator agent plus worker agents that break down a request like "Create a 5-minute explainer for this URL" into steps: script, character, video clips, audio, and final assembly. - A pipeline that generates or reuses a consistent AI character (reference sheet, outfit consistency, background) using Nano Banana / Imagen via Google Gemini or Google Flow. - Integration of Veo 3.x / Gemini video to create short talking clips with lip sync and consistent scene, stitched into a long-form video (2-10 minutes). - A TTS / voice-cloning integration (e.g., ElevenLabs) to keep voice consistent across episodes, with reusable voice profiles. - A content module that takes a target website or process flow and generates a structured script (hook, explanation, examples, CTA). - Logging of every step (JSON + artifacts) with retry/rerun support via OpenClaw Mission Control. Preferred tech & APIs: - OpenClaw Gateway / Mission Control or similar multi-agent orchestration experience - Google Gemini API (Nano Banana/Imagen for images, Veo 3.x for video), Google Flow, Google Vids - ElevenLabs or similar voice-cloning/TTS API - Python or Node.js backend, REST/webhook integrations, queues, storage What we're looking for: - Proven projects involving autonomous agents or multi-step AI pipelines (ideally OpenClaw or similar) - Examples of AI-generated talking-head or explainer videos you've built Initial scope (2 weeks, fixed price $300): - Deliver one end-to-end pipeline that takes a URL + text brief and outputs a 60-120 second talking-head explainer video with consistent AI character and voice. - Provide basic configuration notes to extend to longer 5-10 minute explainers in a future phase. Please include in your proposal: - Any OpenClaw (or similar agent orchestration) projects you've delivered, with links/screenshots if possible - Links to AI-generated videos you've built (talking heads, explainers, AI influencers, faceless channels, etc.) - A short outline of how you'd architect this pipeline (which APIs for images, video, voice, and orchestration/monitoring) If this works well, we're interested in ongoing collaboration to expand our AI content automation.