Back to jobs

AI Image-to-Video Workflow Pipeline

Search - AI Chatbot · local_filter_skipped · UID ~022082385701424423508

Open Job

Job Details

Budget $18.00 - $45.00/hr
ExperienceIntermediate
DurationUnknown
Weekly hoursLess than 30 hrs/week
Client countryAbout the client
Proposals10 to 15
Interviewing0
Invites sent0
First seenWed, Jul 29, 2026 11:47 AM
Last seenWed, Jul 29, 2026 9:14 PM

Description

Summary 1. Project Overview We are seeking a Senior Full-Stack / AI Integration Developer to build an automated AI image-to-video pipeline. The solution allows non-technical users to upload photo(s) of a subject and enter their name via a web form. The system automatically handles prompt templating, generates character-consistent static keyframes, and outputs animated video clips (e.g., folding arms, nodding head). The final output will be a standardized MP4 video file that automatically feeds into our existing Motion After Effects (AE) Web App Workflow for final compositing. 2. High-Level Architecture & Workflow [Web App Form Input] │ ── Upload Photos + Enter Subject Name + Select Motion Template ▼ [Backend Automation Engine] │ ── Injects inputs into pre-configured Templated Prompts ▼ [AI Stage 1: Keyframe & Identity Consistency] ── Google Gemini API ("Nano Banana") │ ── Ingests photo + character prompt ➔ Generates identity-consistent base frame ▼ [AI Stage 2: Video Motion Generation] ── Google Veo 3 API │ ── Image-to-Video motion generation ➔ Renders motion (nodding, arm-folding, etc.) ▼ [Post-Processing & Output] │ ── Normalizes resolution, framerate, aspect ratio, and MP4 encoding ▼ [Handoff to Existing Motion AE Web App Workflow] 3. Detailed Technical Requirements Module A: Web Form & Prompt Engine Frontend Web Form: File upload component for target person’s photos (JPG/PNG). Input fields for Subject Name and selection of predefined motion actions (e.g., Nodding, Folding Arms, Gesturing). Status indicator / job tracker showing generation progress. Dynamic Prompt Engine: System-level templating engine that automatically merges user inputs (Subject Name, uploaded photo references) with pre-designed system prompts to maintain consistent character framing and motion instructions. Module B: Static Keyframe & Identity Preservation (Nano Banana / Gemini API) Goal: Create clean, character-locked static keyframes from user photos. Requirements: Integration with Google Gemini Image API ("Nano Banana" / Gemini Flash Image). Utilize multi-image conditioning / character preservation capabilities to ensure lighting, attire, and facial features match the subject. Generate high-quality static keyframe(s) ready for image-to-video conversion. Module C: Video Generation & Motion Animation (Google Veo 3) Goal: Animate static keyframes into natural motion clips. Requirements: Integration with Google Veo 3 API (via Google AI Studio). Image-to-Video generation using Stage 1 keyframes + motion action prompts. Asynchronous processing with polling/webhooks to handle long-running video generation tasks. Module D: Formatting & AE Workflow Integration Goal: Deliver ready-to-composite video files to the existing AE template engine. Requirements: Automated media standardization matching canvas dimensions. Webhook trigger to send the rendered MP4 file URL and metadata directly to our AE template workflow API. Robust queue management (e.g., BullMQ, Celery, or Cloud Tasks) for handling concurrent user submissions and retries. 4. Required Tech Stack & Developer Profile AI & Generative Models: Deep experience with Google AI Suite (Vertex AI, Gemini Image / Nano Banana API, Google Veo 3 Video API). Backend: Node.js, Python (FastAPI/Django), or similar framework. Task Management: Cloud Tasks, Celery, or Redis-based task queues for handling asynchronous rendering pipelines. Cloud Infrastructure: Google Cloud Platform (GCP Cloud Run, Cloud Functions, Google Cloud Storage). Media & Automation: FFmpeg script automation, RESTful Web APIs, Webhooks. 5. Deliverables Backend Integration Service: Fully functional queue-based pipeline handling prompt assembly, Gemini keyframe generation, and Veo 3 animation. Web Frontend / Form Integration: Clean UI for form submission and status updates. AE Integration Bridge: Validated payload handoff delivering encoded MP4 assets to our existing AE template API. Documentation: Setup instructions, environment variables configuration, and API documentation for editing prompt templates. 6. Submission Guidelines for Applicants Please provide: Examples of previous projects involving Google Cloud AI (Gemini, Vertex AI, or Veo). Prior experience building asynchronous AI media generation workflows (image-to-video / character preservation pipelines). Recommended backend stack and estimated project timeline. Estimated costs for build.

Skills

AI Development Web Application Development Adobe After Effects

Notification History

ChannelTypeStatusSentError
No notifications.

User Actions

ActionActed at
No actions.