Back to jobs

Developer Needed for Private AI Chat System (API-based)

Search - API Integration · local_filter_skipped · UID ~022087172070598778229

Open Job

Job Details

Budget $? - $?/hr
ExperienceExpert
DurationLess than 1 month
Weekly hoursHours to be determined
Client countryAbout the client
Proposals50+
Interviewing24
Invites sent0
First seenWed, Aug 12, 2026 11:44 AM
Last seenWed, Aug 12, 2026 3:14 PM

Description

Summary Developer Needed for Private AI Chat System (API-based) Project Overview I need a private, self-hosted AI assistant that connects in real-time to third-party AI APIs (Anthropic Claude and/or Google Gemini) — NOT a local/offline LLM (no Ollama, no local model weights). The system must be accessible via a web interface from multiple devices (laptop, phone, tablet) and multiple locations, with my data staying under my control. What I Need Built 1. Backend server that securely calls the Anthropic API and/or Google Gemini API - API keys stored server-side only (never exposed to the browser/client) - Support for switching between or combining providers 2. Simple web-based chat interface - Responsive design (usable on desktop and mobile browsers) - User login/authentication (so only I — and anyone I authorize — can access it) 3. Document storage & retrieval (RAG - Retrieval-Augmented Generation) - Ability to upload my own documents/files (PDF, Word, text, etc.) through the web interface - Documents automatically split into chunks and converted into embeddings (numerical representations for semantic search) - Embeddings stored in a self-hosted vector database on my own server (e.g., Qdrant, Weaviate, or pgvector — developer to recommend based on scale) - At query time, the system retrieves only the most relevant chunks (semantic similarity search) and sends those — not the full documents — to the AI API along with my question - Full documents and the vector database must remain on my own infrastructure at all times; only the retrieved text snippets are sent externally to Claude/Gemini per query - Ability to update/delete documents from the knowledge base (re-indexing when content changes) - Source attribution: responses should indicate which document(s) the answer was drawn from 4. Content ingestion pipeline (for large personal knowledge sources) - Ability to bulk-ingest a large personal library of trading education content (e.g., YouTube video transcripts) into the RAG knowledge base - Pipeline should: fetch/extract transcripts, clean up the text (remove filler, fix formatting), chunk by topic/concept rather than fixed length, generate embeddings, and index into the vector database - Should also support ingesting reference/technical documentation (e.g., the official Pine Script language reference) as a separate, always-available knowledge source used specifically to ground any code generation and reduce errors/hallucinations - Reusable pipeline: I should be able to add new sources (new videos, new documents) later without needing the developer each time - Video content must be processed in two parts: (1) extract the spoken transcript with timestamps, and (2) extract screenshots/frames from the video at intervals, (3) content ingestion pipeline, then use a vision-capable AI to describe what's shown in each frame (chart levels, marked zones, price action) - Merge the transcript and the visual descriptions together by timestamp, so that spoken references like "look here" are paired with what was actually shown on screen at that moment 5. Live web search capability - The assistant should be able to search the web in real time for current information (e.g., market news, recent Pine Script/TradingView updates) when a question needs up-to-date data - Implement via a web search tool/plugin connected to the AI API (e.g., Claude's built-in web search tool, or a search API such as Google/Bing/Perplexity Sonar integrated into the backend) - Should be clearly distinguished from the RAG knowledge base: RAG = my own static documents/strategies; web search = live, current information from the internet - Results should include source links so I can verify information 6. Hosting & deployment - Deployed on a VPS (I will provide access — see below) or recommend one with justification - Dockerized setup preferred, for portability and easy maintenance - Must remain accessible 24/7 via a secure URL (HTTPS) 7. Logging & basic security - Request logs (who asked what, when) - Basic protection against unauthorized access (rate limiting, authentication) 8. Documentation - Clear instructions on how to maintain, update, and restart the system - How to add/rotate API keys - How to add new users if needed Requirements for the Developer - Proven experience with backend development (Python/FastAPI or Node.js/Express) - Experience integrating LLM APIs (Anthropic, OpenAI, or Google Gemini API) - Experience with vector databases / RAG pipelines (e.g., Pinecone, Qdrant, Weaviate, or pgvector) - Experience deploying and managing applications on a VPS using Docker - Understanding of basic security practices (secrets management, authentication, HTTPS/SSL setup) - Able to communicate clearly in English and explain technical choices in plain language - Portfolio or examples of similar past projects (chatbots, AI integrations, RAG systems) - Experience with text extraction/ingestion pipelines (e.g., transcript extraction, document parsing, bulk chunking strategies) is a plus - Experience integrating web search tools/APIs (e.g., Claude's web search tool, Google Search API, Bing API, or Perplexity Sonar) is a plus This will be a project based contract (Upwork fixed contract with milestones), where gradual payments will be released upon milestones achievements. It will be considered finalized when everything will be done/completed. This is not a "pay-by-the-hour" contract. The total budget & exact milestones for this project will be discussed and agreed upon.

Skills

API Integration API Python Machine Learning Trading Automation TradingView Python Script +3

Notification History

ChannelTypeStatusSentError
No notifications.

User Actions

ActionActed at
No actions.