Description
Summary Swiss IT services company (est. 1996, ~10 people), clients in legal and finance. We're building on-premise AI servers for our clients: open-source models, RAG over internal documents and email, full data sovereignty — nothing leaves the building. We handle servers, networks and security ourselves. We need someone who has already run this stack in production and can tell us what breaks, not someone building a POC with us. What you'd do: - Hardware/software sizing with precise, justified recommendations: GPU and VRAM, Ollama vs vLLM, Open WebUI configuration, vector DB, ingestion pipeline — not a generic opinion, but numbers and choices we can apply directly - Hardware sizing for on-prem installs at client sites: GPU/VRAM, CPU/RAM balance, storage, rack vs tower, power/cooling in small server rooms - RAG quality in French (non-English) - Access control / multi-user setup - Maintenance, model updates, monitoring over time - 1-2 video calls/week to start, screen-sharing on real deployments, plus async questions by email. Terms: hourly, ongoing, long-term if it works. To apply, answer these three questions — no cover letter, no deck: - One on-premise LLM/RAG deployment you took to production: team size, stack, hardware, what went wrong. Include something verifiable — reference, case study, repo, or other concrete proof. - For a 20-user deployment on 1-2 GPUs: which GPU/VRAM, CPU/RAM, and inference server would you use, and why? - Your availability for recurring calls (we're CET). Applications without a verifiable deployment example will not be considered.