Back to jobs

AI Engineer

Search - AI Chatbot · local_filter_skipped · UID ~022087400556541275055

Open Job

Job Details

Budget Unknown
ExperienceIntermediate
DurationUnknown
Weekly hoursUnknown
Client countryAbout the client
ProposalsLess than 5
Interviewing0
Invites sent0
First seenWed, Aug 12, 2026 12:07 PM
Last seenWed, Aug 12, 2026 6:31 PM

Description

Summary I’m looking for an AI Engineer to help build an AI Safety Evaluation & Governance product powered by open-source models. This is a 1-month, hands-on project with an expected commitment of around 20 hours per week. The goal is to build an MVP that can automatically test AI models, identify safety failures, analyze failure patterns, and support continuous improvement. 🔍 What you’ll work on • Build an automated red-teaming engine that generates test cases across risk domains, severity levels, and attack strategies • Run tests against models such as Gemma, Llama, Qwen, and API-based models • Develop evaluators for jailbreak success, policy violations, over-refusal, under-refusal, and severity • Structure safety policies into consistent taxonomies and evaluation criteria • Turn confirmed failures into reusable eval datasets and regression tests • Build lightweight reporting for model comparison, human review, and policy-version tracking 🧠 What I’m looking for • Experience with open-source LLMs, inference pipelines, prompt optimization, fine-tuning, LoRA/QLoRA, and LLM evaluation • Ability to independently build an end-to-end MVP, including data pipelines, model orchestration, scoring, and reporting • Familiarity with AI safety, red teaming, jailbreaks, content moderation, or Trust & Safety systems • Bonus: experience with model-based evaluators, human-in-the-loop review, agentic testing, or multimodal safety ⏳ Project setup Duration: 1 month Time commitment: Around 20 hours per week Format: Flexible and remote-friendly Stage: Early-stage, 0-to-1 MVP This is not about manually writing red-team prompts one by one. The goal is to build a scalable system that can continuously generate tests, evaluate model behavior, identify safety gaps, and verify whether issues have been resolved. If this sounds like you, please DM me with a brief introduction and examples of relevant work.

Skills

LLM Prompt Engineering Testing Red Team Assessment Open Source

Notification History

ChannelTypeStatusSentError
No notifications.

User Actions

ActionActed at
No actions.