Description
Summary This Is a Paid Trial, Not a Full-Time Role This is a 5-hour paid trial. Build an autonomous AI agent using the Claude Agent SDK that scores 55–63% on the GAIA benchmark (up from a 37–47% baseline). You'll be paid for 5 hours of work. Top performers move on to ongoing, long-term projects with our team. Think of it as a working audition — you get paid, we get signal. The 5 hours cover your hands-on build time. Benchmark runs are unattended and the API tokens are on us (we provide the key) — don't count those against your hours. Use the full window through the deadline for background runs. What You're Building An agent that can research, write and execute code, and solve complex multi-step tasks Runs in E2B sandboxes for secure code execution Tested against the full 165-task GAIA benchmark The current system scores 37–47%. You're building from scratch with the Claude Agent SDK to hit 55–63% — at a sensible cost (see "What Success Looks Like"). Tech Stack (Required) Claude Agent SDK (@anthropic-ai/claude-agent-sdk) — non-negotiable, must build on this framework Model: claude-opus-4-8 — pinned, so every submission is comparable Python 3.14+ E2B sandboxes for code execution Claude Code as your primary development tool Not acceptable: custom orchestration frameworks, LangGraph, or debugging existing code. This is a clean rebuild on the Claude Agent SDK. What We Provide An Anthropic API key for the trial — you don't pay for tokens. Usage runs through our account, which is also how we measure your cost-per-task (below). What Success Looks Like We're hiring for the best score per dollar, not the highest score at any cost. Your submission is evaluated on three things: Benchmark score — accuracy across the full 165-task GAIA set. Cost efficiency — total tokens and cost-per-task (read from the provided key's usage). Between two agents at the same score, the cheaper one wins. Brute-forcing accuracy with runaway token spend is a fail, not a pass. Effective use of Claude Code — how you leverage AI-assisted development. Your score must be reproducible: a one-command run with pinned dependencies, and reported across 3 runs (mean + spread) — not a single cherry-picked number. What You Must Deliver All required for your submission to be considered: Working agent code — clean, documented, runs against the benchmark with one command Git history — full commit history showing your development process Exported Claude Code conversation history — all conversations used Benchmark results — score + cost-per-task, across 3 runs No code + git history + Claude Code logs = no evaluation. We assess both the output and the process. Required Experience Shipped AI agents or systems with measurable performance metrics (link real, verifiable repos) Experience with agent benchmarking (GAIA, etc.) and eval methodology Proficient with Claude Code for development workflows Can deliver under a tight deadline How to Apply Send: Links to agent projects you've built — GitHub with real commit history, or live demos Technical plan (2–3 paragraphs) — how you'd architect this to hit 55%+ efficiently Claude Code experience — how you use it in your workflow No generic proposals. Show relevant, verifiable experience or don't apply. Finalists do a short live screen-share to walk through their submission.