AI Agent Trajectory Annotator and Reviewer
Type: Contract, hourly Location: Remote Hours: 20–30 per week Pay: $20–30/hour, based on experience and language coverage Start: Immediate ABOUT THE ROLE We evaluate how well advanced AI coding agents solve real engineering problems. An agent is given a real open source codebase inside a container and a hard task, then works on its own for 80 to 250 steps. A trajectory is the full record of that run — every command, result, and decision. You will do two jobs, and you should expect either on any given day. •Annotate — Read a trajectory nobody has looked at yet and judge it step by step. • Review — Take an existing annotation, written by our AI tooling or another person, and confirm, correct, or reject it. TASKS YOU'LL SEE • Feature build — Add a working feature to a live library without breaking anything that already worked. • Rebuild — Work out what a compiled tool does by running it, then rebuild it to match its output, exit codes, and file effects. • Bug hunt — Find and fix twenty undocumented bugs across a dozen files with no test suite, then record what caused them. Mostly Python and Go, with some Rust, C, and Ct. A trajectory runs about 100 steps. WHAT YOU JUDGE IN A TRAJECTORY • Was the command right for the state the environment was actually in? • Did the agent read the previous output correctly? • Was the step wrong, or only inefficient — these are scored differently. • Where did the run first go off course — usually earlier than where it visibly broke. • Did the agent notice its own mistake and recover, or keep building on a false assumption? • Did it game the grader instead of solving the task (e.g., weakening a test or hardcoding an expected value)? WHAT WE NEED FROM YOU • Experience — 2 years in software engineering, DevOps, or site reliability, with real debugging in real codebases. • Languages — Strong in Python or Go, and able to read a language you've never used. • Linux — Comfortable with logs, running processes, build failures, and containers. • Workflow — Everyday Git, diffs, pull requests, and issue tracking. • Debugging — Able to work with no test suite and no error message pointing at the cause. • Focus — Able to hold context across a long run, because step 74 can depend on step 12. • Writing — Clear English, since every judgement needs an explanation another engineer can check.
- Manchester, NH
- no content
-