• Home
  • About
  • Contact
Finance
  • Home
  • Jobs Search
  • Finance

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Montpelier, VT

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Evansville, IN

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Barre, VT

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Sterling Heights, MI

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Brookings, SD

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Minot, ND

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Billings, MT

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Wichita, KS

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Scottsdale, AZ

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Sacramento, CA

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Rochester Hills, MI

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Raleigh, NC

  • no content

  • Bespoke Labs
  • «
  • 1
  • …
  • 433
  • 434
  • 435
  • 436
  • 437
  • 438
  • …
  • 4413
  • »
Newsletter

Keep up with our always upcoming jobs and updates. Enter your e-mail and subscribe to our newsletter.

Cities
  • New York
  • San Diego
  • Los Angeles
  • Boston
  • Washington
  • Chicago
Categories
  • Healthcare
  • Automobile Jobs
  • Food Services
  • Construction
  • Logistics
  • Finance

Our network gives you instant access to young talent from over 120 countries and territories from all around the world.

© careerexact.com . Privacy Policy . Terms of Use