• Home
  • About
  • Contact
Healthcare
  • Home
  • Jobs Search
  • Healthcare

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Elizabeth, NJ

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Casper, WY

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Flagstaff, AZ

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Arcata, CA

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Albany, NY

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Seattle, WA

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Richmond, VA

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Anchorage, AK

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Philadelphia, PA

  • no content

  • Bespoke Labs

Long-Horizon Coding Task Expert

Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Shape environments end to end: stateful, resumable systems with snapshotting, checkpointing, and branching rollouts, at multi-node scale where needed Generate and refine ideas for tasks and agentic trajectories across multi-turn, tool-using, and computer-use agent loops with persistent state across turns Design reward structure for sparse-reward settings: milestone and process rewards, subgoal and task decomposition, and credit assignment across long trajectories Validate that the work is correct, hard, and covers the right edge cases, including rubric design for partial credit, contamination avoidance, and defenses against reward hacking Design curricula that ramp task difficulty rather than shipping fixed-difficulty tasks Build automated pipelines that curate high-quality long-horizon RL environments and tasks at scale Write reliable, well-tested Python infrastructure rather than one-off research scripts Work directly with our research team with high ownership, building novel systems rather than maintaining legacy ones

  • Mount Pleasant, SC

  • no content

  • Bespoke Labs

Long-Horizon Coding Task Expert

Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Shape environments end to end: stateful, resumable systems with snapshotting, checkpointing, and branching rollouts, at multi-node scale where needed Generate and refine ideas for tasks and agentic trajectories across multi-turn, tool-using, and computer-use agent loops with persistent state across turns Design reward structure for sparse-reward settings: milestone and process rewards, subgoal and task decomposition, and credit assignment across long trajectories Validate that the work is correct, hard, and covers the right edge cases, including rubric design for partial credit, contamination avoidance, and defenses against reward hacking Design curricula that ramp task difficulty rather than shipping fixed-difficulty tasks Build automated pipelines that curate high-quality long-horizon RL environments and tasks at scale Write reliable, well-tested Python infrastructure rather than one-off research scripts Work directly with our research team with high ownership, building novel systems rather than maintaining legacy ones

  • Athens, GA

  • no content

  • Bespoke Labs

Long-Horizon Coding Task Expert

Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Shape environments end to end: stateful, resumable systems with snapshotting, checkpointing, and branching rollouts, at multi-node scale where needed Generate and refine ideas for tasks and agentic trajectories across multi-turn, tool-using, and computer-use agent loops with persistent state across turns Design reward structure for sparse-reward settings: milestone and process rewards, subgoal and task decomposition, and credit assignment across long trajectories Validate that the work is correct, hard, and covers the right edge cases, including rubric design for partial credit, contamination avoidance, and defenses against reward hacking Design curricula that ramp task difficulty rather than shipping fixed-difficulty tasks Build automated pipelines that curate high-quality long-horizon RL environments and tasks at scale Write reliable, well-tested Python infrastructure rather than one-off research scripts Work directly with our research team with high ownership, building novel systems rather than maintaining legacy ones

  • Bangor, ME

  • no content

  • Bespoke Labs
  • «
  • 1
  • …
  • 4356
  • 4357
  • 4358
  • 4359
  • 4360
  • 4361
  • …
  • 4368
  • »
Newsletter

Keep up with our always upcoming jobs and updates. Enter your e-mail and subscribe to our newsletter.

Cities
  • New York
  • San Diego
  • Los Angeles
  • Boston
  • Washington
  • Chicago
Categories
  • Healthcare
  • Automobile Jobs
  • Food Services
  • Construction
  • Logistics
  • Finance

Our network gives you instant access to young talent from over 120 countries and territories from all around the world.

© careerexact.com . Privacy Policy . Terms of Use