• Home
  • About
  • Contact
Finance
  • Home
  • Jobs Search
  • Finance

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Dekalb, IL

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Madison, WI

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Broken Arrow, OK

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Peoria, IL

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Mesa, AZ

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Logan, UT

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Bakersfield, CA

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Merced, CA

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • The Woodlands, TX

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Milford, DE

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Harrisonburg, VA

  • no content

  • Bespoke Labs

RL Environment Engineer

Work directly with our research team on long-horizon RL environment and task creation for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts Build and shape RL environments in your domain, including the setup, state handling, and tooling an agent interacts with Generate and refine ideas for tasks and agentic trajectories, then validate that the work is correct, hard, and covers the right edge cases Design reward functions, milestones, and rubrics that give meaningful signal across long trajectories, including partial credit where pass/fail is too blunt Identify failure modes, reward hacking, and degenerate solutions before they reach training

  • Pawtucket, RI

  • no content

  • Bespoke Labs
  • «
  • 1
  • …
  • 1460
  • 1461
  • 1462
  • 1463
  • 1464
  • 1465
  • …
  • 4368
  • »
Newsletter

Keep up with our always upcoming jobs and updates. Enter your e-mail and subscribe to our newsletter.

Cities
  • New York
  • San Diego
  • Los Angeles
  • Boston
  • Washington
  • Chicago
Categories
  • Healthcare
  • Automobile Jobs
  • Food Services
  • Construction
  • Logistics
  • Finance

Our network gives you instant access to young talent from over 120 countries and territories from all around the world.

© careerexact.com . Privacy Policy . Terms of Use