• Home
  • About
  • Contact
Automobile
  • Home
  • Jobs Search
  • Automobile

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Tupelo, MS

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Minneapolis, MN

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Norwalk, CT

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Idaho Falls, ID

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Spokane, WA

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Nampa, ID

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Joliet, IL

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Lowell, MA

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Saint Petersburg, FL

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Meridian, ID

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Seaford, DE

  • no content

  • Bespoke Labs

Machine Learning Engineer

Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics Implement methods from recent ML papers quickly and turn them into production-grade systems

  • Austin, TX

  • no content

  • Bespoke Labs
  • «
  • 1
  • …
  • 523
  • 524
  • 525
  • 526
  • 527
  • 528
  • …
  • 4716
  • »
Newsletter

Keep up with our always upcoming jobs and updates. Enter your e-mail and subscribe to our newsletter.

Cities
  • New York
  • San Diego
  • Los Angeles
  • Boston
  • Washington
  • Chicago
Categories
  • Healthcare
  • Automobile Jobs
  • Food Services
  • Construction
  • Logistics
  • Finance

Our network gives you instant access to young talent from over 120 countries and territories from all around the world.

© careerexact.com . Privacy Policy . Terms of Use