Any job page, tailored and filled in with one click.A tailored resume, answers and cover letter for any job page, filled in with one click.

Download the Chrome Extension
AI AdoptionFunded CompaniesJob SimulationCertificationsRoadmapsJobsPricing
Sign In
OneRoadmap

Empowering the next generation with AI education. Custom training for colleges and enterprises.

gaurav.ghai@oneroadmap.in
Delhi NCR, India

Platform

  • AI Roadmaps
  • Free Certifications
  • Learning Resources
  • Pricing

Training

  • AI Adoption Workshops
  • Expert Sessions
  • Upcoming Events
  • Workshop Gallery

Company

  • Blog
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds
  • Delete your data

© 2026 OneRoadmap

Operated by Ghai Technologies, India · International operations through One Roadmap Marketing Management, Dubai, UAE

Built for your next chapter.

Open roles

SOFTWARE · Reinforcement-Learning-Engineering

Reinforcement Learning Engineer

Bright Vision Technologies

Senior · 6+ yrsRemote · USFull Time$100000–150000Listed 4h ago
Apply now
Experience Senior · 6+ yrs (6+ years)

About the role

from listing

Reinforcement Learning Engineer - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. Job Title: Reinforcement Learning Engineer Location: 100% Remote (U.S.) Position Type: Full-time, Direct W2 Salary Range: $100,000–$150,000 Annually Experience Required: 6+ years Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position. Key Responsibilities Design and implement reinforcement learning solutions for sequential decision-making problems in real and simulated environments. Develop, calibrate, and maintain simulation environments suitable for large-scale agent training. Implement and evaluate modern RL algorithms including policy gradient, actor-critic, off-policy, and offline RL methods. Engineer reward functions and shaping strategies that align agent behavior with desired outcomes and safety constraints. Apply offline RL and imitation learning techniques where exploration is costly or unsafe. Use RLHF, DPO, and related techniques for fine-tuning large language models when relevant. Build scalable training infrastructure for distributed RL, including efficient experience collection and replay systems. Optimize training stability and sample efficiency through algorithmic and engineering improvements. Design rigorous evaluation protocols, including out-of-distribution and adversarial test cases. Implement safety mechanisms such as constraint enforcement, conservative policies, and human-in-the-loop oversight. Collaborate with applied scientists and product teams to identify high-value RL use cases. Monitor deployed policies and models in production for drift, regression, and unintended behaviors, building the alerting and dashboards that surface issues before they meaningfully affect users. Document methodology, design decisions, and operational characteristics for internal stakeholders. Stay current with RL research and translate promising techniques into production-ready solutions. Required Qualifications Master’s or PhD in Computer Science, Machine Learning, or a related field; or equivalent applied experience. Six or more years of combined RL research and engineering experience. Strong proficiency in Python and modern deep learning frameworks. Hands-on experience with at least one major RL library or in-house RL stack. Solid understanding of probability, optimization, and the theoretical foundations of RL. Experience designing and tuning reward functions in non-trivial environments. Familiarity with simulation environments and large-scale experience collection. Experience training neural network policies on GPU clusters. Strong written and verbal communication skills. Track record of shipping or publishing impactful RL work. Preferred Qualifications Experience with RLHF for large language models. Familiarity with multi-agent RL or hierarchical RL. Exposure to robotics, control systems, or autonomous driving. Publications in RL or related research venues. Open-source contributions to RL libraries or environments. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at . Bright Vision Technologies is an Equal Opportunity Employer. Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment. Originally posted on Himalayas

Reinforcement-Learning-EngineeringMachine-Learning-EngineeringAI-ResearchDeep-LearningApplied-AIReinforcement-Learning-EngineerMachine-Learning-Engineer-Reinforcement-LearningReinforcement-Learning-SpecialistReinforcement-Learning-Research-ScientistReinforcement-Learning-Scientist
Apply on company site

Meet Ori - your career agent on WhatsApp

Find jobs, get your roadmap, check if you're ready for a role and prepare applications - in chat, any language.

Ask Ori about this role
Checking your fit…

Opportunity details

Deadline
Closing in 60d · 21 Nov

As stated by the source. Anything not shown was not stated.

More at Bright Vision Technologies

Staff Product Manager

Freshworks

Senior · 10+ yrsFunded

On-site · Bengaluru, in

Full-time · 4h ago
View role
Staff Product Manager

Freshworks

Senior · 7+ yrsFunded

On-site · Bengaluru, in

Full-time · 4h ago
View role
Engineering Manager, Data Feeds

Chainlink Labs · Data-Engineering-Manager

Senior · 5+ yrs

Remote · US

SOFTWARE

Full Time · $129000–304000 · 5h ago
View role
L
Senior CMS (WordPress/Shopify) Developer

Lil Horse Lab · CMS-Development

Senior · 5+ yrs

Remote · LATAM

SOFTWARE

Contractor · $1000–2500 · 5h ago
View role
Apply