Tailored answers, filled into supported job forms.Your tailored resume and answers, filled into supported job forms for you to review.

Download the Chrome Extension
AI AdoptionFunded CompaniesJob SimulationCertificationsRoadmapsJobsPricing
Sign In
OneRoadmap

OneRoadmap is a career platform built around ORI, its AI career agent. ORI finds overlooked job opportunities, matches them to your profile and shows the skill gaps to close, with roadmaps, challenges, job simulations and certifications to close them. When you are ready, it prepares a tailored resume, application answers and an application strategy, with a cover letter where the application asks for one. The OneRoadmap Chrome extension fills supported application forms for you to review and submit, and you keep track of every application in one place.

gaurav.ghai@oneroadmap.in
Delhi NCR, India

Platform

  • AI Roadmaps
  • Free Certifications
  • Learning Resources
  • Pricing

Training

  • AI Adoption Workshops
  • Expert Sessions
  • Upcoming Events
  • Workshop Gallery

Company

  • About
  • Blog
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds
  • Delete your data

© 2026 OneRoadmap

Operated by Ghai Technologies, India · International operations through One Roadmap Marketing Management, Dubai, UAE

Built for your next chapter.

Open roles

AI / ML

Staff Engineer, Machine Learning Systems & Reliability - Moveworks

Moveworks

Senior · 5+ yrsOn-site · Mountain view, California, United statesFull-timeListed 4d ago
Apply now

Backed by

Lightspeed India

HQ

🇮🇳 India

Open roles

80

Experience Senior · 6+ yrs (5+ years)

About the role

structured by ORI

Staff Engineer, Machine Learning Systems & Reliability - Moveworks Other Mountain View, CALIFORNIA, United States Full-time Apply for job Company DescriptionMoveworks: the Agentic AI Assistant platform that empowers the entire workforce. Our platform enables employees to converse with all of their business systems…

What you will do

  • Design and build the production path for the complete ML lifecycle: data and feature preparation, training, experiment tracking, evaluation, artifact and model management, serving, monitoring, feedback collection, and retraining.
  • Build continuous-delivery workflows for models, prompts, agent workflows, data dependencies, and supporting services. Establish automated quality, safety, performance, and compatibility checks.
  • Implement safe rollout patterns such as shadow traffic, canaries, progressive delivery, feature flags, versioned artifacts, automated rollback, and operational kill switches.
  • Turn self-learning approaches into controlled production feedback loops. Build systems for collecting outcomes, validating feedback, maintaining lineage, triggering model refreshes, comparing candidates, and promoting changes under explicit guardrails.
  • Define and operate SLIs, SLOs, alerts, and error budgets across infrastructure, data pipelines, inference services, model quality, and product behavior.

What they are looking for

  • A track record of Staff-level technical ownership, typically gained through 7+ years of experience in software engineering, platform engineering, SRE, production engineering, or ML infrastructure.
  • Strong software-engineering skills in Python and at least one production systems language such as Go, Java, C++, or Rust.
  • Experience designing, operating, and troubleshooting distributed production systems, including failure analysis, capacity planning, and performance optimization.
  • Hands-on experience with cloud infrastructure, containers and Kubernetes, infrastructure as code, CI/CD, and modern observability.
  • Practical understanding of the ML lifecycle—including training, evaluation, model deployment, serving, monitoring, versioning, and retraining—and the ability to collaborate effectively with applied ML engineers or researchers.
  • Experience distinguishing service-health problems from data-quality or model-quality problems.
  • Familiarity with SRE practices such as SLIs/SLOs, error budgets, sustainable on-call, incident management, and blameless postmortems.
  • A strong automation and internal-customer mindset: you build platforms that are reliable, understandable, and pleasant for other engineers to use.
  • Excellent technical judgment and communication skills, especially when navigating ambiguity and coordinating across teams during production incidents.
Software EngineeringML InfrastructurePlatform EngineeringSite Reliability EngineeringCI/CDInfrastructure as CodeCapacity PlanningPerformance OptimizationPythonGoJavaC++RustKubernetes
Full posting text

Staff Engineer, Machine Learning Systems & Reliability - Moveworks Other Mountain View, CALIFORNIA, United States Full-time Apply for job Company DescriptionMoveworks: the Agentic AI Assistant platform that empowers the entire workforce. Our platform enables employees to converse with all of their business systems through natural language to quickly find answers and automate tasks. Powered by the world's most advanced LLMs, our proprietary models, and a sophisticated Agentic AI platform, we're transforming how work gets done by allowing AI to take initiative, streamline complex workflows, and continuously learn and adapt.Moveworks is trusted by over 5.5 million employees at more than 350 of the world’s largest companies, including 10% of the Fortune 500, to automate everyday tasks and streamline business operations. Recognized on the Forbes Cloud 100 and AI 50 lists, Moveworks was also named one of Fast Company’s 2025 Most Innovative Companies and Inc’s Best in Business, in the Best in Innovation category. Moveworks was also recognized at Microsoft’s 2025 Partner of the Year and in 2024, received the AI Breakthrough Award. In December 2025, Moveworks was acquired by ServiceNow, marking a pivotal milestone in our journey to create a single front door to work for all business systems. By combining ServiceNow’s leading workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we deliver the AI platform for every person and every workflow. Built to go beyond basic summaries to deliver meaningful business impact. Together, our AI acts across enterprise systems to turn conversations into completed work.By joining our team, you’ll be at the forefront of the AI transformation, backed by the global scale of ServiceNow and the agility of a high-growth company. We are looking for world-class talent to help us extend agentic AI to every employee across every corner of the business. Servicenow It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.Join us to put AI to work for people. Job DescriptionWe are building AI-enabled product capabilities that improve through data, feedback, and real-world use. We need the production systems that make those capabilities dependable: repeatable delivery, measurable quality, controlled learning loops, and reliable operation at scale.We’re looking for a hands-on Staff Engineer who can move machine-learning models, agentic workflows, and self-learning approaches from promising prototypes into secure, observable, continuously deployable production systems.This role sits at the intersection of ML systems, platform engineering, and site reliability engineering. You will partner with ML, data, product, and infrastructure teams to create a paved path from experimentation to production—and take ownership of how those systems perform and evolve once deployed.What you’ll doDesign and build the production path for the complete ML lifecycle: data and feature preparation, training, experiment tracking, evaluation, artifact and model management, serving, monitoring, feedback collection, and retraining.Build continuous-delivery workflows for models, prompts, agent workflows, data dependencies, and supporting services. Establish automated quality, safety, performance, and compatibility checks.Implement safe rollout patterns such as shadow traffic, canaries, progressive delivery, feature flags, versioned artifacts, automated rollback, and operational kill switches.Turn self-learning approaches into controlled production feedback loops. Build systems for collecting outcomes, validating feedback, maintaining lineage, triggering model refreshes, comparing candidates, and promoting changes under explicit guardrails.Define and operate SLIs, SLOs, alerts, and error budgets across infrastructure, data pipelines, inference services, model quality, and product behavior.Connect model analytics and product telemetry with traditional operational signals so teams can understand whether a problem originates in infrastructure, data, model behavior, or the surrounding product.Improve the scalability, availability, latency, and cost efficiency of distributed training, inference, and data-processing workloads. Own capacity planning and resource optimization, including GPU resources where applicable.Participate in production ownership across the service lifecycle: architecture reviews, deployment, on-call, incident response, blameless postmortems, and systemic remediation.Build self-service platforms and automation that reduce operational toil and shorten the time required for ML engineers and data scientists to reach production.Apply LLMs or agentic automation to evaluation, troubleshooting, and operational workflows where they produce reliable, measurable improvements.Establish practical standards for cloud infrastructure, Kubernetes, infrastructure as code, observability, security, and compliance.Provide technical leadership across ML, data, product, and platform teams, mentoring engineers and influencing architecture without relying on formal authority.QualificationsTo be successful in this role you have:A track record of Staff-level technical ownership, typically gained through 7+ years of experience in software engineering, platform engineering, SRE, production engineering, or ML infrastructure.Strong software-engineering skills in Python and at least one production systems language such as Go, Java, C++, or Rust.Experience designing, operating, and troubleshooting distributed production systems, including failure analysis, capacity planning, and performance optimization.Hands-on experience with cloud infrastructure, containers and Kubernetes, infrastructure as code, CI/CD, and modern observability.Practical understanding of the ML lifecycle—including training, evaluation, model deployment, serving, monitoring, versioning, and retraining—and the ability to collaborate effectively with applied ML engineers or researchers.Experience distinguishing service-health problems from data-quality or model-quality problems.Familiarity with SRE practices such as SLIs/SLOs, error budgets, sustainable on-call, incident management, and blameless postmortems.A strong automation and internal-customer mindset: you build platforms that are reliable, understandable, and pleasant for other engineers to use.Excellent technical judgment and communication skills, especially when navigating ambiguity and coordinating across teams during production incidents.Additional InformationWork PersonasWe approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.Equal Opportunity EmployerServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. AccommodationsWe strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact globaltalentss@servicenow.com for assistance. Export Control RegulationsFor positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

Apply on company site

Meet Ori - your career agent on WhatsApp

Find jobs, get your roadmap, check if you're ready for a role and prepare applications - in chat, any language.

Ask Ori about this role
Checking your fit…
Moveworks

India

Backed by Lightspeed India

Company pageWebsite

More at Moveworks

MoveworksFunded

On-site · Hyderabad, Telangana, India

Senior · 5+ yrs · 1d ago

Senior Workday Integration Developer
Workday StudioEIBPrism+9 more
View role
MoveworksFunded

On-site · Mountain view, California, United states

Senior · 5+ yrs · 4d ago

Senior Security Software Engineer, IAM - Moveworks
IAMSecurityLeast privilege+10 more
View role
MoveworksFunded

On-site · Mountain view, California, United states

Senior · 5+ yrs · 4d ago

Senior Software Engineer, Agent Eval Platform
Distributed systems+12 more
View role
MoveworksFunded

On-site · Sandy springs, Georgia, United states

Senior · 5+ yrs · 4d ago

Senior Security Software Engineer, IAM - Moveworks
IAMAI coding toolsRole design+10 more
View role
Apply