Tailored answers, filled into supported job forms.Your tailored resume and answers, filled into supported job forms for you to review.

Download the Chrome Extension
AI AdoptionFunded CompaniesJob SimulationCertificationsRoadmapsJobsPricing
Sign In
OneRoadmap

OneRoadmap is a career platform built around ORI, its AI career agent. ORI finds overlooked job opportunities, matches them to your profile and shows the skill gaps to close, with roadmaps, challenges, job simulations and certifications to close them. When you are ready, it prepares a tailored resume, application answers and an application strategy, with a cover letter where the application asks for one. The OneRoadmap Chrome extension fills supported application forms for you to review and submit, and you keep track of every application in one place.

gaurav.ghai@oneroadmap.in
Delhi NCR, India

Platform

  • AI Roadmaps
  • Free Certifications
  • Learning Resources
  • Pricing

Training

  • AI Adoption Workshops
  • Expert Sessions
  • Upcoming Events
  • Workshop Gallery

Company

  • About
  • Blog
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds
  • Delete your data

© 2026 OneRoadmap

Operated by Ghai Technologies, India · International operations through One Roadmap Marketing Management, Dubai, UAE

Built for your next chapter.

Open roles

SOFTWARE · Data Center Operations

Director, Reliability Engineering (Nashville, TN on-site)

Oracle

Senior · 10+ yrsOn-site · Nashville, TN, United StatesListed 9d ago
Apply now

Backed by

VC portfolio

HQ

🇺🇸 Reading, United Kingdom

Open roles

2049

How to stand out for Director, Reliability Engineering (Nashville, TN on-site) at Oracle

Auto Match agent

Let ORI find you the best jobs.

Set up your Auto Match agent once - your target role, level and where you want to work. It searches every day, scores each opening against your profile and resume, and delivers the ones worth applying to, with a prepared application a click away.

Searches every day Scored against your profile Applications prepared for you
Sign in & set up Auto Match agent

Resume & career call

Get your resume reviewed for this role - 30-minute 1:1 call

Line-by-line resume feedback for this application, how to position your Role Readiness, and a clear plan for what to do next - with a OneRoadmap career coach.

Experience Senior · 6+ yrs (10+ years)

About the role

structured by ORI

You will own the vision, operating model, engineering standards, and portfolio programs that improve infrastructure availability, maintainability, resilience, and lifecycle performance at scale. >>This position will be full-time on-site at Oracle's offices located in Nashville, TN. Relocation assistance may be…

What you will do

  • Lead the Reliability Engineering organization supporting multiple regions, sites, and infrastructure programs across OCI's mission-critical data center portfolio.
  • Define the multi-year reliability engineering strategy, organizational roadmap, operating model, and investment priorities required to improve data center infrastructure resilience and support OCI's continued growth.
  • Establish and govern portfolio-wide reliability engineering standards and methodologies, including FMEA/FMECA, RCA, Reliability-Centered Maintenance (RCM), criticality assessment, defect elimination, reliability growth, and continuous improvement practices.
  • Build and lead a high-performing organization of managers, engineers, analysts, and technical specialists, establishing clear accountability, technical expectations, career development, and succession plans.
  • Own portfolio-level programs that improve reliability across critical data center infrastructure, including electrical distribution, UPS systems, generators, mechanical cooling systems, controls, automation, and supporting facility systems.

What they are looking for

  • 10+ years of progressive engineering, reliability, maintenance, critical facilities, or infrastructure experience, including significant experience directly supporting mission-critical environments.
  • Demonstrated experience working within mission-critical operations or engineering environments where infrastructure availability, redundancy, maintenance execution, and operational risk directly affect service continuity.
  • 5+ years of progressive leadership experience, including responsibility for engineering managers, senior technical professionals, or large multi-site technical organizations and programs.
  • Demonstrated technical knowledge of mission-critical infrastructure, including experience with electrical distribution, UPS systems, generators, mechanical cooling systems, controls/automation, and integrated facility operations.
  • Demonstrated experience developing and implementing reliability, maintenance, asset-management, or operational excellence strategies across mission-critical infrastructure.
  • Strong working knowledge of reliability engineering methodologies, including structured root cause analysis, FMEA/FMECA, RCM, criticality analysis, defect elimination, reliability metrics, and lifecycle risk management.
  • Demonstrated experience evaluating infrastructure failures, operational events, equipment performance, maintenance effectiveness, and systemic reliability risks within mission-critical environments.
  • Demonstrated ability to use operational and engineering data to identify systemic risks, establish priorities, and influence significant technical or business decisions.
  • Experience leading cross-functional initiatives involving Data Center Operations, Engineering, Design, Construction, Commissioning, Procurement, OEMs, vendors, and other technical stakeholders.
  • Experience establishing engineering governance, standards, KPIs, and management mechanisms across multiple sites, regions, or infrastructure programs in mission-critical environments.
  • Demonstrated ability to communicate complex technical risks, tradeoffs, and investment recommendations to senior and executive leadership.
  • Bachelor’s degree in Electrical Engineering, Mechanical Engineering, Industrial Engineering, Systems Engineering, Reliability Engineering, or a related technical discipline; or equivalent relevant industry experience.

Nice to have

  • Direct data center experience supporting mission-critical infrastructure and operations is preferred.
  • Experience leading reliability engineering, critical facilities engineering, or asset-management organizations within hyperscale, colocation, or large-scale enterprise data centers.
  • Experience supporting geographically distributed or global data center portfolios.
  • Deep technical expertise in one or more critical infrastructure domains, with broad working knowledge across electrical distribution, UPS, standby generation, mechanical cooling, controls/automation, and integrated facility operations.
  • Experience developing and scaling predictive maintenance, condition-based monitoring, failure trend analysis, asset health modeling, and equipment risk-ranking programs within data center environments.
  • Advanced knowledge of reliability, availability, and maintainability analysis; maintenance strategy optimization; spare parts planning; lifecycle modeling; and total cost of ownership.
  • Experience governing commissioning, operational acceptance, maintenance program design, and readiness of new or modified mission-critical data center infrastructure.
  • Experience with CMMS/EAM, DCIM, EPMS, BMS, monitoring, telemetry, an
Mission-Critical Technical LeadershipOrganizational LeadershipReliability StrategyTechnical JudgmentSystems ThinkingData-Driven Decision MakingExecutive InfluenceOperational ExcellenceFMEAFMECARCARCMCMMSEAMDCIMEPMS
Full posting text

You will own the vision, operating model, engineering standards, and portfolio programs that improve infrastructure availability, maintainability, resilience, and lifecycle performance at scale.

>>This position will be full-time on-site at Oracle's offices located in Nashville, TN. Relocation assistance may be available in accordance with Oracle’s relocation policies. Candidates should expect a minimum of 25% travel , with additional travel as business needs require.<< As Director of Building Automation , you will provide strategic and organizational leadership for the reliability engineering function supporting Oracle Cloud Infrastructure’s mission-critical data center portfolio. You will own the vision, operating model, engineering standards, and portfolio programs that improve infrastructure availability, maintainability, resilience, and lifecycle performance at scale. This role requires significant hands-on and leadership experience within mission-critical environments. Direct data center experience is preferred. Candidates must have demonstrated experience supporting the electrical, mechanical, controls, and operational systems required to maintain continuous operations in mission-critical environments. You will lead managers, engineers, analysts, and technical programs responsible for reliability engineering, asset performance, predictive maintenance, failure analysis, defect elimination, and lifecycle risk management across OCI's data center infrastructure. You will partner with senior leaders across Data Center Operations, Engineering, Design, Construction, Commissioning, Automation, Procurement, and other infrastructure organizations to translate operational experience and engineering data into long-term reliability strategy. You will ensure lessons learned from incidents, equipment performance, maintenance activities, and portfolio trends result in durable improvements to standards, designs, operating practices, and investment priorities. Success in this role requires the ability to operate at both strategic and technical levels—setting multi-year direction for the reliability organization while maintaining sufficient engineering depth and data center operational knowledge to challenge assumptions, assess complex infrastructure risks, and drive disciplined decision-making across a rapidly scaling global portfolio.

Key Responsibilities Lead the Reliability Engineering organization supporting multiple regions, sites, and infrastructure programs across OCI's mission-critical data center portfolio. Define the multi-year reliability engineering strategy, organizational roadmap, operating model, and investment priorities required to improve data center infrastructure resilience and support OCI's continued growth. Establish and govern portfolio-wide reliability engineering standards and methodologies, including FMEA/FMECA, RCA, Reliability-Centered Maintenance (RCM), criticality assessment, defect elimination, reliability growth, and continuous improvement practices. Build and lead a high-performing organization of managers, engineers, analysts, and technical specialists, establishing clear accountability, technical expectations, career development, and succession plans. Own portfolio-level programs that improve reliability across critical data center infrastructure, including electrical distribution, UPS systems, generators, mechanical cooling systems, controls, automation, and supporting facility systems. Establish a comprehensive reliability measurement framework that provides leadership with visibility into asset health, failure trends, systemic risks, repeat events, corrective actions, maintenance effectiveness, and lifecycle exposure . Define reliability KPIs, targets, governance mechanisms, and executive reporting that enable data-driven prioritization of operational and engineering investments. Establish governance for corrective and preventive actions resulting from incidents, root cause analyses, audits, equipment failures, and reliability trend reviews, ensuring actions are completed, verified for effectiveness, and sustained. Drive systematic identification and elimination of recurring and systemic failure modes across the data center portfolio rather than relying solely on site-specific remediation. Sponsor the development and adoption of predictive and condition-based maintenance capabilities , including monitoring, analytics, automation, asset health modeling, and emerging technologies that improve early detection of equipment degradation and failure risk. Partner with Data Center Operations leadership to continuously improve maintenance strategy, operational readiness, troubleshooting practices, procedures, failure response, and infrastructure risk management. Provide reliability governance and technical leadership for commissioning, acceptance testing, operational handover, major maintenance, retrofits, capacity expansion, and infrastructure lifecycle decisions . Establish portfolio approaches to asset lifecycle management, including equipment health, utilization, failure history, remaining useful life, obsolescence, spare parts strategy, replacement planning, and end-of-life risk. Partner with Design, Construction, Engineering, and Procurement leadership to ensure lessons from operating facilities influence equipment specifications, design standards, redundancy strategies, maintainability requirements, vendor selection, and total cost of ownership. Develop mechanisms to convert site-level events and engineering findings into portfolio-wide standards, design changes, maintenance improvements, and risk-reduction programs . Lead technical and business reviews of significant reliability risks and provide clear recommendations regarding mitigation strategies, priorities, investment requirements, and residual operational risk. Develop strong partnerships with equipment manufacturers, service providers, and technology partners to improve equipment performance, failure intelligence, serviceability, and long-term reliability. Establish effective operating rhythms for the organization, including portfolio reviews, technical reviews, risk escalation, program governance, resource prioritization, and executive communications. Represent Reliability Engineering in senior leadership discussions involving infrastructure risk, operational performance, capacity growth, capital planning, and long-term data center strategy. Minimum Qualifications 10+ years of progressive engineering, reliability, maintenance, critical facilities, or infrastructure experience, including significant experience directly supporting mission-critical environments. Demonstrated experience working within mission-critical operations or engineering environments where infrastructure availability, redundancy, maintenance execution, and operational risk directly affect service continuity. 5+ years of progressive leadership experience, including responsibility for engineering managers, senior technical professionals, or large multi-site technical organizations and programs. Demonstrated technical knowledge of mission-critical infrastructure, including experience with electrical distribution, UPS systems, generators, mechanical cooling systems, controls/automation, and integrated facility operations. Demonstrated experience developing and implementing reliability, maintenance, asset-management, or operational excellence strategies across mission-critical infrastructure. Strong working knowledge of reliability engineering methodologies, including structured root cause analysis, FMEA/FMECA, RCM, criticality analysis, defect elimination, reliability metrics, and lifecycle risk management. Demonstrated experience evaluating infrastructure failures, operational events, equipment performance, maintenance effectiveness, and systemic reliability risks within mission-critical environments. Demonstrated ability to use operational and engineering data to identify systemic risks, establish priorities, and influence significant technical or business decisions. Experience leading cross-functional initiatives involving Data Center Operations, Engineering, Design, Construction, Commissioning, Procurement, OEMs, vendors, and other technical stakeholders. Experience establishing engineering governance, standards, KPIs, and management mechanisms across multiple sites, regions, or infrastructure programs in mission-critical environments. Demonstrated ability to communicate complex technical risks, tradeoffs, and investment recommendations to senior and executive leadership. Bachelor’s degree in Electrical Engineering, Mechanical Engineering, Industrial Engineering, Systems Engineering, Reliability Engineering, or a related technical discipline; or equivalent relevant industry experience. Skills and Competencies Mission-Critical Technical Leadership: Strong understanding of mission-critical infrastructure, operating practices, redundancy, maintenance risk, failure modes, and the interdependencies between electrical, mechanical, controls, and operational systems. Organizational Leadership: Ability to build, develop, and lead managers and senior technical professionals while establishing clear accountability and a strong engineering culture. Reliability Strategy: Ability to translate data center infrastructure performance, business growth, and operational risk into a coherent multi-year reliability strategy and investment roadmap. Technical Judgment: Ability to evaluate complex infrastructure reliability issues, challenge technical assumptions, understand operational consequences, and make decisions under uncertainty. Systems Thinking: Ability to connect individual equipment failures and site-level events to systemic portfolio risks involving design, maintenance, operations, suppliers, processes, or organizational practices. Data-Driven Decision Making: Ability to convert large volumes of operational, maintenance, failure, and asset data into actionable insights and investment priorities. Executive Influence: Ability to communicate technical risk and recommendations clearly to senior leaders and build alignment across organizations with different objectives and priorities. Operational Excellence: Strong commitment to disciplined execution, corrective-action rigor, standardization, measurable improvement, and sustained results. Change Leadership: Ability to introduce and scale new engineering methods, technologies, processes, and operating models across a large and geographically distributed data center organization. Talent Development: Demonstrated ability to develop engineering leaders and technical talent, establish career paths, strengthen organizational capability, and build succession depth. Preferred Qualifications Direct data center experience supporting mission-critical infrastructure and operations is preferred. Experience leading reliability engineering, critical facilities engineering, or asset-management organizations within hyperscale, colocation, or large-scale enterprise data centers. Experience supporting geographically distributed or global data center portfolios. Deep technical expertise in one or more critical infrastructure domains, with broad working knowledge across electrical distribution, UPS, standby generation, mechanical cooling, controls/automation, and integrated facility operations. Experience developing and scaling predictive maintenance, condition-based monitoring, failure trend analysis, asset health modeling, and equipment risk-ranking programs within data center environments. Advanced knowledge of reliability, availability, and maintainability analysis; maintenance strategy optimization; spare parts planning; lifecycle modeling; and total cost of ownership. Experience governing commissioning, operational acceptance, maintenance program design, and readiness of new or modified mission-critical data center infrastructure. Experience with CMMS/EAM, DCIM, EPMS, BMS, monitoring, telemetry, an

Apply on company site

Meet Ori - your career agent on WhatsApp

Find jobs, get your roadmap, check if you're ready for a role and prepare applications - in chat, any language.

Ask Ori about this role
Checking your fit…

Opportunity details

Deadline
Closing in 171d · 21 Mar

As stated by the source. Anything not shown was not stated.

Oracle

IT Services and IT Consulting

Oracle is a global leader in AI, delivering the cloud infrastructure, data, and applications that organizations across the world trust to successfully achieve business outcomes at scale. Oracle Cloud Infrastructure (OCI) provides fast, flexible, scalable AI infrastructure. With superior compute performance and network design, a comprehensive choice of AI services for developing and orchestrating agentic AI workflows at scale, and unrivaled data control, security, privacy, and governance, OCI is designed for AI workloads. It also gives customers the flexibility to run their workloads wherever t

Company pageWebsite

More at Oracle

OracleSoftware EngineeringPreferredFunded

Remote · US

Senior · 22h ago

Senior Platform Software Engineer
DebuggingTroubleshooting+5 more
View role
OraclePreferredFunded

On-site · United States

22h ago

APEX Developer
DevOps engineering+3 more
View role
OracleTechnical SupportPreferredFunded

On-site · DUBAI, United Arab Emirates

Senior · 5+ yrs · 22h ago

Senior Support Engineer -Nursing Integration/Phys Doc/Clinical AI Agent
HCIT supportTroubleshooting+6 more
View role
OraclePreferredFunded

On-site · Morocco

Senior · 22h ago

Senior Principal Software Engineer LiveLabs Platform
View role
Apply