About the role
from listingDrive reliability and scalability for AI/ML Data Platforms through innovative solutions and collaborative teamwork.
Join a dynamic team where your expertise in site reliability engineering will shape the future of AI/ML data platforms. Unlock opportunities for growth and impact as you help build resilient, market-leading solutions. As a Site Reliability Engineer III at JPMorgan Chase within the AI/ML Data Platforms team, you will play a pivotal role in developing scalable and resilient data solutions. You will engage in root cause analysis, production changes, and strategic initiatives that drive operational excellence. Your experience will help mentor team members and foster collaboration across global teams. Together, we create innovative solutions that support the firm’s mission and community. Job responsibilities Build and support scalable, resilient AI/ML data solutions Coordinate incident management coverage for effective application issue resolution Collaborate with cross-functional teams to perform root cause analysis and implement production changes Develop and support AI/ML solutions for troubleshooting and incident resolution Mentor and guide team members to drive strategic change Manage budgetary considerations and staffing challenges Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements. Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes. Partner with colleagues across global teams to deliver impactful results Required qualifications, capabilities and skills Formal training or certification on site reliability engineering concepts and 3+ years applied experience Proficient in site reliability culture and principles, with familiarity in implementing site reliability within an application or platform Proficiency in running production incident calls and managing incident resolution Experience in observability including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and others Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements Strong understanding of SLI/SLO/SLA, Error Budgets and Proficiency in Python or PySpark for AI/ML modeling Must be able to reduce toil by building new tools to automate repeated tasks Hands-on experience in system design, resiliency, testing, operational stability, and disaster recovery Awareness of risk controls and compliance with departmental and company-wide standards Ability to work collaboratively in teams and build meaningful relationships to achieve common goals Preferred qualifications, capabilities and skills 4+ years in an SRE or production support role with AWS Cloud, Databricks, Snowflake or similar technologies AWS and Databricks certifications