About the role
structured by ORIWill be joining the GBP Production Engineering team - Production-focused engineer with strong troubleshooting, observability, incident-management, cloud, and operational-automation experience. This position focuses on production support and governance for operational excellence, including: - Production support,…
What you will do
- Production support, observability, and production governance/management
- Designing operating models and process controls
- Participating in audits and compliance activities
- Designing dashboards and views
- Proficient use of Copilot to improve day-to-day efficiency
What they are looking for
- 6+ years of IT troubleshooting, resolution, and maintenance experience.
- Experience with large-scale on-premises and public-cloud environments, preferably AWS .
- Strong observability and monitoring knowledge, including Grafana, Dynatrace, Prometheus, Datadog, and Splunk .
- Familiarity with CI/CD and infrastructure tools such as Jenkins, GitLab, and Terraform .
- Familiarity with EKS, Kubernetes, Docker , and common networking issues.
- A proactive approach to learning and recommending new technologies
Full posting text
Will be joining the GBP Production Engineering team - Production-focused engineer with strong troubleshooting, observability, incident-management, cloud, and operational-automation experience.
This position focuses on production support and governance for operational excellence, including: - Production support, observability, and production governance/management - Designing operating models and process controls - Participating in audits and compliance activities - Designing dashboards and views - Proficient use of Copilot to improve day-to-day efficiency - Providing production stability metrics and insights to management - Acting as a gatekeeper for GBP production processes and standards - Leading meetings, driving decisions, and coordinating cross-functional stakeholders Key technical skills required: - AWS - Logging and monitoring tools: Splunk, Dynatrace, Datadog, CloudWatch - ServiceNow - Kafka - API design/integration and troubleshooting - Incident, problem, and change management (ITIL-aligned) - SRE/operations metrics (availability, reliability, MTTR, SLA/SLO reporting) - Root cause analysis and post-incident governance - SQL and data analysis for operational reporting - CI/CD awareness and release governance - Strong communication, stakeholder management, and documentation skills Required experience and skills 6+ years of IT troubleshooting, resolution, and maintenance experience. Experience with large-scale on-premises and public-cloud environments, preferably AWS . Strong observability and monitoring knowledge, including Grafana, Dynatrace, Prometheus, Datadog, and Splunk . Familiarity with CI/CD and infrastructure tools such as Jenkins, GitLab, and Terraform . Familiarity with EKS, Kubernetes, Docker , and common networking issues. A proactive approach to learning and recommending new technologies