About the role
structured by ORISite Reliability Engineer part of Oracle Analytics Service Excellence (OASE) UK GOV SRE team partnering with Oracle Analytics development teams to improve the reliability, availability, performance, operational support and maturity of Oracle Analytics Cloud services. Site Reliability Engineer part of Oracle Analytics…
What you will do
- Perform SRE activities supporting Oracle Analytics customers, engineering teams, and release cycles in both pre-production and production environments.
- Participate in a follow-the-sun model providing 24x7 operational support for Oracle Analytics services.
- Respond to incidents, troubleshoot complex service issues, drive mitigation to completion, and contribute to root-cause analysis and post-incident actions.
- Read, understand, and troubleshoot existing application and service code to diagnose complex production issues, identify root causes, and support safe remediation.
- Build, maintain, and improve operational tooling, automation, dashboards, monitoring, run-books, and knowledge-base documentation.
What they are looking for
- BS or MS in Computer Science, Engineering, or equivalent practical experience.
- Experience supporting cloud infrastructure, networking, applications, services, tools, and operational processes.
- Strong understanding of networking and TCP/IP fundamentals, including DNS, HTTP/HTTPS, TLS, load balancing, and service connectivity.
- Linux/Unix system-administration experience, including troubleshooting processes, memory, CPU, filesystem, and network issues.
- Experience developing, operating, or supporting cloud services and large-scale distributed applications in production.
- Demonstrated ability to troubleshoot complex technical issues methodically, including investigation of existing applications and code.
- Experience creating and maintaining technical documentation, runbooks, knowledge articles, and operational guides.
- Experience working in agile development and operational environments.
- Strong written and verbal communication skills, including the ability to work effectively with remote global teams.
- Ability to work independently, manage competing priorities, and participate in on-call, after-hours maintenance, and weekend support as needed.
Nice to have
- Experience with Oracle Analytics Cloud, Oracle Analytics Server, OBIS, BI Publisher, Oracle Database, Autonomous Database, MySQL, SQL Server, or NoSQL technologies.
- Two to four years of experience operating large-scale, customer-facing web applications or cloud services.
- Experience with OCI, AWS, Azure, or GCP compute, storage, networking, monitoring, and operational tooling.
- Programming and scripting experience with Python, Bash, JavaScript/Node.js, Ansible, and related technologies; Java experience is a plus.
- Ability to read, understand, troubleshoot, and safely modify existing enterprise application code.
- Familiarity with AI-assisted development tools, such as Codex and Claude Code, for software development, automation, investigation, and documentation.
- Experience with CI/CD and infrastructure automation tools such as Ansible, Puppet, Chef, Git, and deployment pipelines.
- Experience with cloud-native applications, containers, Kubernetes, microservices, and independently scalable services.
Full posting text
Site Reliability Engineer part of Oracle Analytics Service Excellence (OASE) UK GOV SRE team partnering with Oracle Analytics development teams to improve the reliability, availability, performance, operational support and maturity of Oracle Analytics Cloud services.