AI AdoptionFunded CompaniesJob SimulationCertificationsRoadmapsJobsPricing
Sign In
OneRoadmap

Empowering the next generation with AI education. Custom training for colleges and enterprises.

gaurav.ghai@oneroadmap.in
Delhi NCR, India

Platform

  • AI Roadmaps
  • Free Certifications
  • Learning Resources
  • Pricing

Training

  • AI Adoption Workshops
  • Expert Sessions
  • Upcoming Events
  • Workshop Gallery

Company

  • Blog
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds
  • Delete your data

© 2026 OneRoadmap

Operated by Ghai Technologies, India · International operations through One Roadmap Marketing Management, Dubai, UAE

Built for your next chapter.

Open roles

· Engineering Services

Senior Site Reliability Engineer

Oracle

Senior · 4+ yrsOn-site · BENGALURU, KARNATAKA, IndiaListed 1d ago
Apply now

Backed by

VC portfolio

HQ

🇺🇸 Reading, United Kingdom

Open roles

2042

Experience Senior · 6+ yrs (4–8 years)

About the role

from listing

Takes proactive steps to design and architect infrastructure and service to ensure reliability and functionality.

Job Responsibilities Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components. Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions. Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems. Build automation and tooling to reduce operational toil and improve production safety. Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements. Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response. Support upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts. Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes. Contribute to incident-management practices, operational readiness, and service ownership improvements. Share technical knowledge and support team members through documentation, reviews, and collaboration. Participate in a 12x7 on-call rotation and support response to customer-impacting incidents.

Mandatory Skills 4–8 years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role. Experience operating and improving highly available production systems. Strong programming or scripting skills in Python, Java, Go , or similar languages. Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage. Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing. Experience owning or improving service SLIs, SLOs, KPIs , and operational procedures. Strong incident troubleshooting, RCA, debugging, and problem-solving skills. Experience with deployment pipelines, release validation, automation, and change-management practices. Understanding of distributed systems, service dependencies, capacity planning, and performance tuning. Ability to work independently on technical problems and collaborate effectively with engineering teams. Strong written and verbal communication skills. Preferred Skills Experience with OCI and cloud infrastructure services. Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation. Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD. Experience with service migrations, fleet maintenance, upgrades, patching, or production rollouts. Experience with architecture reviews, operational-readiness reviews, and post-incident improvements. Experience contributing to technical initiatives, knowledge sharing, code reviews, or operational improvements within the team. Familiarity with security, compliance, and access-control practices in production environments. Self-Test Questions Do you have 4–8 years of relevant SRE, Production Engineering, Cloud Operations, or Systems Engineering experience? Have you independently operated or improved a production service, system, or infrastructure component? Can you investigate production incidents and contribute to mitigation, recovery, RCA, and follow-up actions? Do you have hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage? Are you proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting? Have you built or improved automation, deployment validation, CI/CD pipelines, or operational tooling? Do you have experience with monitoring, alerting, logs, metrics, tracing, and service health indicators such as SLOs or KPIs? Can you work independently on assigned technical problems, collaborate with partner teams, and participate in a 12x7 on-call rotation ?

Career Level - IC3

hot job
Apply on company site

Meet Ori - your career agent on WhatsApp

Find jobs, get your roadmap, check if you're ready for a role and prepare applications - in chat, any language.

Ask Ori about this role
Checking your fit…
Oracle

IT Services and IT Consulting

Oracle is a global leader in AI, delivering the cloud infrastructure, data, and applications that organizations across the world trust to successfully achieve business outcomes at scale. Oracle Cloud Infrastructure (OCI) provides fast, flexible, scalable AI infrastructure. With superior compute performance and network design, a comprehensive choice of AI services for developing and orchestrating agentic AI workflows at scale, and unrivaled data control, security, privacy, and governance, OCI is designed for AI workloads. It also gives customers the flexibility to run their workloads wherever t

Company pageWebsite

More at Oracle

Principal Applied Scientist

Oracle · Data and Applied Science

Senior · 5+ yrsFunded

On-site · United States / Seattle, WA, United States / Santa Clara, CA, United States / Austin, TX, United States / Nashville, TN, United States

AI / ML

Full time · 1d ago
View role
Senior Machine Learning Engineer

Oracle · Data and Applied Science

Senior · 5+ yrsFunded

On-site · United States / Seattle, WA, United States / Santa Clara, CA, United States / Austin, TX, United States / Nashville, TN, United States

AI / ML

Full time · 1d ago
View role
Senior Manager, Core Infrastructure Engineering

Oracle

Senior · 5+ yrsFunded

Remote · US

1d ago
View role
Software Developer 3

Oracle · Software Engineering

Senior · 3+ yrsFunded

Remote · US

SOFTWARE

Full time · 1d ago
View role
Apply