Jobs / Oracle / Senior Site Reliability Engineer

Senior Site Reliability Engineer

Oracle · ✓ Verified company
📍 BENGALURU, KARNATAKA, India · On-site · Full-time

About the role

Job Responsibilities - Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components. - Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions. - Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems. - Build automation and tooling to reduce operational toil and improve production safety. - Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements. - Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response. - Support upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts. - Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes. - Contribute to incident-management practices, operational readiness, and service ownership improvements. - Share technical knowledge and support team members through documentation, reviews, and collaboration. - Participate in a 12x7 on-call rotation and support response to customer-impacting incidents. Mandatory Skills - 4–8 years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role. - Experience operating and improving highly available production systems. - Strong programming or scripting skills in Python, Java, Go, or similar languages. - Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage. - Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing. - Experience owning or improving service SLIs, SLOs, KPIs, and operational procedures. - Strong incident troubleshooting, RCA, debugging, and problem-solving skills. - Experience with deployment pipelines, release validation, automation, and change-management practices. - Understanding of distributed systems, service dependencies, capacity planning, and performance tuning. - Ability to work independently on technical problems and collaborate effectively with engineering teams. - Strong written and verbal communication skills. Preferred Skills - Experience with OCI and cloud infrastructure services. - Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation. - Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD. - Experience with service migrations, fleet maintenance, upgrades, patching, or production rollouts. - Experience with architecture reviews, operational-readiness reviews, and post-incident improvements. - Experience contributing to technical initiatives, knowledge sharing, code reviews, or operational improvements within the team. - Familiarity with security, compliance, and access-control practices in production environments. Self-Test Questions  - Do you have 4–8 years of relevant SRE, Production Engineering, Cloud Operations, or Systems Engineering experience? - Have you  independently operated or improved a production service, system, or infrastructure component? - Can you investigate production incidents and contribute to mitigation, recovery, RCA, and follow-up actions? - Do you have hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage? - Are you proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting? - Have you built or improved automation, deployment validation, CI/CD pipelines, or operational tooling? - Do you have experience with monitoring, alerting, logs, metrics, tracing, and service health indicators such as SLOs or KPIs? - Can you work independently on assigned technical problems, collaborate with partner teams, and participate in a 12x7 on-call rotation? Career Level - IC3