Senior Consultant | ITSM | Bengaluru | Engineering | Platform Development & Integration
Deloitte
· ✓ Verified company
📍 Bengaluru, India · On-site · Full-time
About the role
Senior Consultant | ITSM | Bengaluru | Engineering | Platform Development & Integration • Job requisition ID : 109992 • Location: Bengaluru • Entity: Deloitte Touche Tohmatsu India LLP Senior Consultant | Bangalore | Platform Development & Integration | ITSM The team Our Enterprise Technology & Performance team helps organizations build, operate, and optimize resilient cloud-native platforms. We are looking for an experienced Lead Production Incident Manager (IM) to lead enterprise production operations, incident management, and cloud infrastructure reliability across large-scale AWS environments. The ideal candidate should possess strong expertise in production support, Site Reliability Engineering (SRE), cloud technologies, and ITIL-based service management, preferably with experience in Banking & Financial Services, specifically Cards & Payments and Mobile Applications. Enterprise technology has to do much more than keep the wheels turning; it is the engine that drives functional excellence and the enabler of innovation and long-term growth. Learn more about: Customer Role Summary - We are seeking an experienced Lead Production Incident Manager (IM) to lead enterprise production operations, major incident management, and cloud infrastructure reliability across large-scale AWS environments. The ideal candidate will be responsible for ensuring secure, scalable, highly available, and cost-effective cloud operations while driving operational excellence, service reliability, and continuous improvement across mission-critical enterprise applications. This role requires strong expertise in ITIL-based service management, Site Reliability Engineering (SRE), AWS cloud technologies, production support, automation, and stakeholder management. Experience in the Banking & Financial Services domain, particularly Cards & Payments, Mobile Applications, and Cloud-Native Solutions, will be highly advantageous. - Lead enterprise-wide Incident, Problem, and Change Management activities aligned with ITIL best practices. - Own the end-to-end lifecycle of production incidents (P1–P4), ensuring timely identification, escalation, communication, resolution, and closure. - Drive service recovery, Root Cause Analysis (RCA), Post Incident Reviews (PIR), and Corrective & Preventive Actions (CAPA) to improve service reliability and reduce MTTR. - Lead 24x7 production support operations, managing L2/L3 application and infrastructure support across Linux-based environments. - Oversee production deployments, release management, infrastructure operations, and Site Reliability Engineering (SRE) initiatives. - Manage Disaster Recovery (DR), High Availability (HA) architecture, and Active-Active/Active-Passive failover strategies. - Drive automation, cloud infrastructure optimization, capacity planning, and operational excellence using AWS, Kubernetes, Docker, and CI/CD pipelines. - Monitor and optimize production environments using observability tools including Grafana, Kibana, ELK, Splunk, CloudWatch, Prometheus, Nagios, Zenduty, and Site24x7. - Support REST API-based applications, perform production troubleshooting using SQL, and collaborate with engineering teams to ensure highly resilient and secure production environments. - Lead and mentor SRE and Production Support teams while ensuring SLA, SLO, KPI, and customer satisfaction targets are consistently achieved. Key Skills Required - Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related discipline. - 10–15+ years of overall IT experience with at least 8+ years specializing in Enterprise Production Support, Incident Management, Cloud Infrastructure, and Site Reliability Engineering (SRE). - Strong experience in Banking & Financial Services, preferably supporting Cards & Payments, Mobile Applications, and Cloud-Native solutions. - Proven expertise in Incident, Problem & Change Management (ITIL), Major Incident Management, Service Recovery, Root Cause Analysis (RCA), and Production Operations. - Experience leading 24x7 production support teams and managing L2/L3 application and infrastructure support across Linux environments. - Strong knowledge of AWS services including EC2, S3, RDS, Lambda, VPC, IAM, DynamoDB, CloudWatch, and secure cloud architecture. - Experience designing highly available, scalable, secure, and disaster recovery-enabled cloud solutions. - Hands-on experience with Docker, Kubernetes, Jenkins, CI/CD pipelines, Infrastructure as Code (Terraform, AWS CloudFormation, Ansible), and automation using Python or Bash. - Strong understanding of Linux administration, networking, load balancing, SQL/NoSQL databases, REST APIs, and cloud security best practices. - Experience with monitoring and observability tools including Grafana, Kibana, ELK Stack, Splunk, CloudWatch, Prometheus, Nagios, Zenduty, and Site24x7. - Proven ability to drive SLA/SLO/KPI compliance, reduce MTTR, improve service reliability, and implement automation and continuous improvement initiatives. - Strong leadership, stakeholder management, client communication, analytical, and problem-solving skills with the ability to coordinate cross-functional teams during critical incidents. - ITIL Foundation Certification is preferred. - AWS Certified Solutions Architect – Associate or Professional certification is highly preferred.