Head - Incident Management, Monitoring & Operational Resilience
Poonawalla Fincorp · India
Job Description
Role Overview
We are looking for a senior leader to head the Incident Management function, ensuring high availability, regulatory compliance, and operational resilience of all critical systems, including lending platforms, payment integrations, customer channels, collections systems, and risk platforms.
This role will be pivotal in safeguarding business continuity, customer trust, and regulatory adherence while driving the organization's transition toward proactive, intelligence-driven operations through advanced monitoring, observability, automation, and governance.
Key Responsibilities
1. Major Incident Management (Financial Services Criticality)
- Own end-to-end Major Incident Management (MIM) across all production systems.
- Ensure rapid resolution of business-critical incidents impacting loan origination, collections, payments, customer servicing, and digital channels.
- Lead 24x7 incident command center (war room) activities for high-severity incidents.
- Establish and continuously improve incident severity classification, escalation paths, and response playbooks.
- Drive Root Cause Analysis (RCA), Problem Management, and implementation of corrective and preventive actions.
- Ensure executive stakeholders receive timely and accurate communication during critical incidents.
2. Monitoring & Observability
- Define and execute enterprise monitoring and observability strategy across:
- Core lending systems
- Payment gateways and partner integrations
- Customer digital channels (Web, Mobile Apps, APIs)
- Infrastructure, Cloud, Database, and Middleware platforms
- Implement real-time monitoring and alerting frameworks with business-impact correlation.
- Drive adoption and optimization of observability platforms including Grafana, Dynatrace, AppDynamics, Splunk, and cloud-native monitoring solutions.
- Establish service health dashboards, synthetic monitoring, end-user experience monitoring, and transaction observability.
- Improve alert quality and reduce false positives through continuous tuning and automation.
3. Regulatory Compliance & Risk Management
- Ensure adherence to RBI guidelines, cybersecurity directives, IT governance standards, and audit requirements.
- Maintain audit-ready documentation for:
- Incident logs
- RCA reports
- Regulatory incidents
- SLA breaches and corrective actions
- Partner with Risk, Compliance, Information Security, Internal Audit, and External Audit teams.
- Ensure technology incidents are handled in alignment with regulatory reporting obligations and cybersecurity policies.
- Drive operational risk reduction through proactive identification and mitigation of technology vulnerabilities.
4. SLA, Availability & Business Continuity
- Define, monitor, and enforce SLAs, OLAs, and operational KPIs for mission-critical systems.
- Drive high-availability practices supporting uptime targets of 99.9% and above.
- Integrate Incident Management with:
- Disaster Recovery (DR)
- Business Continuity Planning (BCP)
- Crisis Management Frameworks
- Lead periodic DR drills, resilience testing, and crisis simulation exercises.
- Ensure readiness for large-scale outages, cyber incidents, and third-party service disruptions.
5. ITSM & Process Governance
- Drive ITSM processes aligned with ITIL best practices.
- Integrate Incident, Problem, Change, Event, and Knowledge Management processes.
- Establish governance forums, performance reviews, and service improvement programs.
- Standardize incident lifecycle management and stakeholder communication protocols.
- Drive automation of incident detection, routing, escalation, and resolution workflows.
6. Stakeholder & Business Alignment
- Act as the senior escalation point during critical outages impacting customers, operations, or regulatory commitments.
- Provide real-time updates to leadership, business teams, regulators, and external partners where required.
- Collaborate closely with Technology, Operations, Risk, Compliance, Product, and Customer Service teams.
- Ensure technology operations remain aligned with business priorities and customer experience objectives.
7. Transformation & Automation
- Lead the transformation from reactive support models to proactive and predictive operations.
- Drive implementation of event correlation, anomaly detection, predictive alerting, and auto-remediation capabilities.
- Leverage AI-driven operational insights to improve service reliability and operational efficiency.
- Reduce incident volumes, improve Mean Time to Detect (MTTD), and improve Mean Time to Resolution (MTTR).
8. Team Leadership & Capability Building
- Build and lead high-performing 24x7 Incident Management, Monitoring, and NOC teams.
- Develop leadership pipeline and succession planning within operations teams.
- Upskill teams on:
- Incident Command Practices
- ITIL and ITSM Processes
- Monitoring & Observability
- Financial Systems Criticality
- Customer Impact Management
- Foster a culture of ownership, resilience, collaboration, and customer-centricity.
Required Qualifications & Experience
- Bachelor's degree in engineering, Computer Science, Information Technology, or related discipline.
- 18–25 years of experience in IT Operations, Service Management, Infrastructure Operations, or Production Support.
- Minimum 8–10 years of leadership experience managing enterprise-scale Incident Management and Monitoring functions.
- Strong experience managing major incidents in high-availability production environments.
- Proven leadership in service governance, operational excellence, and enterprise technology transformation.
- Deep understanding of ITIL, ITSM frameworks, and operational resilience principles.
- Mandatory experience within Banking, NBFC, FinTech, Payments, or broader BFSI environments.
- Experience handling regulatory, audit, compliance, and risk management requirements.
Details
| Company | Poonawalla Fincorp |
| Location | India |
| Type | FULL TIME |
| Niche | general |
