21 Sep
|
HCLTech
|
Victoria
Job Description
SIAM MIM & PM Manager / Endurance Manager
n
WeareHCLTech,oneofthefastest-growinglargetechcompaniesintheworldandhometo219,000+peopleacross54countries,superchargingprogressthroughindustry-leadingcapabilitiescenteredaroundDigital,EngineeringandCloud.
n
Thedrivingforcebehindthatwork,ourpeople,arediverse,creative,andpassionate,raisingthebarforexcellenceonaregularbasis. We,inturn,workhardtobringoutthebestinthemaswestrivetohelpthemfindtheirsparkandbecomethebestversionofthemselvesthattheycanbe.
n
Areyoureadytobeanimportantpartofthisever-transformationaljourney?
n
Job Description
n
Duration:6MonthsFTCRole
n
HybridRole:3daysfromofficeismandatory
n
Position Overview
n
WearelookingforanexperiencedSIAMMIM&PMManager/EnduranceManagertoleadMajorIncidentManagement,ProblemManagement,andserviceenduranceactivitiesacrosscriticalITservices. Therolewillberesponsibleforensuringrapidandeffectiverestorationofservicesduringmajorincidents,drivingpermanentresolutionofrecurringissues,andimprovingtheoverallresilience,stability,andoperationalenduranceoftheITenvironment.
n
ThesuccessfulcandidatewillworkcloselywithServiceManagement,Infrastructure,Applications,Engineering,Security,Cloud,Vendors,andBusinessstakeholderstominimizebusinessimpactandcontinuouslyimproveservicereliability.
n
Key Responsibilities
n
n
- Own and lead the end-to-end Major Incident Management process.
n
- Ensure timely incident declaration, stakeholder communication, escalation, coordination, and service restoration.
n
- Drive incident resolution within an agreed Service Level Agreements (SLAs) and business expectations.
n
- Coordinate bridge calls, war rooms, technical investigations, and executive communications.
n
- Review major incidents and ensure comprehensive Post-Incident Reviews (PIRs) are completed.
n
- Track corrective and preventive actions resulting from major incidents.
n
- Maintain and continuously improve MIM procedures, playbooks, escalation matrices, and communication templates.
n
n
Problem Management (PM)
n
n
- Own and manage the end-to-end Problem Management lifecycle.
n
- Identify recurring incidents, trends, systemic failures, and operational risks.
n
- Lead Root Cause Analysis (RCA) for major and recurring incidents.
n
- Drive permanent resolution rather than repeated incident restoration.
n
- Work with engineering and technical team to eliminate known errors and recurring failures.
n
- Track Problem Management KPIs and demonstrate reduction in recurring incidents.
n
- Identify trends and proactively highlight potential service risks before they become major incidents.
n
n
Endurance & Service Resilience Management
n
n
- Assess service resilience against operational failures, capacity constraints, technology risks, and recurring incidents.
n
- Identify single points of failure and coordinate remediation plans.
n
- Drive service improvement initiatives based on incident and problem trends.
n
- Ensure critical services have appropriate monitoring, alerting, recovery procedures, support models, and operational documentation.
n
- Partner with technology teams to improve service availability, reliability, recoverability, and stability.
n
- Ensure lessons learned from incidents are incorporated into service design and operational practices.
n
- Support operational readiness for new applications, platforms, infrastructure, and major technology changes.
n
n
Governance & Reporting
n
n
- Produce regular operational reports and management dashboards.
n
- Mean TimetoAcknowledge(MTTA)
n
- RCAcompletionandquality
n
- Ensure compliance with IT Service Management policies, processes, and audit requirements.
n
n
Stakeholder & Vendor Management
n
n
- Build solid relationships with technology, business, service management, and senior leadership teams.
n
- Provide clear and concise communication during high-pressure situations.
n
- Manage escalations across internal teams and external service providers.
n
- Hold vendors accountable for incident resolution, RCA quality, SLA performance, and corrective actions.
n
- Facilitate cross-functional collaboration to resolve complex and business-critical issues.
n
- Act as a trusted point of escalation for service stability and resilience concerns.
n
n
Leadership Responsibilities
n
n
- Conduct regular performance and capability reviews.
n
- Develop team skills through coaching, knowledge sharing, simulations, and incident exercises.
n
- Ensure adequate coverage and on-call/escalation arrangements for critical services.
n
n
Required Experience & Skills
n
n
- 8–12+ years of experience in IT Service Management, Major Incident Management, Problem Management, or IT Operations.
n
- Strong experience leading Major Incident Management for complex enterprise IT environments.
n
- Experience with service resilience, availability, reliability, disaster recovery, or operational risk management.
n
- Experience working with multiple technology towers such as:
n
- Infrastructure
n
- Cloud
n
- Networks
n
- Applications
n
- Databases
n
- Cybersecurity
n
- End-user computing
n
- Strong analytical and problem-solving capabilities.
n
- Experience working with third-party vendors and managed service providers.
n
n
Preferred Qualifications
n
n
- Experience with ITSM platforms such as ServiceNow, BMC Remedy, Jira, or similar.
n
- Experience with Agile/DevOps/SRE operating models.
n
- Experience with operational resilience, service continuity, or disaster recovery.
n
- Experience with automation, observability, monitoring, and service reliability practices.
n
- Relevant certifications in Problem Management, SRE, Service Management, or Business Continuity are desirable.
n
📌 SIAM MIM & PM Manager / Endurance Manager (Victoria)
🏢 HCLTech
📍 Victoria