22 Sep
|
HCLTech
|
Victoria
Job Description
SIAM MIM & PM Manager / Endurance Manager
n
WeareHCLTech,oneofthefastest
-
growinglargetechcompaniesintheworldandhometo219,000+peopleacross54countries,superchargingprogressthroughindustry
-
leadingcapabilitiescenteredaroundDigital,EngineeringandCloud.
n
Thedrivingforcebehindthatwork,ourpeople,arediverse,creative,andpassionate,raisingthebarforexcellenceonaregularbasis.
We,inturn,workhardtobringoutthebestinthemaswestrivetohelpthemfindtheirsparkandbecomethebestversionofthemselvesthattheycanbe.
n
Areyoureadytobeanimportantpartofthisever
-
transformationaljourney?
n
Job Description
n
Duration:6MonthsFTCRole
n
HybridRole:3daysfromofficeismandatory
n
Position Overview
n
WearelookingforanexperiencedSIAMMIM
&
PMManager/EnduranceManagertoleadMajorIncidentManagement,ProblemManagement,andserviceenduranceactivitiesacrosscriticalITservices.
Therolewillberesponsibleforensuringrapidandeffectiverestorationofservicesduringmajorincidents,drivingpermanentresolutionofrecurringissues,andimprovingtheoverallresilience,stability,andoperationalenduranceoftheITenvironment.
n
ThesuccessfulcandidatewillworkcloselywithServiceManagement,Infrastructure,Applications,Engineering,Security,Cloud,Vendors,andBusinessstakeholderstominimizebusinessimpactandcontinuouslyimproveservicereliability.
n
Key Responsibilities
n
n
Own and lead the end-to-end Major Incident Management process.
n
Ensure timely incident declaration, stakeholder communication, escalation, coordination, and service restoration.
n
Drive incident resolution within an agreed Service Level Agreements (SLAs) and business expectations.
n
Coordinate bridge calls, war rooms, technical investigations, and executive communications.
n
Review major incidents and ensure comprehensive Post-Incident Reviews (PIRs) are completed.
n
Track corrective and preventive actions resulting from major incidents.
n
Maintain and continuously improve MIM procedures, playbooks, escalation matrices, and communication templates.
n
n
Problem Management (PM)
n
n
Own and manage the end-to-end Problem Management lifecycle.
n
Identify recurring incidents, trends, systemic failures, and operational risks.
n
Lead Root Cause Analysis (RCA) for major and recurring incidents.
n
Drive permanent resolution rather than repeated incident restoration.
n
Work with engineering and technical team to eliminate known errors and recurring failures.
n
Track Problem Management KPIs and demonstrate reduction in recurring incidents.
n
Identify trends and proactively highlight potential service risks before they become major incidents.
n
n
Endurance & Service Resilience Management
n
n
Assess service resilience against operational failures, capacity constraints, technology risks, and recurring incidents.
n
Identify single points of failure and coordinate remediation plans.
n
Drive service improvement initiatives based on incident and problem trends.
n
Ensure critical services have appropriate monitoring, alerting, recovery procedures, support models, and operational documentation.
n
Partner with technology teams to improve service availability, reliability, recoverability, and stability.
n
Ensure lessons learned from incidents are incorporated into service design and operational practices.
n
Support operational readiness for new applications, platforms, infrastructure, and major technology changes.
n
n
Governance & Reporting
n
n
Produce regular operational reports and management dashboards.
n
Mean
TimetoAcknowledge(MTTA)
n
RCAcompletionandquality
n
Ensure compliance with IT Service Management policies, processes, and audit requirements.
n
n
Stakeholder & Vendor Management
n
n
Build strong relationships with technology, business, service management, and senior leadership teams.
n
Provide clear and concise communication during high-pressure situations.
n
Manage escalations across internal teams and external service providers.
n
Hold vendors accountable for incident resolution, RCA quality, SLA performance, and corrective actions.
n
Facilitate cross-functional collaboration to resolve complex and business-critical issues.
n
Act as a trusted point of escalation for service stability and resilience concerns.
n
n
Leadership Responsibilities
n
n
Conduct regular performance and capability reviews.
n
Develop team skills through coaching, knowledge sharing, simulations, and incident exercises.
n
Ensure adequate coverage and on-call/escalation arrangements for critical services.
n
n
Required Experience & Skills
n
n
8–12+ years of experience in IT Service Management, Major Incident Management, Problem Management, or IT Operations.
n
Robust experience leading Major Incident Management for complex enterprise IT environments.
n
Experience with service resilience, availability, reliability, disaster recovery, or operational risk management.
n
Experience working with multiple technology towers such as:
n
Infrastructure
n
Cloud
n
Networks
n
Applications
n
Databases
n
Cybersecurity
n
End-user computing
n
Strong analytical and problem-solving capabilities.
n
Experience working with third-party vendors and managed service providers.
n
n
Preferred Qualifications
n
n
Experience with ITSM platforms such as ServiceNow, BMC Remedy, Jira, or similar.
n
Experience with Agile/DevOps/SRE operating models.
n
Experience with operational resilience, service continuity, or disaster recovery.
n
Experience with automation, observability, monitoring, and service reliability practices.
n
Relevant certifications in Problem Management, SRE, Service Management, or Business Continuity are desirable.
n
📌 Siam Mim & Pm Manager / Endurance Manager (Victoria)
🏢 HCLTech
📍 Victoria