20 Sep
|
HCLTech
|
Melbourne
SIAM MIM & PM Manager / Endurance Manager
WeareHCLTech,oneofthefastest-growinglargetechcompaniesintheworldandhometo219,000+peopleacross54countries,superchargingprogressthroughindustry-leadingcapabilitiescenteredaroundDigital,EngineeringandCloud.
Thedrivingforcebehindthatwork,ourpeople,arediverse,creative,andpassionate,raisingthebarforexcellenceonaregularbasis.We,inturn,workhardtobringoutthebestinthemaswestrivetohelpthemfindtheirsparkandbecomethebestversionofthemselvesthattheycanbe.
Areyoureadytobeanimportantpartofthisever-transformationaljourney?
Job Description
Duration:6MonthsFTCRole
HybridRole:3daysfromofficeismandatory
Position Overview
WearelookingforanexperiencedSIAMMIM&PMManager;/EnduranceManagertoleadMajorIncidentManagement,ProblemManagement,andserviceenduranceactivitiesacrosscriticalITservices.Therolewillberesponsibleforensuringrapidandeffectiverestorationofservicesduringmajorincidents,drivingpermanentresolutionofrecurringissues,andimprovingtheoverallresilience,stability,andoperationalenduranceoftheITenvironment.
ThesuccessfulcandidatewillworkcloselywithServiceManagement,Infrastructure,Applications,Engineering,Security,Cloud,Vendors,andBusinessstakeholderstominimizebusinessimpactandcontinuouslyimproveservicereliability.
Key Responsibilities
Own and lead the end-to-end Major Incident Management process.
Ensure timely incident declaration, stakeholder communication, escalation, coordination, and service restoration.
Drive incident resolution within an agreed Service Level Agreements (SLAs) and business expectations.
Coordinate bridge calls, war rooms, technical investigations, and executive communications.
Review major incidents and ensure comprehensive Post-Incident Reviews (PIRs) are completed.
Track corrective and preventive actions resulting from major incidents.
Maintain and continuously improve MIM procedures, playbooks, escalation matrices, and communication templates.
Problem Management (PM)
Own and manage the end-to-end Problem Management lifecycle.
Identify recurring incidents, trends, systemic failures, and operational risks.
Lead Root Cause Analysis (RCA) for major and recurring incidents.
Drive permanent resolution rather than repeated incident restoration.
Work with engineering and technical team to eliminate known errors and recurring failures.
Track Problem Management KPIs and demonstrate reduction in recurring incidents.
Identify trends and proactively highlight potential service risks before they become major incidents.
Endurance & Service Resilience Management
Assess service resilience against operational failures, capacity constraints, technology risks, and recurring incidents.
Identify single points of failure and coordinate remediation plans.
Drive service improvement initiatives based on incident and problem trends.
Ensure critical services have appropriate monitoring, alerting, recovery procedures, support models, and operational documentation.
Partner with technology teams to improve service availability, reliability, recoverability, and stability.
Ensure lessons learned from incidents are incorporated into service design and operational practices.
Support operational readiness for new applications, platforms, infrastructure, and major technology changes.
Governance & Reporting
Produce regular operational reports and management dashboards.
Mean TimetoAcknowledge(MTTA)
RCAcompletionandquality
Ensure compliance with IT Service Management policies, processes, and audit requirements.
Stakeholder & Vendor Management
Build strong relationships with technology, business, service management, and senior leadership teams.
Provide clear and concise communication during high-pressure situations.
Manage escalations across internal teams and external service providers.
Hold vendors accountable for incident resolution, RCA quality, SLA performance, and corrective actions.
Facilitate cross-functional collaboration to resolve complex and business-critical issues.
Act as a trusted point of escalation for service stability and resilience concerns.
Leadership Responsibilities
Conduct regular performance and capability reviews.
Develop team skills through coaching, knowledge sharing, simulations, and incident exercises.
Ensure adequate coverage and on-call/escalation arrangements for critical services.
Required Experience & Skills
8–12+ years of experience in IT Service Management, Major Incident Management, Problem Management, or IT Operations.
Strong experience leading Major Incident Management for complex enterprise IT environments.
Experience with service resilience, availability, reliability, disaster recovery, or operational risk management.
Experience working with multiple technology towers such as:
Infrastructure
Cloud
Networks
Applications
Databases
Cybersecurity
End-user computing
Robust analytical and problem-solving capabilities.
Experience working with third-party vendors and managed service providers.
Preferred Qualifications
Experience with ITSM platforms such as ServiceNow, BMC Remedy, Jira, or similar.
Experience with Agile/DevOps/SRE operating models.
Experience with operational resilience, service continuity, or disaster recovery.
Experience with automation, observability, monitoring, and service reliability practices.
Relevant certifications in Problem Management, SRE, Service Management, or Business Continuity are desirable.
#J-*****-Ljbffr
📌 Siam Mim & Pm Manager / Endurance Manager (Melbourne)
🏢 HCLTech
📍 Melbourne