21 Sep
|
HCLTech
|
Melbourne
SIAM MIM & PM Manager / Endurance Manager
WeareHCLTech,oneofthefastest-growinglargetechcompaniesintheworldandhometo219,000+peopleacross54countries,superchargingprogressthroughindustry-leadingcapabilitiescenteredaroundDigital,EngineeringandCloud.
Thedrivingforcebehindthatwork,ourpeople,arediverse,creative,andpassionate,raisingthebarforexcellenceonaregularbasis.We,inturn,workhardtobringoutthebestinthemaswestrivetohelpthemfindtheirsparkandbecomethebestversionofthemselvesthattheycanbe.
Areyoureadytobeanimportantpartofthisever-transformationaljourney?
Job Description
Duration:6MonthsFTCRole
HybridRole:3daysfromofficeismandatory
Position Overview
WearelookingforanexperiencedSIAMMIM&PMManager;/EnduranceManagertoleadMajorIncidentManagement,ProblemManagement,andserviceenduranceactivitiesacrosscriticalITservices.Therolewillberesponsibleforensuringrapidandeffectiverestorationofservicesduringmajorincidents,drivingpermanentresolutionofrecurringissues,andimprovingtheoverallresilience,stability,andoperationalenduranceoftheITenvironment.
ThesuccessfulcandidatewillworkcloselywithServiceManagement,Infrastructure,Applications,Engineering,Security,Cloud,Vendors,andBusinessstakeholderstominimizebusinessimpactandcontinuouslyimproveservicereliability.
Key Responsibilities
- Own and lead the end-to-end Major Incident Management process.
- Ensure timely incident declaration, stakeholder communication, escalation, coordination, and service restoration.
- Drive incident resolution within an agreed Service Level Agreements (SLAs) and business expectations.
- Coordinate bridge calls, war rooms, technical investigations, and executive communications.
- Review major incidents and ensure comprehensive Post-Incident Reviews (PIRs) are completed.
- Track corrective and preventive actions resulting from major incidents.
- Maintain and continuously improve MIM procedures, playbooks, escalation matrices, and communication templates.
Problem Management (PM)
- Own and manage the end-to-end Problem Management lifecycle.
- Identify recurring incidents, trends, systemic failures, and operational risks.
- Lead Root Cause Analysis (RCA) for major and recurring incidents.
- Drive permanent resolution rather than repeated incident restoration.
- Work with engineering and technical team to eliminate known errors and recurring failures.
- Track Problem Management KPIs and demonstrate reduction in recurring incidents.
- Identify trends and proactively highlight potential service risks before they become major incidents.
Endurance & Service Resilience Management
- Assess service resilience against operational failures, capacity constraints, technology risks, and recurring incidents.
- Identify single points of failure and coordinate remediation plans.
- Drive service improvement initiatives based on incident and problem trends.
- Ensure critical services have appropriate monitoring, alerting, recovery procedures, support models, and operational documentation.
- Partner with technology teams to improve service availability, reliability, recoverability, and stability.
- Ensure lessons learned from incidents are incorporated into service design and operational practices.
- Support operational readiness for new applications, platforms, infrastructure, and major technology changes.
Governance & Reporting
- Produce regular operational reports and management dashboards.
- Mean TimetoAcknowledge(MTTA)
- RCAcompletionandquality
- Ensure compliance with IT Service Management policies, processes, and audit requirements.
Stakeholder & Vendor Management
- Build strong relationships with technology, business, service management, and senior leadership teams.
- Provide clear and concise communication during high-pressure situations.
- Manage escalations across internal teams and external service providers.
- Hold vendors accountable for incident resolution, RCA quality, SLA performance, and corrective actions.
- Facilitate cross-functional collaboration to resolve complex and business-critical issues.
- Act as a trusted point of escalation for service stability and resilience concerns.
Leadership Responsibilities
- Conduct regular performance and capability reviews.
- Develop team skills through coaching, knowledge sharing, simulations, and incident exercises.
- Ensure adequate coverage and on-call/escalation arrangements for critical services.
Required Experience & Skills
- 8–12+ years of experience in IT Service Management, Major Incident Management, Problem Management, or IT Operations.
- Strong experience leading Major Incident Management for complex enterprise IT environments.
- Experience with service resilience, availability, reliability, disaster recovery, or operational risk management.
- Experience working with multiple technology towers such as:
- Infrastructure
- Cloud
- Networks
- Applications
- Databases
- Cybersecurity
- End-user computing
- Solid analytical and problem-solving capabilities.
- Experience working with third-party vendors and managed service providers.
Preferred Qualifications
- Experience with ITSM platforms such as ServiceNow, BMC Remedy, Jira, or similar.
- Experience with Agile/DevOps/SRE operating models.
- Experience with operational resilience, service continuity, or disaster recovery.
- Experience with automation, observability, monitoring, and service reliability practices.
- Relevant certifications in Problem Management, SRE, Service Management, or Business Continuity are desirable.
#J-18808-Ljbffr
📌 SIAM MIM & PM Manager / Endurance Manager (Melbourne)
🏢 HCLTech
📍 Melbourne