10 Sep
|
HCL Australia Services
|
Hawthorn East
10 Sep
HCL Australia Services
Hawthorn East
We are HCLTech, one of the fastest-growing large tech companies in the world and home to 219,000+ people across 54 countries, supercharging progress through industry-leading capabilities centered around Digital, Engineering and Cloud.
The driving force behind that work, our people, are diverse, creative, and passionate, raising the bar for excellence on a regular basis. We, in turn, work hard to bring out the best in them as we strive to help them find their spark and become the best version of themselves that they can be.
At HCLTech Australia, we value the unique perspective and contributions of all individual and we actively encourage applications from Aboriginal and Torres Strait Islander people to apply for this role.
Are you ready to be an important part of this ever-transformational journey?
Job Description
Duration: 6 Months FTC Role
Hybrid Role: 3 days from office is mandatory
Position Overview
We are looking for an experienced SIAM MIM & PM Manager / Endurance Manager to lead Major Incident Management, Problem Management, and service endurance activities across critical IT services. The role will be responsible for ensuring rapid and effective restoration of services during major incidents, driving permanent resolution of recurring issues, and improving the overall resilience, stability, and operational endurance of the IT workplace.
The successful candidate will work closely with Service Management, Infrastructure, Applications, Engineering, Security, Cloud, Vendors, and Business stakeholders to minimize business impact and continuously improve service reliability.
Key Responsibilities
1. Major Incident Management (MIM)
· Own and lead the end-to-end Major Incident Management process.
· Act as the Major Incident Manager during high-severity and business-critical incidents.
· Ensure timely incident declaration, stakeholder communication, escalation, coordination, and service restoration.
· Establish clear command and control during major incidents and ensure appropriate technical teams are engaged.
· Drive incident resolution within agreed Service Level Agreements (SLAs) and business expectations.
· Coordinate bridge calls, war rooms, technical investigations, and executive communications.
· Ensure accurate and timely communication to business stakeholders, leadership, and customers.
· Review major incidents and ensure comprehensive Post-Incident Reviews (PIRs) are completed.
· Track corrective and preventive actions resulting from major incidents.
· Identify opportunities to improve incident response processes, automation, monitoring, and operational readiness.
· Maintain and continuously improve MIM procedures, playbooks, escalation matrices, and communication templates.
2. Problem Management (PM)
· Own and manage the end-to-end Problem Management lifecycle.
· Identify recurring incidents, trends, systemic failures, and operational risks.
· Lead Root Cause Analysis (RCA) for major and recurring incidents.
· Ensure corrective and preventive actions are identified, assigned, tracked, and closed.
· Drive permanent resolution rather than repeated incident restoration.
· Maintain the Problem Management backlog and prioritize problems based on business impact and risk.
· Work with engineering and technical teams to eliminate known errors and recurring failures.
· Challenge incomplete or low-quality RCA and ensure evidence-based root cause identification.
· Track Problem Management KPIs and demonstrate reduction in recurring incidents.
· Identify trends and proactively highlight potential service risks before they become major incidents.
3. Endurance & Service Resilience Management
· Own the endurance and operational resilience agenda for critical IT services.
· Assess service resilience against operational failures, capacity constraints, technology risks, and recurring incidents.
· Identify single points of failure and coordinate remediation plans.
· Drive service improvement initiatives based on incident and problem trends.
· Ensure critical services have appropriate monitoring, alerting, recovery procedures, support models, and operational documentation.
· Partner with technology teams to improve service availability, reliability, recoverability, and stability.
· Establish and maintain service endurance metrics and dashboards.
· Conduct regular service health and resilience reviews for critical services.
· Ensure lessons learned from incidents are incorporated into service design and operational practices.
· Support operational readiness for recent applications, platforms, infrastructure, and major technology changes.
4. Governance & Reporting
· Develop and maintain MIM, Problem Management, and Endurance governance frameworks.
· Produce regular operational reports and management dashboards.
· Track KPIs/KRIs including:
o Major Incident volume and severity
o Mean Time to Acknowledge (MTTA)
o Mean Time to Restore/Resolve (MTTR)
o Repeat/recurring incidents
o Problem backlog and aging
o RCA completion and quality
o Corrective action closure
o Service availability and stability
o Operational risk and resilience indicators
· Present key trends, risks, improvement opportunities, and action plans to senior management.
· Ensure compliance with IT Service Management policies, processes, and audit requirements.
5. Stakeholder & Vendor Management
· Build solid relationships with technology, business, service management, and senior leadership teams.
· Provide clear and concise communication during high-pressure situations.
· Manage escalations across internal teams and external service providers.
· Hold vendors accountable for incident resolution, RCA quality, SLA performance, and corrective actions.
· Facilitate cross-functional collaboration to resolve complex and business-critical issues.
· Act as a trusted point of escalation for service stability and resilience concerns.
Leadership Responsibilities
· Lead and mentor MIM, Problem Management, and operational resilience teams.
· Establish explicit roles, responsibilities, and escalation paths.
· Build a culture of accountability, continuous improvement, and proactive risk management.
· Conduct regular performance and capability reviews.
· Develop team skills through coaching, knowledge sharing, simulations, and incident exercises.
· Ensure adequate coverage and on-call/escalation arrangements for critical services.
Required Experience & Skills
· 8–12+ years of experience in IT Service Management, Major Incident Management, Problem Management, or IT Operations.
· Strong experience leading Major Incident Management for complex enterprise IT environments.
· Proven experience in Problem Management and Root Cause Analysis.
· Experience managing critical production environments and high-impact business services.
· Strong understanding of ITIL/ITSM processes and service management frameworks.
· Experience with service resilience, availability, reliability, disaster recovery, or operational risk management.
· Robust stakeholder management and executive communication skills.
· Experience working with multiple technology towers such as:
o Infrastructure
o Cloud
o Networks
o Applications
o Databases
o Cybersecurity
o End-user computing
· Strong analytical and problem-solving capabilities.
· Experience working with third-party vendors and managed service providers.
· Ability to remain calm and provide effective leadership during high-pressure incidents.
· Strong reporting, governance, and presentation skills.
Preferred Qualifications
· ITIL 4 Foundation or higher.
· Experience with ITSM platforms such as ServiceNow, BMC Remedy, Jira, or similar.
· Experience with Agile/DevOps/SRE operating models.
· Knowledge of cloud platforms such as AWS, Azure, or GCP.
· Experience with operational resilience, service continuity, or disaster recovery.
· Experience with automation, observability, monitoring, and service reliability practices.
· Relevant certifications in Problem Management, SRE, Service Management, or Business Continuity are desirable.
Key Competencies
· Major Incident Leadership
· Problem Solving & Root Cause Analysis
· Operational Resilience
· Service Reliability
· Risk Management
· Stakeholder Management
· Executive Communication
· Vendor Management
· Data & Trend Analysis
· Continuous Improvement
· Decision Making Under Pressure
· Team Leadership
📌 SIAM MIM & PM Manager / Endurance Manager (Hawthorn East)
🏢 HCL Australia Services
📍 Hawthorn East