24 Sep
|
Sharon AI
|
Sydney
?️ Full-time | Permanent | Start ASAP
? Sydney
? Hybrid working
About Sharon AI
Sharon AI is an Australian neocloud, delivering trusted AI infrastructure organisations need to build, train and run AI at scale. We support customers across the full AI lifecycle, from training through to inference and agentic AI, drawing on a strong ecosystem of technology and co-location partners to deliver capability where it's needed. As the first neocloud to join NVIDIA's AI Compute Program, we're growing quickly, scaling our AI Factory platform to meet rising demand for advanced compute.
The Role
As Sharon AI continues to grow, we're looking for a talented Incident & Problem Manager to join our NOC Service Management team and lead the day-to-day execution of Incident Management, Major Incident Management, Problem Management and Service Level Management across our Australian and global AI Factory and Data Centre environments, reporting to our NOC Manager.
In this role, the Incident & Problem Manager will act as the primary focal point for major and ongoing incidents, coordinating customer teams, the NOC Service Desk, L2/L3 technical support, Engineering, Build, Service Operations, Data Centre Operations, SOC, Customer Success and technology vendors through to service restoration and closure. You'll own Post-Incident Reviews, Root Cause Analysis and SLA performance, working closely with our Data Analytics & Reporting team, and participate in an on-call roster to support after-hours P1 and escalated P2 incidents.
What You'll Be Doing
- Own and execute end-to-end Incident and Major Incident Management, ensuring incidents are classified, prioritised, escalated and driven through to service restoration and closure
- Maintain strong governance across technical bridges, assessing customer, service and business impact and rapidly engaging L2/L3, Engineering, Build, Service Operations, Data Centre and vendor teams
- Lead P1/P2 incident coordination, establishing ownership, restoration priorities and communication cadence for leadership, stakeholders and customers
- Coordinate and govern Root Cause Analysis and Post-Incident Reviews, ensuring corrective and preventive actions are documented, assigned and tracked to resolution
- Perform detailed analysis of incidents, root causes and recurring issues, applying Lean Six Sigma and structured problem-solving techniques such as DMAIC, 5 Whys, Fishbone and Pareto analysis
- Own operational governance of SLA and OLA commitments, proactively monitoring performance and driving corrective actions ahead of breach risk
- Maintain and continuously improve Incident, Major Incident, Problem and Service Level Management processes, playbooks and escalation matrices, including periodic ticket quality audits
- Use incident, problem and SLA insights to identify recurring risks and improvement opportunities, sharing lessons learned across L1/L2/L3 Operations
What We're Looking For
- 5+ years' relevant experience in Service Assurance, Incident/Major Incident Management, Problem Management and Service Level Management within complex, business-critical infrastructure and operational environments
- Proven experience leading high-priority incidents and coordinating technical teams, stakeholders, escalations and communications under SLA-driven conditions
- Strong experience in Problem Management, RCA/PIR, corrective actions, known errors and ITIL-based operational governance
- Ability to lead and coordinate major incidents, make timely decisions and drive actions and communications with wider teams operating across multiple time zones
- Strong technical understanding and analytical problem-solving ability, with strong attention to detail and accuracy
- Customer-first mindset with a clear focus on service availability, SLA performance and customer outcomes
- Robust stakeholder communication, collaboration and active listening skills across customers, leadership and technical teams
- Ability to remain calm, focused, composed and effective in a fast-paced, complex and highly technical environment
- ITIL Foundation certification
Nice to have:
- Experience within Data Centre, Cloud, Infrastructure, Network or AI/HPC operational environments
- Lean Six Sigma or other relevant Service Management / Quality certifications
- Tertiary qualification in Information Technology, Engineering, Telecommunications, Business or a related discipline
Why Join Sharon AI? You'll be joining a rapidly growing Australian technology business at an exciting stage of its journey, with the opportunity to work directly with the infrastructure and technology powering the next generation of AI.
? Hybrid working – flexibility between our office and working from home
? Birthday leave – take some extra time to celebrate your day
? Employee Assistance Program (EAP) – confidential support when you need it
? Exposure to AI and next-generation technology – work in one of the fastest-moving areas of technology
? Learning & development – we support your career growth with approved conferences, professional memberships & courses
? Novated leasing – a tax-effective way to finance and run your car
? Bounty referral program – generous rewards for successfully referring new talent to Sharon AI
? Employee of the month – recognition plus a $500 gift card
? Growing global business – be part of an Australian technology company with an expanding international footprint
Our Values
Integrity | Innovation | Collaboration | Wellbeing | Inclusion
Apply today and help us build the infrastructure powering the next generation of AI.
Due to the high volume of applications we receive, we're unfortunately not always able to provide individual feedback to unsuccessful candidates. We appreciate your understanding and want to assure you that every application will be reviewed with care.
📌 Incident & Problem Manager - AI Operations (Sydney)
🏢 Sharon AI
📍 Sydney