25 Aug
|
HCLTech
|
South Australia
25 Aug
HCLTech
South Australia
Job Description
Senior DevOps/ Site Reliability Engineer
n
WeareHCLTech,oneofthefastest-growinglargetechcompaniesintheworldandhometo219,000+peopleacross54countries,superchargingprogressthroughindustry-leadingcapabilitiescenteredaroundDigital,EngineeringandCloud.
n
Thedrivingforcebehindthatwork,ourpeople,arediverse,creative,andpassionate,raisingthebarforexcellenceonaregularbasis.We,inturn,workhardtobringoutthebestinthemaswestrivetohelpthemfindtheirsparkandbecomethebestversionofthemselvesthattheycanbe.
n
Areyoureadytobeanimportantpartofthisever-transformationaljourney?
n
Duration: Perm Role
n
Role Overview
n
We are seeking a highly skilled and hands-on
n
Senior DevOps / Site Reliability Engineer
n
to design, build, and maintain highly available, scalable, and resilient platform services across cloud and on-prem environments. This role requires deep expertise in SRE practices, CI/CD pipelines, infrastructure as code, observability, incident management, and modern cloud-native technologies. The successful candidate will drive reliability improvements, reduce operational toil, and champion engineering excellence across the platform.
n
Key Responsibilities
n
n
Design, implement, and maintain CI/CD pipelines for non-prod and prod environments using Azure DevOps and Git.
n
Define, measure, and maintain Service Level Objectives (SLOs), Service Level Indicators(SLIs), and error budgets for critical platform services.
n
Manage hybrid infrastructure across cloud and on-prem environments with a focus on automation, scalability, and reliability.
n
Drive platform security best practices and compliance across environments.
n
Design and maintain microservices frameworks and container orchestration using Kubernetes and AKS.
n
Develop and maintain infrastructure as code using Terraform and Ansible.
n
Implement and continuously improve monitoring and observability solutions using AppDynamics, Azure Application Insights, Splunk, Splunk Observability Cloud, and Datadog.
n
Lead incident response, perform root cause analysis, and drive post-incident reviews to prevent recurrence.
n
Identify and eliminate operational toil through automation and self-healing systems.
n
Perform capacity planning and proactive performance tuning across platform services.
n
Troubleshoot production issues across AIX and Linux systems, including performance analysis and system diagnostics.
n
Collaborate with development, security, and operations teams to ensure seamless platform integration.
n
Manage and track work using JIRA and ServiceNow.
n
Provide technical leadership in Azure services, including networking, compute, storage, identity, and cost optimisation.
n
Maintain and optimise application servers such as IBM WebSphere Application Server and IBM HTTP Server.
n
Write and maintain scripts in Shell and Python for operational and automation tasks.
n
Participate in on-call rotation and ensure operational readiness for production systems.
n
Contribute to chaos engineering practices and resilience testing to validate system reliability.
n
Mentor junior engineers and contribute to a culture of collaboration, continuous improvement, and blameless learning.
n
n
Required Skills and Experience
n
n
8+ years of experience in DevOps, SRE, platform engineering, or infrastructure roles.
n
Strong hands-on experience with Azure, Kubernetes/AKS, Terraform, Ansible, and CI/CDtools.
n
Deep understanding of SRE principles, including SLOs, SLIs, error budgets, toil reduction, andincident management.
n
Proven experience with incident response, root cause analysis, post-incident reviews, and on-call operations.
n
Strong understanding of cloud-native architecture and hybrid infrastructure.
n
Proficiency in scripting languages such as Shell and Python.
n
Experience with observability and APM tools including AppDynamics, Splunk, SplunkObservability Cloud, Azure Application Insights, and Datadog.
n
Strong understanding of distributed tracing, log analytics, and alerting strategies.
n
Experience implementing alert noise reduction and intelligent routing, including PagerDutyintegration.
n
Familiarity with AIX and Linux systems,
including performance analysis and troubleshooting.
n
Experience with ticketing and ITSM systems such as JIRA and ServiceNow.
n
Strong understanding of application server workflows, including IBM WebSphere ApplicationServer and IBM HTTP Server.
n
Knowledge of networking fundamentals such as DNS, load balancing, CDN, and TLS/SSL.
n
Experience with cloud cost optimisation and governance practices.
n
Excellent communication, documentation, and stakeholder management skills.
n
SRE Focussed
n
Experience with chaos engineering tools and practices such as Azure Chaos Studio, Gremlin,or LitmusChaos.
n
Knowledge of GitOps practices and tools such as ArgoCD or Flux.
n
Experience with secrets management tools such as HashiCorp Vault and Azure Key Vault.
n
Exposure to container security and vulnerability scanning tools such as Trivy, Frogbot, orJFrog Xray.
n
Experience with API gateways and ingress controllers such as Traefik, NGINX, or Azure API Management.
n
Familiarity with CDN and WAF solutions such as Imperva or Akamai.
n
Experience with database reliability, DB2, SQL performance tuning, and connection pool management.
n
Knowledge of automation orchestration tools such as Control-M or Harness.
n
Experience with Open Telemetry and modern telemetry pipelines.
n
Familiarity with FinOps practices and Azure cost governance tooling.
n
Experience contributing to platform migration projects such as ingress controller migrations or monitoring platform consolidation.
n
n
Preferred Certifications
n
n
Azure Solutions Architect Expert
n
Google Cloud Skilled SRE or equivalent SRE certification
n
HashiCorp Terraform Associate
n
ITIL Foundation
n
n
Representing165nationalitiesacrosstheglobe,weprideourselvesonbeinganequalopportunityemployer,committedtoprovidequalemploymentopportunitiestoallapplicantsandemployeesregardlessofrace,religion,sex,color,age,nationalorigin,pregnancy,sexualorientation,physicaldisabilityorgeneticinformation,militaryorveteranstatus,AboriginalandTorresStraitIslanderpeopleoranyotherprotectedclassification,inaccordancewithfederal,state,and/orlocallaw.
n
CandidateDataPrivacyNotice|HCLTechnologies
n
:...
#J-*****-Ljbffr
📌 Senior Devops/ Site Reliability Engineer (South Australia)
🏢 HCLTech
📍 South Australia