21 Aug
|
HCLTech
|
Williamstown
21 Aug
HCLTech
Williamstown
Senior DevOps/ Site Reliability Engineer
WeareHCLTech,oneofthefastest-growinglargetechcompaniesintheworldandhometo219,000+peopleacross54countries,superchargingprogressthroughindustry-leadingcapabilitiescenteredaroundDigital,EngineeringandCloud.
Thedrivingforcebehindthatwork,ourpeople,arediverse,creative,andpassionate,raisingthebarforexcellenceonaregularbasis.We,inturn,workhardtobringoutthebestinthemaswestrivetohelpthemfindtheirsparkandbecomethebestversionofthemselvesthattheycanbe.
Areyoureadytobeanimportantpartofthisever-transformationaljourney?
Duration: Perm Role
Role Overview
We are seeking a highly skilled and hands-on
Senior DevOps / Site Reliability Engineer
to design, build, and maintain highly available, scalable, and resilient platform services across cloud and on-prem environments. This role requires deep expertise in SRE practices, CI/CD pipelines, infrastructure as code, observability, incident management, and modern cloud-native technologies. The successful candidate will drive reliability improvements, reduce operational toil, and champion engineering excellence across the platform.
Key Responsibilities
- Design, implement, and maintain CI/CD pipelines for non-prod and prod environments using Azure DevOps and Git.
- Define, measure, and maintain Service Level Objectives (SLOs), Service Level Indicators(SLIs), and error budgets for critical platform services.
- Manage hybrid infrastructure across cloud and on-prem environments with a focus on automation, scalability, and reliability.
- Drive platform security best practices and compliance across environments.
- Design and maintain microservices frameworks and container orchestration using Kubernetes and AKS.
- Develop and maintain infrastructure as code using Terraform and Ansible.
- Implement and continuously improve monitoring and observability solutions using AppDynamics, Azure Application Insights, Splunk, Splunk Observability Cloud, and Datadog.
- Lead incident response, perform root cause analysis, and drive post-incident reviews to prevent recurrence.
- Identify and eliminate operational toil through automation and self-healing systems.
- Perform capacity planning and proactive performance tuning across platform services.
- Troubleshoot production issues across AIX and Linux systems,
including performance analysis and system diagnostics.
- Collaborate with development, security, and operations teams to ensure seamless platform integration.
- Manage and track work using JIRA and ServiceNow.
- Provide technical leadership in Azure services, including networking, compute, storage, identity, and cost optimisation.
- Maintain and optimise application servers such as IBM WebSphere Application Server and IBM HTTP Server.
- Write and maintain scripts in Shell and Python for operational and automation tasks.
- Participate in on-call rotation and ensure operational readiness for production systems.
- Contribute to chaos engineering practices and resilience testing to validate system reliability.
- Mentor junior engineers and contribute to a culture of collaboration, continuous improvement, and blameless learning.
Required Skills and Experience
- 8+ years of experience in DevOps, SRE, platform engineering, or infrastructure roles.
- Strong hands-on experience with Azure, Kubernetes/AKS, Terraform, Ansible, and CI/CDtools.
- Deep understanding of SRE principles, including SLOs, SLIs, error budgets, toil reduction, andincident management.
- Proven experience with incident response, root cause analysis, post-incident reviews, and on-call operations.
- Solid understanding of cloud-native architecture and hybrid infrastructure.
- Proficiency in scripting languages such as Shell and Python.
- Experience with observability and APM tools including AppDynamics, Splunk, SplunkObservability Cloud, Azure Application Insights, and Datadog.
- Strong understanding of distributed tracing, log analytics, and alerting strategies.
- Experience implementing alert noise reduction and intelligent routing, including PagerDutyintegration.
- Familiarity with AIX and Linux systems, including performance analysis and troubleshooting.
- Experience with ticketing and ITSM systems such as JIRA and ServiceNow.
- Strong understanding of application server workflows, including IBM WebSphere ApplicationServer and IBM HTTP Server.
- Knowledge of networking fundamentals such as DNS, load balancing, CDN, and TLS/SSL.
- Experience with cloud cost optimisation and governance practices.
- Excellent communication, documentation, and stakeholder management skills.
- SRE Focussed
- Experience with chaos engineering tools and practices such as Azure Chaos Studio, Gremlin,or LitmusChaos.
- Knowledge of GitOps practices and tools such as ArgoCD or Flux.
- Experience with secrets management tools such as HashiCorp Vault and Azure Key Vault.
- Exposure to container security and vulnerability scanning tools such as Trivy, Frogbot, orJFrog Xray.
- Experience with API gateways and ingress controllers such as Traefik, NGINX, or Azure API Management.
- Familiarity with CDN and WAF solutions such as Imperva or Akamai.
- Experience with database reliability, DB2, SQL performance tuning, and connection pool management.
- Knowledge of automation orchestration tools such as Control-M or Harness.
- Experience with Open Telemetry and modern telemetry pipelines.
- Familiarity with FinOps practices and Azure cost governance tooling.
- Experience contributing to platform migration projects such as ingress controller migrations or monitoring platform consolidation.
Preferred Certifications
- Azure Solutions Architect Expert
- Google Cloud Professional SRE or equivalent SRE certification
- HashiCorp Terraform Associate
- ITIL Foundation
Representing165nationalitiesacrosstheglobe,weprideourselvesonbeinganequalopportunityemployer,committedtoprovidequalemploymentopportunitiestoallapplicantsandemployeesregardlessofrace,religion,sex,color,age,nationalorigin,pregnancy,sexualorientation,physicaldisabilityorgeneticinformation,militaryorveteranstatus,AboriginalandTorresStraitIslanderpeopleoranyotherprotectedclassification,inaccordancewithfederal,state,and/orlocallaw.
CandidateDataPrivacyNotice|HCLTechnologies
Wearecommittedtorespectingyourprivacyandfortheprotectionofyourpersonaldata.Yourpersonaldatawillbecollectedandprocessedinlinewithourcandidateprivacynotice:https://www.hcltech.com/candidate-privacy-notice.Thisprivacynoticewillhelpyout…
#J-18808-Ljbffr
📌 Senior DevOps/ Site Reliability Engineer (Williamstown)
🏢 HCLTech
📍 Williamstown