20 Aug
|
Arista Networks
|
New South Wales
20 Aug
Arista Networks
New South Wales
Job Description
Site Reliability Engineer (SRE/DevOps) – Engineering Productivity – Sydney
n
Join Arista Networks' Engineering Productivity (EngProd) team to design, build, and operate secure, scalable, and fault‐tolerant infrastructure in a hybrid cloud environment.
n
What You'll Do
n
n
Build, deploy safely and incrementally and operate critical production systems with focus on scalability, reliability, observability, performance and security.
n
Monitor, support and enhance developer experience across services.
n
Build automation to remove toil and efficiently operate production systems.
n
Proactively monitor, respond to, and enhance alerts and set up automated alert handling.
n
Create and maintain incident response runbooks.
n
Triage
platform/infrastructure
issues and help Arista software engineers; engage with 3rd‐party vendor support.
n
Write post‐mortem documents and build solutions to prevent incident recurrence.
n
Plan and communicate maintenance windows on production systems.
n
Work with product development teams to identify and resolve infrastructural bottlenecks.
n
Survey and adopt best practices around
infrastructure/platform
for secure, scalable, fault‐tolerant systems.
n
Study the design and implementation details of OSS systems for better triage and fix resolution.
n
n
Qualifications
n
n
At least BSc in Computer Science or Engineering + 3 years of experience, or equivalent.
n
Knowledge of Go, Python,
or shell scripting for automation workflows.
n
Experience with Linux (UNIX) administration and debugging.
n
Hands‐on experience operating infrastructure at scale.
n
Server provisioning experience, especially with storage and networking.
n
Strong problem‐solving and software troubleshooting skills.
n
Experience with infrastructure‐as‐code (e.g., Ansible).
n
n
Desired Skills
n
n
Managing databases (mariadb, postgres, mongodb).
n
Docker and virtualization (kvm, qemu, kata‐containers).
n
Monitoring stack (Prometheus, Loki, Tempo, InfluxDB, Grafana, Thanos).
n
ElasticSearch cluster management.
n
Artifactory, docker registry management.
n
CI/CD systems (ArgoCD, Spinnaker).
n
Version control (Perforce, Gerrit).
n
Infrastructure‐as‐code frameworks (Ansible).
n
Large Java application management.
n
Storage infrastructure (NAS, SAN, Ceph).
n
n
Additional Information
n
Please note: We are not engaging external recruiters for this role. Only direct applications will be considered.
n
Australian Work Rights
n
Only candidates with Australian Citizenship, Australian Permanent Residency, or another demonstrable legal entitlement to work in Australia for the duration of employment, will be considered for this role.
n
Employment Details
n
Location: Sydney, Current South Wales, Australia
n
Type: Full‐time
n
Salary: A$150,****** – A$170,******
#J-*****-Ljbffr
📌 Site Reliability Engineer (Sre/ Devops) - Engineering Productivity - Sydney (New South Wales)
🏢 Arista Networks
📍 New South Wales