21 Aug
|
Synechron
|
Melbourne
21 Aug
Synechron
Melbourne
We are seeking an experienced Platform Engineer to build, operate, and continuously improve cloud-native platforms hosted on AWS. The ideal candidate will have strong expertise in AWS infrastructure, containerized workloads, CI/CD automation, observability, and production support.
This role requires hands-on experience managing Amazon ECS environments, troubleshooting complex production incidents, implementing Infrastructure as Code (IaC), and driving platform reliability, scalability, and automation initiatives. The successful candidate will work closely with engineering teams to ensure high availability, performance, security, and operational excellence across enterprise platforms.
Key Responsibilities
- Design, implement, and maintain highly available, scalable, and secure AWS-based platform infrastructure.
- Manage and support containerized applications running on Amazon ECS (EC2 and/or Fargate).
- Monitor platform health and respond to production incidents, ensuring minimal downtime and rapid resolution.
- Troubleshoot application performance issues, latency bottlenecks, infrastructure failures, and deployment-related incidents.
- Optimize ECS clusters, Auto Scaling Groups, load balancers, and cloud resources for performance and cost efficiency.
- Develop and maintain Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar technologies.
- Build and enhance CI/CD pipelines to support automated deployment, testing, and release processes.
- Configure and maintain monitoring, logging, alerting, and observability platforms to improve operational visibility.
- Perform Root Cause Analysis (RCA) and implement preventative measures to improve platform stability.
- Collaborate with software engineering teams to enhance application reliability, deployment practices, and operational readiness.
- Implement cloud security, governance, compliance, and operational best practices.
- Drive platform automation initiatives to reduce manual intervention and improve operational efficiency.
Required Skills & Experience
- 5+ years of experience as a Platform Engineer, DevOps Engineer, Cloud Engineer, or Site Reliability Engineer (SRE).
- Strong hands-on experience with AWS services including:
- Amazon ECS
- EC2
- CloudWatch
- Auto Scaling
- IAM
- VPC and Networking
Containers & Orchestration
- Strong experience managing containerized applications using:
- Docker
- Amazon ECS
Production Support & Troubleshooting
- Proven experience investigating and resolving production incidents related to:
- High latency
- Application crashes
- Memory leaks
- Resource bottlenecks
- Capacity and scaling issues
- Strong understanding of:
- Application performance tuning
- Resource optimization
- Container performance management
Infrastructure & Automation
- Expertise with Infrastructure as Code (IaC):
- Terraform
- CloudFormation
- Robust scripting and automation skills using:
- Python
- Bash
- PowerShell
DevOps & CI/CD
- Experience implementing and supporting CI/CD pipelines using:GitHub Actions
- Jenkins
Monitoring & Observability
- Hands-on experience with monitoring and observability platforms including:
- CloudWatch
- Grafana
- Splunk
- ELK Stack
Desired Experience
- Strong experience supporting production ECS workloads and container lifecycle management.
- Knowledge of Out-of-Memory (OOM) analysis, JVM tuning, and container resource optimization.
- Experience implementing Blue-Green and Canary deployment strategies.
- Exposure to Site Reliability Engineering (SRE) principles and operational best practices.
- Experience implementing autoscaling strategies using CPU, memory, throughput, and application-level metrics.
- Understanding of FinOps practices and cloud cost optimization.
- Experience with incident management, post-incident reviews, and operational excellence frameworks.
#J-18808-Ljbffr
📌 Platform Engineer (AWS, ECS & DevOps) (Melbourne)
🏢 Synechron
📍 Melbourne