Senior Production Reliability Engineer – Cloud Infrastructure & Automation
Westbury Partners Sydney, AustraliaSenior Production Reliability Engineer – Cloud Infrastructure & Automation
Westbury Partners Sydney, Australia
Drive reliability across critical production systems by engineering scalable cloud infrastructure, automating operations, strengthening observability, leading incident response, and enabling engineering teams to deliver safer, resilient services.
What You'll Do:
- Design, operate, and continuously improve highly available, scalable, secure production infrastructure.
- Build automation for provisioning, deployments, configuration, and operational workflows.
- Establish meaningful SLOs, SLIs, reliability metrics, and error budgets.
- Strengthen monitoring, logging, alerting, tracing, and overall observability.
- Troubleshoot complex production issues and participate in on-call incident response.
- Lead post-incident reviews and turn recurring problems into lasting improvements.
- Improve deployment safety, rollback strategies, change management, and release processes.
- Identify infrastructure bottlenecks, reliability risks, technical debt, and opportunities for optimisation.
- Support capacity planning, disaster recovery, performance testing, and resilience initiatives.
- Create infrastructure-as-code, runbooks, documentation, and operational best practices.
Your responsibilities will include:
- You’ll take ownership of critical production environments while partnering with software, security, and infrastructure teams. You’ll automate repetitive operational work, improve system resilience, reduce toil, and help engineering teams adopt reliability-focused practices.
- You’ll work across cloud platforms, Linux, containers, Kubernetes, infrastructure-as-code, CI/CD, observability, networking, databases, and distributed systems. Your work will directly contribute to improved availability, faster incident recovery, safer deployments, scalability, and infrastructure efficiency.
Why Join Us:
- This is an opportunity to solve challenging production infrastructure problems while having a direct influence on how systems are designed, deployed, monitored, and operated.
- You’ll be empowered to build meaningful automation, improve engineering practices, strengthen reliability, and create tools that make life easier for development teams while delivering a better experience for customers.
About You:
- You’re an experienced reliability, infrastructure, DevOps, or production engineer with strong Linux and cloud expertise. You’re comfortable writing code or scripts in Python, Go, Bash, or similar languages and have hands-on experience with technologies such as Kubernetes, Docker, Terraform, and modern CI/CD platforms.
- You understand networking, DNS, HTTP/TLS, load balancing, databases, distributed systems, monitoring, and observability. You’re analytical, collaborative, calm during incidents, and naturally driven to automate problems rather than repeatedly work around them.
- You take ownership, communicate clearly, enjoy solving complex technical challenges, and continuously look for ways to make production systems safer, simpler, and more reliable.
#SiteReliabilityEngineering
#SRE
#DevOps
#CloudInfrastructure
#Kubernetes
#Terraform
#InfrastructureAsCode
#CloudEngineering
#Observability
#ProductionEngineering
#Automation
#IncidentResponse
#PlatformEngineering
#ReliabilityEngineering
#CloudNative
Job ID 8646
More Jobs From Westbury Partners
Westbury Partners
Sydney, Australia
Westbury Partners
Sydney, Australia
Westbury Partners
Singapore
Westbury Partners
Sydney, Australia
Westbury Partners
Sydney, Australia
Westbury Partners
Sydney, Australia
Westbury Partners
Sydney, Australia
Westbury Partners
Sydney, Australia
Westbury Partners
Sydney, Australia
Boost your career
Find thousands of job opportunities by signing up to eFinancialCareers today.More Jobs Like This
Morgan Stanley
Tokyo, Japan
Northern Trust
Pune, India
CoreWeave
Northfield, United States
JPMorgan Chase & Co.
London, United Kingdom