You will design and maintain CI/CD pipelines for both infrastructure and application deployment, automate build, test, release, and rollback, and integrate security scans, quality gates, and infrastructure validation into those pipelines. You will implement monitoring, alerting, and logging, participate in incident response and root cause analysis, and drive the post-incident improvements that actually stop the same thing happening twice. You will work closely with Software Engineering, QA, Security, and Architecture as a shared owner of production rather than a downstream service desk.
- 4+ years of hands-on DevOps or platform engineering
- Strong Azure: AKS, virtual networks, load balancers, storage accounts, Azure Monitor, Application Insights, IAM
- Terraform building and maintaining your own modules, not only running someone else's
- Kubernetes in production, ideally AKS: deployments, scaling, rolling upgrades, resiliency, and real troubleshooting
- CI/CD pipeline design and ownership, including automated rollback
- Monitoring, alerting, and logging, plus genuine incident response experience
- Willingness to participate in an on-call rotation
- Collaborative instincts across engineering, QA, security, and architecture