Principal Software Engineer

Overview
The Azure Kubernetes Service AKS team is building a world-class managed Kubernetes platform for Linux and Windows workloads across cloud and edge environments. Our mission is to define the next generation of secure, reliable, and scalable cloud-native infrastructure on Azure.
We are seeking a Principal Software Engineer to lead the architecture and evolution of the AKS control plane. You will solve complex distributed-systems challenges across Kubernetes API servers, etcd, scheduling, controllers, networking, and control-plane lifecycle management. Your work will improve availability, performance, scalability, security, and operational resilience for millions of clusters and some of the world’s largest production and AI workloads.

You will drive designs that eliminate single points of failure, reduce customer-impacting incidents, improve failure isolation and automated recovery, and enable control planes to scale safely under rapid growth and demanding workloads. You will establish reliability engineering practices using meaningful SLIs and SLOs, capacity modeling, fault injection, observability, and data-driven incident prevention. You will also lead cross-team technical initiatives, influence upstream Kubernetes, and guide engineers through architecture, implementation, rollout, and production operations.

As AI transforms infrastructure requirements, AKS is enabling increasingly large GPU and accelerator fleets, high-throughput training and inference workloads, and AI-native applications. You will help ensure that the AKS control plane remains predictable, efficient, and resilient as these workloads push Kubernetes to new limits of scale.
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Responsibilities
- Define and drive the architecture of highly available, secure, and scalable AKS control-plane systems.
- Advance Kubernetes and critical control-plane components—including API server, etcd, and controllers.
- Improve reliability through clear SLIs and SLOs, capacity planning, performance engineering, fault testing, and production telemetry.
- Design automated detection, mitigation, recovery, and safe rollout mechanisms that prevent failures from becoming customer-impacting incidents.
- Lead investigations of complex distributed-system failures and translate findings into durable platform improvements.
- Improve engineering velocity through reusable frameworks, testing infrastructure, deployment automation, and operational tooling.
- Balance long-term architectural vision with pragmatic, incremental delivery and measurable customer outcomes.
- Provide technical leadership across teams, mentor enginee




Looking for a job?

Visit Careersaas to find millions of new and unfilled roles, including thousands of remote and hybrid positions.