Software Engineer II

Microsoft — Redmond, WA, US

Apply on employer website
Overview
The AI Infrastructure team is responsible for building and operating the large-scale, reliable, and efficient GPU-based clustering infrastructure that powers Microsoft’s AI/ML ecosystem. We host the training and inference platforms behind many of Microsoft’s flagship AI offerings, including Azure OpenAI Service, M365 Copilot & Copilot Tuning, GitHub Copilot, Azure AI Foundry’s inference and fine-tuning services for both OpenAI and open-source models; as well as the mature Azure ML Services, which provide data scientists and developers a rich experience for defining, training, fine-tuning, deploying, monitoring, and consuming machine learning models. Our infrastructure enables AI innovation at hyperscale and supports some of the most demanding workloads and business groups across the company. We engage directly with some of the major internal research and applied AI/ML groups using these services, including Microsoft Research, M365, Microsoft Security, and the Bing WebXT team. The AI Infra team is looking for a talented Software Engineer II, with initial focus on the Scheduler subsystem. The scheduler is the “brains” of the AI Infra control plane. It governs access to the GPU, NPU and CPU capacity of the platform according to a complex system of workload preference rules, placement constraints, optimization objectives, and dynamically interacting policies aimed to maximize hardware utilization and fulfill greatly varying needs of users and the AI platform partner services in terms of workload types, prioritization, and capacity targeting flexibility. The scheduler’s set of capabilities is broad and ambitions. It manages quota, capacity reservations, SLA tiers, preemption, auto-scaling, and a wide range of configurable policies. It is both a workload-aware and topology-aware scheduler down to the level of cluster racks and nodes. Global scheduling is a distinctive major feature that overcomes the regional segmentation of the Azure compute fleet by treating the GPU capacity as a single global virtual pool, which greatly increases capacity availability and utilization for major classes of AI/ML workload. We have achieved this capability by avoiding a significant global single point of failure, based on regional instances of the scheduler service interacting via peer-to-peer protocols for sharing capacity inventory and coordinating handoff of jobs for scheduling. Our system manages significant amount of GPU capacity even outside Azure datacenters, through a unified model and operational process and highly generalized, flexible workload scheduling capabilities. To be able to manage the inherent complexity of the Scheduler subsystem and enable it to meet the stringent expectations of high service reliability, availability, and throughput, we emphasize rigorous engineering, utmost precision and quality, and strong ownership—from feature design to livesite. Quality mindset, attention to detail, development process rigor, and dat




Looking for a job?

Visit Careersaas to find millions of new and unfilled roles, including thousands of remote and hybrid positions.