We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
New

Principal Software Engineer

Microsoft
$142,800.00 - $274,800.00 / yr
United States, California, Mountain View
Aug 06, 2026
Overview

The FIT Infrastructure team builds the foundational accelerated compute platforms that power largescale AI training and inference across Azure. Our mission is to deliver secure, reliable, and highly efficient GPU infrastructure that enables multitenant AI systems at global scale while maximizing utilization, performance, and developer productivity.

This role sits at the intersection ofcloud infrastructure, systems software, virtualization, and container platforms, working closely with Azure Infrastructure, OS, Networking, and Hardware teams to deliver end-to-end platform capabilities.

Impact & Growth:

  • You will work onmission critical infrastructurethat directly powers largescale AI systems.
  • Influence the future ofcloud GPU platformsused by internal and external customers.

  • Collaborate with experts acrossOS, hardware, networking, and AI platform teams.

  • Opportunity to grow as atechnical leader, shaping long term platform strategy.



Responsibilities
  • Design and buildGPU accelerated infrastructurefor training and inference workloads, spanning bare metal, virtual machines, and containerized environments.

  • Develop systems forGPU device management, scheduling, isolation, and sharing (e.g., partial GPU allocation, multitenant usage).

  • Build and operateadvanced orchestration and resource governance scenariosusing platforms such asAKS,Dynamic Resource Allocation (DRA), and related Kubernetes ecosystem capabilities to enable fair sharing, isolation, and efficient utilization of accelerated resources.

  • Build and evolvevirtualization and container stacksto support modern AIworkloads, including secure and confidential compute scenarios.

  • Optimizeperformance, reliability, and utilizationacross large GPU fleets, including scaleup and scale out configurations.

  • Partner with networking and storage teams to enablehigh performanceinterconnects(e.g., RDMA/InfiniBand class networking) for distributed workloads.

  • Driveend-to-end platform featuresfrom design through production, including observability, diagnostics, and operational excellence.

  • Influence platform architecture and technical direction across teams through design reviews and technical leadership.



Qualifications

Required Qualifications:

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience.

Preferred Qualifications:

  • Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience.
  • Proven ability to design and operate largescale, production infrastructure with high reliability and performance requirements.
  • Strong problem-solving skills and the ability to debug complex, cross layer systems issues.
  • Demonstrated technical leadership, including mentoring engineers and driving cross team architectural alignment.
  • Hands-on experience withvirtualization and/or container platforms(e.g., VMs, Kubernetes, container runtimes).
  • Strong collaboration and communication skills, with the ability to work across organizational boundaries.
  • Experience in building or operatingmultitenant AI platformsin cloud environments.
  • Familiarity withhigh performancenetworkingand low latency communication stacks.
  • Familiarity withGPU virtualization, passthrough, or partitioningtechnologies.

#AIPLATFORM

Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.

Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Applied = 0

(web-77cf7d65c7-jdxdg)