1000 KLA Corporation · Ann Arbor, MI

HPC / AI Software Infrastructure Lead (E) at 1000 KLA Corporation — Ann Arbor, MI

Full-timeAnn Arbor, MI$151,100–$256,900/yearPosted 2026-07-29Apply on Workday

Full job description

Company Overview

HPC/AI Software Infrastructure Leads are core to KLA’s technology, while we do not currently have an opening, we are always building our HPC/AI Software Infrastructure Lead Engineering talent community, we are interested in learning about your background.

Apply to this posting for Future Opportunities with KLA.

At KLA, we’re pushing the boundaries of semiconductor inspection through advanced AI and high-performance computing. We are looking for a hands-on technical leader to architect and scale the next generation of AI/HPC infrastructure powering our most critical imaging and data platforms. This role is ideal for someone who thrives at the intersection of distributed systems, GPU computing, and real-world AI workloads, and who enjoys building and mentoring high-performing engineering teams while driving technical excellence.

What You’ll Do

  • Lead the architecture and development of large-scale HPC and AI infrastructure supporting cutting-edge image processing and machine learning workloads
  • Design scalable, high-performance distributed systems that unify traditional image processing with modern AI/Deep Learning pipelines
  • Drive GPU-accelerated computing strategies, optimizing performance across compute, storage, and networking layers
  • Partner cross-functionally with hardware, algorithms, and product teams to deliver robust, production-ready platforms
  • Establish engineering best practices (code quality, CI/CD, observability, performance tuning) for mission-critical systems
  • Mentor and develop engineers, providing technical guidance, coaching, and growth opportunities for junior team members
  • Serve as a technical leader and decision-maker, influencing architecture and long-term platform strategy

What You Bring

Experience

  • 10+ years in software engineering, including leading and scaling technical teams
  • Proven success building distributed systems in HPC, AI/ML, or cloud-native environments
  • Track record of delivering performance-critical infrastructure at scale
  • Experience mentoring and growing early- and mid-career engineers

Technical Expertise

  • Deep understanding of distributed systems, parallel computing, and Linux systems programming
  • Strong programming skills in C++, Python, or similar systems-level languages
  • Experience with GPU computing (CUDA, ROCm) and modern AI frameworks (PyTorch, TensorFlow, etc.)
  • Familiarity with high-performance storage systems, networking, and data pipelines
  • Strong foundation in CI/CD, DevOps, and production system reliability

Bonus Experience

  • Background in image processing, computer vision, or scientific computing
  • Experience supporting hybrid HPC + AI workloads in production environments

Leadership & Impact

  • Passion for developing talent and building inclusive, high-performing teams
  • Ability to operate as both a hands-on engineer and strategic technical leader
  • Strong communication skills with the ability to influence across engineering and product stakeholders

Why KLA / Why Ann Arbor

  • Work on real-world AI systems at scale, not just experiments
  • Collaborate across hardware, software, and algorithm teams in a deeply technical environment
  • Join a growing engineering presence in Ann Arbor, with access to top talent and a strong technical community
  • Opportunity to shape the direction of AI infrastructure in a core product domain

Minimum Qualifications

Doctorate (Academic) Degree and related work experience of 5 years; Master's Level Degree and related work experience of 8 years; Bachelor's Level Degree and related work experience of 12 years

Base Pay Range: $151,100.00 - $256,900.00

Primary Location: USA-MI-Ann Arbor-KLA