Asteralabs · San Jose, CA

Fabric Modeling and Analysis Engineer for Scale Up Fabric (2026-325) at Asteralabs — San Jose, CA

Full-timeSan Jose, CAPosted 2026-07-30Apply on Greenhouse

Full job description

Astera Labs (NASDAQ: ALAB) provides rack-scale AI infrastructure through purpose-built connectivity solutions. By collaborating with hyperscalers and ecosystem partners, Astera Labs enables organizations to unlock the full potential of modern AI. Astera Labs’ Intelligent Connectivity Platform integrates CXL®, Ethernet, NVLink, PCIe®, and UALinkTM semiconductor-based technologies with the company’s COSMOS software suite to unify diverse components into cohesive, flexible systems that deliver end-to-end scale-up, and scale-out connectivity. The company’s custom connectivity solutions business complements its standards-based portfolio, enabling customers to deploy tailored architectures to meet their unique infrastructure requirements. Discover more at www.asteralabs.com.

Fabric Modeling and Analysis Engineer- Scale Up Fabric

Role Overview

Astera Labs is powering the connectivity behind rack-scale AI, and our Scorpio Scale-Up Fabric is central to how tomorrow’s GPU clusters scale. As a Fabric Modeling and Analysis Engineer, you will own the performance models that shape this fabric — quantifying bandwidth, latency, and throughput ceilings, and predicting how Scorpio hardware behaves under real AI/ML workloads long before silicon exists.

This is a high-impact role for an engineer who thrives where architecture, performance analysis, and software converge. Your models will shape the next generate products and enable informed architectural decisions ahead of tape-out, surface bottlenecks that only appear at scale, and translate directly into roadmap and IP decisions. As Astera Labs continues its hyper-growth, you’ll be the internal authority on fabric modeling.

Key Responsibilities

  • Simulation Infrastructure & Model Correlation
  • Design, implement, and maintain features in AI/ML system simulators and fabric models, extending support for collective communication, transport, topology, congestion, routing, buffering, scheduling, and traffic management
  • Improve simulator fidelity, scalability, debuggability, and runtime performance across packet-level, flow-level, and analytical backends, and build reusable abstractions and APIs for end-to-end simulation flows
  • Continuously calibrate and validate models against benchmark data from the Performance Engineering team, owning the correlation between model predictions and measured silicon across scale up fabric generations
  • Analytical & Fabric Performance Modeling
  • Design and implement transaction-level or cycle-approximate system models of the Scale up fabric, enabling rigorous evaluation of new architectural ideas early in the design cycle, well before tape-out
  • Develop and own theoretical roofline and analytical models that establish bandwidth, latency, and throughput ceilings for the Scale-Up fabric, mapping AI/ML workload demands against fabric capabilities to identify compute-bound vs. fabric-bound regimes
  • Model AI/ML collective communication patterns (AllReduce, AllGather, ReduceScatter) and next-generation features such as advanced congestion control, In-Network Computing, and novel resiliency at scale
  • Workload Characterization, Scalability & Topology Analysis

Required skills

  • next.js
  • artificial intelligence
  • machine learning
  • communication