Manager, Site Reliability Engineering (Auth0) at Okta — New York, DC
Full job description
Secure Every Identity, from AI to Human
### The SRE Leadership Team
The SRE Leadership Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great infrastructure is invisible—it just works. Our team champions a culture of continuous learning, data-driven decision-making, and blameless incident response. We work at the intersection of product engineering, architecture, and operations to ensure Auth0 remains the trusted authentication platform for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with a focus on scalability, resilience, and empowering engineers to grow as technical leaders.
### What You'll Be Doing
- Lead the SRE team's technical direction, translating organizational vision into actionable roadmaps while driving complex, cross-functional initiatives across product and platform teams
- Operate at scale through hands-on participation in 24/7 on-call rotations (follow-the-sun weekdays, shared weekends), directly troubleshooting and remediating incidents on critical systems
- Build infrastructure resilience, designing and implementing monitoring, alerting, and automation improvements that reduce toil and elevate operational efficiency
- Champion reliability best practices, establishing policies and cultural standards that embed observability, resilience, and software engineering rigor into all engineering efforts
- Mentor and develop SRE talent, elevating team capabilities through pair programming, design discussions, and code reviews while fostering a culture of continuous learning
- Represent reliability as a senior technical leader in architectural reviews and strategic planning, ensuring reliability is a core consideration in major engineering decisions
### What You'll Bring to the Role
- 3+ years of hands-on team leadership in SRE or software engineering roles within cloud-native environments, combined with 8+ years of total industry experience
- Deep expertise in cloud platforms (AWS, Azure) and infrastructure as code (Terraform), with proven experience managing cloud-native architectures including containers, Kubernetes, microservices, and databases
- Strong programming skills in Go or Python, with a track record of building and maintaining production-grade tools, automation, and infrastructure solutions
- Data-driven mindset grounded in SRE principles: blameless culture, systematic problem-solving, and the ability to apply software engineering approaches to operational challenges
- Exceptional communication skills—both verbal and written—enabling you to drive clarity during high-pressure incidents and articulate complex concepts to diverse stakeholders
- Proven ability to build and lead high-performing teams in globally distributed, remote-first environments with strong interpersonal and collaboration skills
- Strategic vision and technical depth, combining leadership acumen with hands-on technical excellence and a passion for mentoring senior engineers and shaping team direction
### Extra Credit
- Experience leading reliability initiatives that directly improved system uptime and reduced incident response times at scale
- Contributions to open-source infrastructure or observability tooling
- Experience designing and implementing comprehensive incident response programs and runbook automation
Additional requirements:
- This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.
P13036
$182,000—$250,800 USD
The Okta Experience
- Supporting Your Well-Being
- Driving Social Impact
- Developing Talent and Fostering Connection + Community
Required skills
- team leadership
- microservices
- kubernetes
- artificial intelligence
- communication
- driving
- machine learning
- terraform
- python
- cross-functional