CrowdStrike, Inc. · USA - Sunnyvale, CA

Manager, Engineering - Dev Ops/SRE (Hybrid) at CrowdStrike, Inc. — USA - Sunnyvale, CA

Full-timeUSA - Sunnyvale, CA$140,000–$215,000/yearPosted 2026-07-16Apply on Workday

Full job description

About the Role

At CrowdStrike, Site Reliability Engineering (SRE) is at the forefront of ensuring the reliability and scalability of our cloud-native security platform. In this role, you'll manage a team of talented engineers, providing technical leadership on key projects and empowering them to excel in their roles.

As an SRE Manager, you will lead a team of SRE engineers ensuring the reliability, scalability, and performance of CrowdStrike's cloud-native security platform. You'll provide technical leadership and mentorship, owning both reliability engineering and software delivery pipelines - driving engineering velocity while maintaining zero tolerance for downtime in security-critical infrastructure.

What you will Do

  • Define and enforce SLOs, SLIs, and error budgets across distributed systems processing millions of events per second
  • Drive system reliability by blending software engineering principles with AI-driven automation, moving from reactive firefighting to proactive, automated operations
  • Lead major incident response and facilitate blameless postmortems, driving systemic reliability improvements
  • Own capacity planning, traffic management, and load shedding strategies for high-throughput distributed systems
  • Own the end-to-end software delivery pipeline strategy — designing, building, and maintaining scalable, reliable pipelines using Jenkins, GitLab CI, and Bitbucket Pipelines
  • Build and maintain observability frameworks including metrics, distributed tracing, and log aggregation across the full stack
  • Champion chaos engineering and resilience validation practices for security-critical systems
  • Lead and grow a high-performing SRE team, mentoring engineers and fostering a culture of continuous learning and operational excellence
  • Partner with cross-functional engineering teams to embed reliability practices early in the software development lifecycle

What You'll Need

Experience & Leadership

  • Proven track record of building, growing, and retaining high-performing SRE/DevOps engineering teams in a fast-paced, high-growth environment
  • 10+ years of software engineering experience with significant focus on reliability engineering, platform infrastructure, and production operations at scale
  • 3+ years of hands-on management experience overseeing SRE/DevOps engineering teams, including incident command and reliability ownership
  • Bachelor's degree in Computer Science or related field, or equivalent work experience

Reliability Engineering

  • Deep understanding of SRE principles including SLOs, SLAs, SLIs, and error budgeting strategies applied to large-scale distributed systems
  • Proven experience owning reliability for high-throughput distributed systems processing millions of events per second, including capacity planning, traffic management, and load shedding strategies
  • Strong incident management facilitating blameless postmortems, and driving system reliability improvements
  • Demonstrated ability to build, operationalize, and maintain highly scalable, security-critical microservices-based distributed systems with zero tolerance for data loss or downtime.
  • Advanced observability experience including Prometheus, Grafana, distributed tracing (Jaeger/OpenTelemetry), and large-scale log aggregation (ELK/Splunk) with a focus on building custom SLO dashboards and reliability scorecards.
  • Experience owning disaster recovery strategies including backup automation, failover testing, and business continuity planning for stateful distributed systems

Platform and Delivery Engineering

  • Proficiency in Python and/or Golang for automation, tooling, and platform services
  • Hands-on experience designing and managing scalable software delivery pipelines using Jenkins, GitLab CI, Bitbucket Pipelines, or equivalent
  • Strong proficiency in Infrastructure as Code (IaC) - Terraform, Ansible, Pulumi, or equivalent
  • Familiarity with GitOps workflows using ArgoCD or Flux for managing infrastructure deployments at scale

Cloud and Big Data Exposure

  • Proficiency in at least one cloud environment (AWS, Azure, GCP) with emphasis on multi-region architecture, cloud-native reliability patterns, and security-first cloud design
  • Strong experience with Kubernetes at scale - managing large cluster fleets, workload orchestration, and container lifecycle management
  • Familiarity with distributed data systems including relational databases (PostgreSQL), NoSQL (Cassandra), OLAP (Pinot), Indexing(OpenSearch) and real-time streaming platforms (Kafka, Flink)
  • Exposure to Big Data and analytics technologies like Spark,Storm.

#LI-AP1

Benefits of Working at CrowdStrike:

  • Market leader in compensation and equity awards
  • Comprehensive physical and mental wellness programs
  • Competitive vacation and holidays for recharge
  • Paid parental and adoption leaves
  • Professional development opportunities for all employees regardless of level or role
  • Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
  • Vibrant office culture with world class amenities
  • Great Place to Work CertifiedTM across the globe

Notice of E-Verify Participation

Right to Work