Manager, Engineering - Dev Ops/SRE (Hybrid) at CrowdStrike, Inc. — USA - Sunnyvale, CA
Full job description
About the Role
At CrowdStrike, Site Reliability Engineering (SRE) is at the forefront of ensuring the reliability and scalability of our cloud-native security platform. In this role, you'll manage a team of talented engineers, providing technical leadership on key projects and empowering them to excel in their roles.
As an SRE Manager, you will lead a team of SRE engineers ensuring the reliability, scalability, and performance of CrowdStrike's cloud-native security platform. You'll provide technical leadership and mentorship, owning both reliability engineering and software delivery pipelines - driving engineering velocity while maintaining zero tolerance for downtime in security-critical infrastructure.
What you will Do
- Define and enforce SLOs, SLIs, and error budgets across distributed systems processing millions of events per second
- Drive system reliability by blending software engineering principles with AI-driven automation, moving from reactive firefighting to proactive, automated operations
- Lead major incident response and facilitate blameless postmortems, driving systemic reliability improvements
- Own capacity planning, traffic management, and load shedding strategies for high-throughput distributed systems
- Own the end-to-end software delivery pipeline strategy — designing, building, and maintaining scalable, reliable pipelines using Jenkins, GitLab CI, and Bitbucket Pipelines
- Build and maintain observability frameworks including metrics, distributed tracing, and log aggregation across the full stack
- Champion chaos engineering and resilience validation practices for security-critical systems
- Lead and grow a high-performing SRE team, mentoring engineers and fostering a culture of continuous learning and operational excellence
- Partner with cross-functional engineering teams to embed reliability practices early in the software development lifecycle
What You'll Need
Experience & Leadership
- Proven track record of building, growing, and retaining high-performing SRE/DevOps engineering teams in a fast-paced, high-growth environment
- 10+ years of software engineering experience with significant focus on reliability engineering, platform infrastructure, and production operations at scale
- 3+ years of hands-on management experience overseeing SRE/DevOps engineering teams, including incident command and reliability ownership
- Bachelor's degree in Computer Science or related field, or equivalent work experience
Reliability Engineering
- Deep understanding of SRE principles including SLOs, SLAs, SLIs, and error budgeting strategies applied to large-scale distributed systems
- Proven experience owning reliability for high-throughput distributed systems processing millions of events per second, including capacity planning, traffic management, and load shedding strategies
- Strong incident management facilitating blameless postmortems, and driving system reliability improvements
- Demonstrated ability to build, operationalize, and maintain highly scalable, security-critical microservices-based distributed systems with zero tolerance for data loss or downtime.
- Advanced observability experience including Prometheus, Grafana, distributed tracing (Jaeger/OpenTelemetry), and large-scale log aggregation (ELK/Splunk) with a focus on building custom SLO dashboards and reliability scorecards.
- Experience owning disaster recovery strategies including backup automation, failover testing, and business continuity planning for stateful distributed systems
Platform and Delivery Engineering
- Proficiency in Python and/or Golang for automation, tooling, and platform services
- Hands-on experience designing and managing scalable software delivery pipelines using Jenkins, GitLab CI, Bitbucket Pipelines, or equivalent
- Strong proficiency in Infrastructure as Code (IaC) - Terraform, Ansible, Pulumi, or equivalent
- Familiarity with GitOps workflows using ArgoCD or Flux for managing infrastructure deployments at scale
Cloud and Big Data Exposure
- Proficiency in at least one cloud environment (AWS, Azure, GCP) with emphasis on multi-region architecture, cloud-native reliability patterns, and security-first cloud design
- Strong experience with Kubernetes at scale - managing large cluster fleets, workload orchestration, and container lifecycle management
- Familiarity with distributed data systems including relational databases (PostgreSQL), NoSQL (Cassandra), OLAP (Pinot), Indexing(OpenSearch) and real-time streaming platforms (Kafka, Flink)
- Exposure to Big Data and analytics technologies like Spark,Storm.
#LI-AP1
Benefits of Working at CrowdStrike:
- Market leader in compensation and equity awards
- Comprehensive physical and mental wellness programs
- Competitive vacation and holidays for recharge
- Paid parental and adoption leaves
- Professional development opportunities for all employees regardless of level or role
- Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
- Vibrant office culture with world class amenities
- Great Place to Work CertifiedTM across the globe
Notice of E-Verify Participation
Right to Work