Senior Staff Software Engineer – SRE, Release & Test Platforms at ServiceNow — Santa Clara, CA
Full job description
Join us to build the next generation of cloud-native reliability, release, and test platforms that enable engineering excellence, developer productivity, and high-confidence ServiceNow releases through automation, observability, and AI-driven operations.
What you get to do in this role:
- Build and operate cloud-native engineering platforms for software validation, release qualification, and operational readiness.
- Design production-like release and test environments that improve release confidence and deployment readiness.
- Develop automated quality gates to assess release health, operational risk, and production readiness.
- Integrate automated testing, observability, reliability signals, and deployment intelligence into CI/CD pipelines.
- Build reusable test frameworks, self-service environments, test data, mock services, and developer productivity tooling.
- Advance shift-left engineering through automated validation, continuous verification, and quality gates.
- Automate failure detection, policy validation, deployment verification, security checks, and reliability assessments.
- Lead Kubernetes-based platform evolution for scalable test infrastructure, release automation, and developer self-service.
- Resolve recurring infrastructure issues through sustainable software, systems, and networking solutions.
- Partner with engineering teams on design reviews, architecture standards, and automation-first reliability practices.
To be successful in this role you have:
- Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
- 12+ years of experience in software, systems, platform, or reliability engineering with a Bachelor's degree; or 8 years and a Master's degree; or a PhD with 5 years experience; or equivalent experience.
- Deep Kubernetes expertise across architecture, operations, networking, storage, security, autoscaling, and multi-cluster environments.
- Experience building and operating large-scale Kubernetes platforms for cloud-native, mission-critical services.
- Experience integrating Kubernetes with CI/CD, GitOps, automated testing, and deployment validation.
- Experience designing cloud-native platforms for ephemeral environments, release qualifications, and automated validation.
- Proven ability to lead engineering excellence across developer productivity, platform engineering, release confidence, and modernization.
- Experience with progressive delivery, including canary releases, feature flags, automated rollback, and deployment verification.
- Experience with chaos engineering, resilience validation, disaster recovery, and reliability assessments.
- Expertise designing, authoring, testing, and debugging code in a team setting using languages such as Python, Go, Java, or Ruby.
- Experience using AI-assisted engineering for intelligent testing, release risk analysis, incident diagnostics, and operational automation.
- Strong coding, observability, SLO, and cross-team collaboration skills to improve reliability, performance, and engineering standards.
Good to have:
- Expertise in observability and monitoring applications, services, and networks at scale.
- Experience with DevOps automation, CI/CD pipelines, and agile methodologies, including GitLab CI/CD or similar tools.
- Experience building enterprise-scale test automation frameworks such as Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit/TestNG, or equivalent technologies.
- Experience with test orchestration, test impact analysis, flaky test detection, parallel execution, and intelligent regression testing.
- Experience with service virtualization, contract testing, synthetic testing, and test data management.
- Experience building engineering platforms that support developer self-service and release engineering.
- Experience with infrastructure configuration management tools such as Ansible.
- Expertise with Kubernetes ecosystem technologies such as Helm, Argo CD, Argo Workflows, Kustomize, Istio/Linkerd, Gateway API/Ingress, Prometheus, OpenTelemetry, and container runtimes.
- Experience implementing GitOps using Argo CD, Flux, or similar technologies.
- Experience operating Kubernetes across AWS (EKS), Azure (AKS), and Google Cloud (GKE).
For positions in this location, we offer a base pay of $190,900 - $334,100, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.
Work Personas
Equal Opportunity Employer
Accommodations
Export Control Regulations