Site Reliability Engineer Jobs in Santa Clara, CA
Site Reliability Engineer jobs in Santa Clara are concentrated in the Central Expressway corridor, the Great America Parkway tech campus zone, and the Mission College district, with demand driven by semiconductor companies, cloud infrastructure providers, and enterprise software firms among NVIDIA, Palo Alto Networks, and d-Matrix. Kubernetes reliability, incident response, and observability tooling are the specialties most in demand right now. See the openings below and apply to the ones that match your experience.
Find JobsOverview
Showing 5 of 196+ Site Reliability Engineer jobs











Our Mission
At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.
Who We Are
In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!
We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.Job Summary
Key Responsibilities
Lead, mentor, and develop a team of Site Reliability/Production Engineers, providing technical direction, coaching, and career development.
Own the reliability, availability, and operational health of critical Cortex services and infrastructure.
Drive improvements in monitoring, alerting, incident management, and observability to identify and resolve issues before they impact customers.
Lead major production incidents, ensure effective root cause analysis, and drive corrective and preventive actions.
Partner with Engineering teams to improve service architecture, production readiness, scalability, and operability.
Drive automation and self-healing solutions to reduce operational toil and improve engineering efficiency.
Establish clear priorities, operational processes, and measurable reliability goals for the team.
Collaborate with Production Engineering teams across regions to strengthen follow-the-sun operations and consistent operational practices.
Evaluate new technologies and drive adoption of solutions that improve reliability, scalability, and operational efficiency.
Qualifications
Required Qualifications
10+ years of experience in Site Reliability Engineering, DevOps, Production Engineering, Cloud Infrastructure, or related areas, including experience leading or managing engineering teams.
Strong technical knowledge of cloud platforms, preferably Google Cloud Platform (GCP).
Strong experience with Kubernetes, containerized environments, and large-scale distributed systems.
Experience with observability technologies such as Prometheus, Thanos, Grafana, OpenTelemetry, and cloud-native monitoring solutions.
Strong understanding of incident management, monitoring, alerting, SLOs, and reliability engineering practices.
Experience with automation and Infrastructure as Code using technologies such as Python, Terraform, Ansible, and GitOps.
Demonstrated ability to lead engineers, set priorities, drive execution, and manage multiple operational and technical initiatives.
Strong communication skills with the ability to collaborate and influence across engineering teams and global organizations.
Preferred Qualifications
Experience managing SRE, DevOps, Production Engineering, or infrastructure teams supporting large-scale SaaS platforms.
Experience operating highly available, multi-region cloud environments.
Experience driving automation, self-healing, or AI-assisted operational capabilities.
Strong operational leadership with experience managing high-severity production incidents.
Proven ability to develop engineers, improve team processes, and drive a culture of ownership and continuous improvement.
Compensation Disclosure
The compensation offered for this position will depend on qualifications, experience, and work location. For candidates who receive an offer at the posted level, the starting base salary (for non-sales roles) or base salary + commission target (for sales/com-missioned roles) is expected to be the annual range listed below. The offered compensation may also include restricted stock units and a bonus.
- /yr
Our Commitment
We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.
We are committed to providing reasonable accommodations for all qualified individuals with a disability. If you require assistance or accommodation due to a disability or special need, please contact us at accommodations@paloaltonetworks.com.
Palo Alto Networks is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.
All your information will be kept confidential according to EEO guidelines.
Is role eligible for Immigration Sponsorship? No. Please note that we will not sponsor applicants for work visas for this position.See All 196+ Site Reliability Engineer Jobs in Santa Clara
Find roles in Santa Clara that match your experience and apply in just a few clicks.
Find JobsSite Reliability Engineer Job Market in Santa Clara
Who's Hiring
- NVIDIA51

- Palo Alto Networks51

- d-Matrix26

- IonQ9

- Everpure9

Top Industries Hiring
- Technology & Software60
- Agriculture & Farming17
- Law & Legal Services9
Site Reliability Engineer Jobs in Santa Clara: Frequently Asked Questions
How do I get a site reliability engineer job in Santa Clara?
Focus your search on Santa Clara's semiconductor, cloud infrastructure, and enterprise software employers, which cluster along Great America Parkway, the Central Expressway corridor, and the Mission College area. Hands-on experience with distributed systems, on-call incident management, and infrastructure-as-code tools like Terraform or Ansible gives candidates a clear edge. Chip and hardware companies in Santa Clara also value SREs who can support embedded or mixed on-prem and cloud environments, so that combination opens doors most candidates miss.
Which companies hire site reliability engineers in Santa Clara?
Employers hiring site reliability engineers in Santa Clara right now include NVIDIA, Palo Alto Networks, and d-Matrix, based on current listings on Migrate Mate as of August 2026. Santa Clara's employer mix leans heavily toward semiconductor manufacturers, large cloud platform operators, and enterprise IT vendors, so site reliability engineers here often work at established tech giants or well-funded infrastructure-focused companies rather than early-stage startups.
Are there remote site reliability engineer jobs in Santa Clara?
Yes, though remote availability depends on the role: SRE positions tied to physical data centers or on-site hardware in Santa Clara tend to require in-person presence, while roles focused on cloud operations, monitoring, and automation pipelines are often hybrid or fully remote. About 33% of site reliability engineer openings tied to Santa Clara are remote or hybrid as of August 2026, reflecting the mix of hardware-adjacent and cloud-native employers here. Observability engineering and automation-heavy roles carry the highest remote flexibility locally.
How can I get a site reliability engineer job in Santa Clara with little or no experience?
The most realistic entry path in Santa Clara is landing a junior DevOps, systems administrator, or production support role at one of the city's mid-size enterprise software or networking companies, then building toward SRE responsibilities over time. Santa Clara employers in the chip and semiconductor space regularly hire associate infrastructure engineers who can grow into reliability roles. Certifications in cloud platforms like AWS or Google Cloud, paired with a public GitHub showing automation scripts or monitoring dashboards, significantly improve candidacy for these entry-level openings.
Which industries hire the most site reliability engineers in Santa Clara?
Most site reliability engineer openings in Santa Clara sit in Technology & Software, Agriculture & Farming, and Law & Legal Services, per current listings on Migrate Mate as of August 2026. Santa Clara's position as a global center for semiconductor design, corporate networking hardware, and cloud infrastructure means SRE demand is especially deep in companies building or operating the physical and virtual layers of the internet, which is unusual compared to other Bay Area cities that skew more toward consumer software.
Related Jobs in California
See All 196+ Site Reliability Engineer Jobs in Santa Clara
Find roles in Santa Clara that match your experience and apply in just a few clicks.
Find Jobs