Site Reliability Engineer Jobs in San Francisco, CA
Site Reliability Engineer jobs in San Francisco are in high demand, concentrated in SoMa, the Financial District, and Mission Bay across cloud infrastructure, fintech, and enterprise software. Employers hiring right now include Braze, Wayve, and Harvey. Scan the live roles below and apply to whichever ones fit.
Find JobsOverview
Showing 5 of 212+ Site Reliability Engineer jobs








Location
San Francisco
Employment Type
Full time
Department
Engineering
Compensation
- $180K – $250K • Offers Equity
fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
About this role
You are a seasoned SRE who keeps production infrastructure running at scale. You own the reliability and availability of customer-facing systems — from Kubernetes clusters to deployment pipelines to the networking layer that connects it all. You think in SLOs, automate ruthlessly, and treat every incident as a chance to make the system better.
Key Responsibilities
Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads
Build and maintain CI/CD pipelines and deployment infrastructure
Leverage AI to an extreme level to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability
Build dashboards, alerting, and anomaly detection across our systems
Define and enforce SLOs and build out incident response processes
Manage and improve our networking, load balancing, and service mesh configurations
Drive reliability improvements across the stack through automation, runbooks, and chaos engineering
Requirements
5+ years experience in managing critical production systems and software development workflows
Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible)
Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS
Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD)
Proficiency in Python and either Go or Bash for tooling and automation
Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog)
Excellent communication and ability to drive technical decisions across teams
Self-starter who executes quickly, takes ownership, and constantly seeks improvement
Nice to have
Experience with managing GPU and AI/ML workloads
Experience with kernel-based monitoring and routing (eBPF, XDP)
Experience with security tooling (Falco, Coroot, SIEM)
Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB)
Experience with distributed storage systems (Ceph, Longhorn, etc.)
Location
San Francisco, CA
What we offer at fal
Interesting and challenging work
A lot of learning and growth opportunities
We are currently hiring in downtown San Francisco.
Health, dental, and vision insurance (US)
Regular team events and offsites
U.S. EQUAL EMPLOYMENT OPPORTUNITY INFORMATION:
fal provides equal employment opportunities to applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability, or any other classification protected by applicable law.
Compensation Range: $180K - $250KSee All 212+ Site Reliability Engineer Jobs in San Francisco
Find roles in San Francisco that match your experience and apply in just a few clicks.
Find JobsSite Reliability Engineer Job Market in San Francisco
Who's Hiring
- Braze10

- Wayve10

- Harvey10

- Lambda10

- MongoDB10

Top Industries Hiring
- Technology & Software58
- Social Media5
- Cybersecurity5
- Marketing & Advertising5
- Banking & Financial Services5
Site Reliability Engineer Jobs in San Francisco: Frequently Asked Questions
How do I get a site reliability engineer job in San Francisco?
Target San Francisco's density of cloud-native companies in SoMa and Mission Bay, plus the fintech and enterprise software firms in the Financial District. Hands-on experience with Kubernetes, Terraform, and observability tooling stands out in this market. Contributing to open-source infrastructure projects and networking through local DevOps and SRE meetups in the city gives candidates a concrete edge over remote applicants.
Which companies hire site reliability engineers in San Francisco?
San Francisco site reliability engineer roles are posted by Braze, Wayve, and Harvey and others right now, based on current listings on Migrate Mate as of September 2026. The city's hiring base is especially strong among high-growth cloud platforms, consumer tech companies, and financial technology firms headquartered in SoMa and the Financial District.
Are there remote site reliability engineer jobs in San Francisco?
Yes, though many SRE roles require at least some on-site presence given hands-on infrastructure and incident response responsibilities. About 68% of site reliability engineer openings tied to San Francisco are remote or hybrid as of September 2026, skewing toward hybrid schedules. Automation, monitoring, and on-call engineering tasks are most commonly performed remotely within San Francisco-based teams.
How can I get a site reliability engineer job in San Francisco with little or no experience?
The most realistic entry path in San Francisco is moving laterally from a systems administration or DevOps engineer role at one of the city's mid-size SaaS or fintech companies, which tend to hire junior SRE talent more actively than large enterprises. Building skills in Linux, cloud platforms like AWS or GCP, and basic scripting, then applying to associate infrastructure or platform engineer roles, is the clearest route into the field locally.
Which industries hire the most site reliability engineers in San Francisco?
San Francisco site reliability engineer roles concentrate in Technology & Software, Social Media, and Cybersecurity, based on current listings on Migrate Mate as of September 2026. San Francisco's position as a global hub for cloud-native software, financial technology, and consumer platforms means these sectors consistently generate the highest volume of SRE openings in the city.
Related Jobs in California
See All 212+ Site Reliability Engineer Jobs in San Francisco
Find roles in San Francisco that match your experience and apply in just a few clicks.
Find Jobs