Site Reliability Engineer Jobs
Site Reliability Engineer jobs are open across cloud infrastructure, fintech, healthtech, e-commerce, and enterprise software, from junior SRE to staff and principal level, with specializations in platform engineering, chaos engineering, and observability. Find a role that fits from the openings below and apply directly.
Find JobsLooking for remote work? View remote site reliability engineer jobs →Student or new grad? View site reliability engineer internships →Overview
Showing 5 of 851+ Site Reliability Engineer jobs











INTRODUCTION
The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power our products. We manage a massive, distributed environment built on technologies like Kubernetes, Redis, MySQL, and Message Queue. Our work is not about building features, but about engineering the resilience and performance of the underlying platform that all product teams depend on. We are the guardians of production, ensuring our data systems run smoothly, nonstop. This role includes participation in a rotational on-call schedule to ensure nonstop coverage for our critical data infrastructure. You will be expected to respond to, troubleshoot, and resolve production incidents. Our team collaborates across multiple time zones, and you will engage in rigorous change management and post-incident review processes to maintain system stability. As a Site Reliability Engineer, you will be on the front lines of keeping our large-scale data systems running reliably and efficiently. You will focus on hands-on operational work, from responding to alerts and managing production changes to automating routine tasks. This role is an excellent opportunity to develop deep expertise in modern infrastructure technologies and SRE practices while working alongside senior engineers to solve challenging problems.
Responsibilities
- Incident response and postmortems: Act as an incident commander for critical production issues, guiding the team through triage and resolution. Drive deep, blameless post-incident reviews and ensure that follow-up actions are implemented to prevent recurrence.
- SLO/SLA and error budgets: Define, negotiate, and maintain Service Level Objectives (SLOs) for critical data services. Champion the use of error budgets to balance reliability work with feature development.
- Capacity and cost optimization: Lead initiatives in capacity planning, performance tuning, and resource management. Develop strategies and automation to ensure our infrastructure scales efficiently and stays within budget.
- Pragmatic automation and AI orchestration: Design and build automation and leverage AI Agents to eliminate operational toil, improve deployment safety, and enhance overall operational efficiency. Focus on creating maintainable, robust tools and intelligent workflows that make the entire team more effective.
- Operational excellence and change management: Uphold and improve our standards for production operations, including runbooks, monitoring, and alerting. Vet complex changes and deployments to ensure they meet our bar for production readiness.
- Data Center and AI Infrastructure: Lead the construction, maintenance, and optimization of data centers and specialized AI infrastructure, ensuring high availability and peak performance for complex AI-driven workloads.
- Cross-team influence and mentorship: Act as a subject matter expert on reliability, consulting with application development and other infrastructure teams. Mentor junior SREs, helping them develop their technical and operational skills.
MINIMUM QUALIFICATIONS
- Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
- 5+ years of experience in a Site Reliability Engineering, Production Engineering, or similar role.
- Strong proficiency in a programming or scripting language (e.g., Go, Python, Bash) for automation and tool development.
- Deep understanding of Linux/Unix operating systems, networking fundamentals (TCP/IP, DNS), and distributed systems.
PREFERRED QUALIFICATIONS
- Extensive hands-on experience managing large-scale data infrastructure (e.g., MySQL, Redis, Kafka, Flink).
- Proven experience with container orchestration technologies, particularly Kubernetes, in a production environment.
- Expertise in designing, analyzing, and troubleshooting large-scale distributed systems.
- A systematic problem-solving approach, coupled with strong communication skills and a sense of ownership.
- Experience leading incident response for complex, high-impact events.
- Experience in the operation and construction of Data Centers is a big plus.
ABOUT TIKTOK
TikTok is the leading destination for short-form mobile video. At TikTok, our mission is to inspire creativity and bring joy. TikTok's global headquarters are in Los Angeles and Singapore, and we also have offices in New York City, London, Dublin, Paris, Berlin, Dubai, Jakarta, Seoul, and Tokyo.
WHY JOIN US
Inspiring creativity is at the core of TikTok's mission. Our innovative product is built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and bring joy – a mission we work towards every day. We strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. Every challenge is an opportunity to learn and innovate as one team. We're resilient and embrace challenges as they come. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users. When we create and grow together, the possibilities are limitless. Join us.
DIVERSITY & INCLUSION
TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.
TIKTOK ACCOMMODATION
TikTok is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request
JOB INFORMATION
【For Pay Transparency】Compensation Description (Annually) The base salary range for this position in the selected city is $202160 - $368220 annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units. Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).
The Company reserves the right to modify or change these benefits programs at any time, with or without notice.
For Los Angeles County (unincorporated) Candidates: Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:
1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and
3. Exercising sound judgment.
Site Reliability Engineer Jobs by Experience Level
Top Cities Hiring Site Reliability Engineers
Explore site reliability engineer openings in the cities hiring most right now.
See All 851+ Site Reliability Engineer Jobs
Find roles that match your experience and apply in just a few clicks.
Find JobsSite Reliability Engineer Job Market
Who's Hiring
- Apple29

- JPMorganChase19

- MongoDB17

- Rad Hires14

- Google13

Top Industries Hiring
- Technology & Software178
- Electronics & Hardware26
- Banking & Financial Services25
- Consulting & Professional Services23
- Investment & Asset Management18
What Employers Look For
The qualifications that appear most often in site reliability engineer jobs.
- Proficiency with container orchestration platforms such as Kubernetes and Docker
- Experience with infrastructure-as-code tools including Terraform or Pulumi
- Hands-on background with cloud platforms such as AWS, GCP, or Azure
- Fluency in at least one scripting or programming language such as Python or Go
- Experience designing and maintaining observability stacks using tools like Prometheus, Grafana, or Datadog
- Familiarity with CI/CD pipelines and deployment automation tooling
Tips for Your Site Reliability Engineer Job Search
Quantify your reliability impact clearly
Recruiters and hiring managers scan for SLO, SLA, and SLI metrics you've owned or improved. Include uptime percentages you maintained, incident response times you reduced, and toil-automation wins. Numbers tied to reliability engineering stand out far more than general infrastructure experience.
Tailor your resume to the stack
Site reliability engineer postings vary widely between Kubernetes-heavy shops, Terraform-centric teams, and observability-first orgs. Read each job description for the exact tooling mentioned and mirror that language in your resume. A generic SRE resume loses to one matched to the hiring team's actual stack.
Apply early to roles that fit
Migrate Mate lists site reliability engineer openings from across the United States in one place, so you can find roles that match and apply directly to each listing.
Highlight on-call ownership and postmortems
Many candidates skip on-call history, but hiring teams treat it as proof you've operated systems under real pressure. Note the scale of systems you were on-call for, any blameless postmortem culture you contributed to, and runbooks or playbooks you authored or standardized.
Prepare for system design and failure scenarios
SRE interviews almost always include a distributed systems design question and at least one incident simulation or troubleshooting walkthrough. Practice narrating your diagnostic reasoning out loud, covering how you'd isolate a latency spike or cascading failure across dependent services.
Negotiate scope alongside compensation
When evaluating an offer, ask specifically about on-call rotation frequency, escalation paths, and headcount on the SRE team. Understaffed teams mean heavier rotations. Understanding operational load before you accept is as important as any other term in the offer.
Site Reliability Engineer Jobs: Frequently Asked Questions
Which companies are hiring the most site reliability engineers?
The companies hiring the most site reliability engineers right now include Apple, JPMorganChase, and MongoDB, with the largest share of openings in California, Texas, and New York, based on current listings on Migrate Mate as of August 2026. Demand is consistently high at large cloud-dependent organizations and high-growth SaaS companies.
How many site reliability engineer jobs are remote?
About 59% of site reliability engineer openings are fully remote or hybrid as of August 2026, reflecting the infrastructure-as-code shift that makes most SRE work location-independent. Roles focused on platform engineering, automation, and observability tend to carry the highest share of remote options, while on-site demand is more common for regulated industries like finance and healthcare.
How do you become a site reliability engineer?
Most site reliability engineers start with a strong foundation in either software engineering or systems administration before moving into SRE roles. Building hands-on experience with Linux, networking, and at least one cloud platform is essential. From there, learning infrastructure-as-code, container orchestration, and observability tooling rounds out the core skill set. Contributing to open-source reliability projects or building a home lab with realistic failure scenarios accelerates the transition significantly.
Can you get a site reliability engineer job with little experience?
Breaking into SRE with limited experience is possible by targeting junior or associate SRE roles, which often value demonstrated systems curiosity over years on a resume. Building a public portfolio that shows you've automated something, monitored something, and broken something on purpose is more compelling than certifications alone. Cloud platform certifications from AWS or Google can also help hiring managers assess readiness when your professional history is short.
What does the site reliability engineer interview process look like?
The site reliability engineer interview process typically begins with a recruiter screen focused on your background with cloud infrastructure and on-call experience. Technical rounds usually include a coding exercise in Python or Go, a systems design question covering distributed reliability, and a troubleshooting or incident walkthrough where you diagnose a staged failure scenario. Final rounds often include a conversation with engineering leadership about SLO philosophy and how you'd approach reducing toil on the team.
Where can I find and apply to site reliability engineer jobs?
You can find and apply to site reliability engineer jobs on Migrate Mate, which lists current openings from across the United States. Search the listings to find roles that match your experience and stack, then apply directly to each one that fits. No additional steps are needed between finding a role and submitting your application.
See All 851+ Site Reliability Engineer Jobs
Find roles that match your experience and apply in just a few clicks.
Find Jobs