Site Reliability Engineer Jobs
Site Reliability Engineer jobs are open across cloud infrastructure, fintech, healthtech, e-commerce, and enterprise software, from junior SRE to staff and principal level, with specializations in platform engineering, chaos engineering, and observability. Find a role that fits from the openings below and apply directly.
Find JobsLooking for remote work? View remote site reliability engineer jobs →Student or new grad? View site reliability engineer internships →Overview
Showing 5 of 804+ Site Reliability Engineer jobs











We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time.
Position Summary
The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, scalability, and performance of CVS Health's Retail and Pharmacy platforms. This role combines software engineering, operations, observability, and automation practices to proactively identify and resolve issues, improve system resilience, and support critical store operations.
As part of the SRE organization, you will partner with application development, infrastructure, observability, and store operations teams to drive operational excellence, implement reliability engineering best practices, and enable highly scalable deployments across thousands of retail and pharmacy locations.
******This position requires working in shifts and on weekends, with compensatory time off provided on weekdays.
Key Responsibilities:
Observability & Monitoring
- Develop and implement proactive monitoring, alerting, and dashboarding strategies to detect issues before they impact store operations or customer experience.
- Design and maintain operational dashboards using enterprise observability platforms.
- Define, monitor, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), Service Level Agreements (SLAs), and error budgets for critical business services.
- Analyze platform telemetry, logs, traces, and metrics to improve service reliability and reduce operational risk.
- Drive continuous improvements in observability maturity across Edge applications and services.
Reliability Engineering & Incident Management
- Lead major incident response, recovery, and post-incident reviews to minimize customer impact and prevent recurring issues.
- Perform root cause analysis and drive corrective and preventive actions through structured Problem Management practices.
- Improve key operational metrics including Mean Time to Detect (MTTD), Mean Time to Resolve (MTTR), and service availability.
- Collaborate with engineering teams to build reliability into applications throughout the Software Development Lifecycle (SDLC).
- Drive automation initiatives to reduce operational toil and improve system resiliency.
Performance & Platform Optimization
- Identify and eliminate bottlenecks in development, testing, and deployment workflows.
- Support performance tuning and capacity planning for Edge retail and pharmacy applications.
- Analyze system behavior and implement improvements that enhance scalability, stability, and efficiency.
- Partner with infrastructure teams to maintain highly available and resilient platform services.
Edge Platform Operations
- Support business-critical applications deployed across CVS retail and pharmacy locations.
- Collaborate with store operations and engineering teams to ensure seamless operation of Edge platforms.
- Participate in on-call rotations and provide technical leadership during production incidents.
- Ensure operational readiness, deployment validation, and production support for new platform capabilities.
Cloud, Microservices & Deployment Engineering
- Champion cloud-native technologies and container-based architectures.
- Support and optimize Kubernetes and OpenShift environments operating in hybrid cloud ecosystems.
- Leverage CI/CD pipelines and Infrastructure-as-Code principles to enable automated, scalable deployments.
- Promote best practices for microservices architecture, resiliency, and deployment automation.
- Work closely with development teams to ensure services are observable, scalable, and production ready.
Required Qualifications
- 5+ years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Infrastructure Engineering, or related technology roles.
- 3+ years of experience delivering and supporting large-scale distributed systems utilizing reliability and resiliency concepts.
- 2+ years of experience with one or more programming languages such as Java, Python, Go, or JavaScript.
- 2+ years of experience with cloud platforms including AWS, Microsoft Azure, or Google Cloud Platform.
- Hands-on experience with Kubernetes, OpenShift, Docker, Rancher, and containerized workloads.
- Experience implementing and supporting CI/CD pipelines using tools such as GitHub, Bitbucket, Jenkins, GitLab, or similar platforms.
- Experience with observability and monitoring tools such as Splunk, Dynatrace, Datadog, Prometheus, Grafana, OpenTelemetry, or similar technologies.
- Strong scripting and automation skills using Shell, Python, PowerShell, or equivalent technologies.
- Experience supporting microservices-based and cloud-native architectures.
- Working knowledge of Incident Management, Problem Management, Change Management, and ITIL-based operational practices.
- Excellent analytical, troubleshooting, communication, and collaboration skills.
Preferred Qualifications
- Experience supporting retail, pharmacy, healthcare, or large-scale edge computing environments.
- Experience designing and implementing SLO, SLA, and Error Budget frameworks.
- Knowledge of distributed tracing and observability best practices.
- Experience driving platform modernization, reliability engineering initiatives, and operational excellence programs.
- Certifications such as AWS Certified Solutions Architect, Kubernetes (CKA/CKAD), Google SRE, or related cloud certifications.
Education
Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
Anticipated Weekly Hours
40Time Type
Full timePay Range
The typical pay range for this role is:
$92,700.00 - $203,940.00This pay range represents the base hourly rate or base annual full-time salary for all positions in the job grade within which this position falls. The actual base salary offer will depend on a variety of factors including experience, education, geography and other relevant factors. This position is eligible for a CVS Health bonus, commission or short-term incentive program in addition to the base pay range listed above.
Our people fuel our future. Our teams reflect the customers, patients, members and communities we serve and we are committed to fostering a workplace where every colleague feels valued and that they belong.
Great benefits for great people
We take pride in offering a comprehensive and competitive mix of pay and benefits that reflects our commitment to our colleagues and their families.
This full‑time position is eligible for a comprehensive benefits package designed to support the physical, emotional, and financial well‑being of colleagues and their families. The benefits for this position include medical, dental, and vision coverage, paid time off, retirement savings options, wellness programs, and other resources, based on eligibility.
Additional details about available benefits are provided during the application process and on Benefits Moments.
We anticipate the application window for this opening will close on: 10/30/2026
Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state and local laws.
Site Reliability Engineer Jobs by Experience Level
Top Cities Hiring Site Reliability Engineers
Explore site reliability engineer openings in the cities hiring most right now.
See All 804+ Site Reliability Engineer Jobs
Find roles that match your experience and apply in just a few clicks.
Find JobsSite Reliability Engineer Job Market
Who's Hiring
- Apple25

- Google23

- MongoDB19

- SpaceX17

- Tesla16

Top Industries Hiring
- Technology & Software128
- Electronics & Hardware18
- Banking & Financial Services14
- Consulting & Professional Services14
- Investment & Asset Management13
What Employers Look For
The qualifications that appear most often in site reliability engineer jobs.
- Proficiency with container orchestration platforms such as Kubernetes and Docker
- Experience with infrastructure-as-code tools including Terraform or Pulumi
- Hands-on background with cloud platforms such as AWS, GCP, or Azure
- Fluency in at least one scripting or programming language such as Python or Go
- Experience designing and maintaining observability stacks using tools like Prometheus, Grafana, or Datadog
- Familiarity with CI/CD pipelines and deployment automation tooling
Tips for Your Site Reliability Engineer Job Search
Quantify your reliability impact clearly
Recruiters and hiring managers scan for SLO, SLA, and SLI metrics you've owned or improved. Include uptime percentages you maintained, incident response times you reduced, and toil-automation wins. Numbers tied to reliability engineering stand out far more than general infrastructure experience.
Tailor your resume to the stack
Site reliability engineer postings vary widely between Kubernetes-heavy shops, Terraform-centric teams, and observability-first orgs. Read each job description for the exact tooling mentioned and mirror that language in your resume. A generic SRE resume loses to one matched to the hiring team's actual stack.
Apply early to roles that fit
Migrate Mate lists site reliability engineer openings from across the United States in one place, so you can find roles that match and apply directly to each listing.
Highlight on-call ownership and postmortems
Many candidates skip on-call history, but hiring teams treat it as proof you've operated systems under real pressure. Note the scale of systems you were on-call for, any blameless postmortem culture you contributed to, and runbooks or playbooks you authored or standardized.
Prepare for system design and failure scenarios
SRE interviews almost always include a distributed systems design question and at least one incident simulation or troubleshooting walkthrough. Practice narrating your diagnostic reasoning out loud, covering how you'd isolate a latency spike or cascading failure across dependent services.
Negotiate scope alongside compensation
When evaluating an offer, ask specifically about on-call rotation frequency, escalation paths, and headcount on the SRE team. Understaffed teams mean heavier rotations. Understanding operational load before you accept is as important as any other term in the offer.
Site Reliability Engineer Jobs: Frequently Asked Questions
Which companies are hiring the most site reliability engineers?
The companies hiring the most site reliability engineers right now include Apple, Google, and MongoDB, with the largest share of openings in California, Texas, and New York, based on current listings on Migrate Mate as of September 2026. Demand is consistently high at large cloud-dependent organizations and high-growth SaaS companies.
How many site reliability engineer jobs are remote?
About 66% of site reliability engineer openings are fully remote or hybrid as of September 2026, reflecting the infrastructure-as-code shift that makes most SRE work location-independent. Roles focused on platform engineering, automation, and observability tend to carry the highest share of remote options, while on-site demand is more common for regulated industries like finance and healthcare.
How do you become a site reliability engineer?
Most site reliability engineers start with a strong foundation in either software engineering or systems administration before moving into SRE roles. Building hands-on experience with Linux, networking, and at least one cloud platform is essential. From there, learning infrastructure-as-code, container orchestration, and observability tooling rounds out the core skill set. Contributing to open-source reliability projects or building a home lab with realistic failure scenarios accelerates the transition significantly.
Can you get a site reliability engineer job with little experience?
Breaking into SRE with limited experience is possible by targeting junior or associate SRE roles, which often value demonstrated systems curiosity over years on a resume. Building a public portfolio that shows you've automated something, monitored something, and broken something on purpose is more compelling than certifications alone. Cloud platform certifications from AWS or Google can also help hiring managers assess readiness when your professional history is short.
What does the site reliability engineer interview process look like?
The site reliability engineer interview process typically begins with a recruiter screen focused on your background with cloud infrastructure and on-call experience. Technical rounds usually include a coding exercise in Python or Go, a systems design question covering distributed reliability, and a troubleshooting or incident walkthrough where you diagnose a staged failure scenario. Final rounds often include a conversation with engineering leadership about SLO philosophy and how you'd approach reducing toil on the team.
Where can I find and apply to site reliability engineer jobs?
You can find and apply to site reliability engineer jobs on Migrate Mate, which lists current openings from across the United States. Search the listings to find roles that match your experience and stack, then apply directly to each one that fits. No additional steps are needed between finding a role and submitting your application.
See All 804+ Site Reliability Engineer Jobs
Find roles that match your experience and apply in just a few clicks.
Find Jobs