Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Seattle, WA
Financial Services
Leads software engineering and site reliability efforts, delivering secure, scalable production systems. Responsibilities include system design, coding, testing, troubleshooting, automation of recurring remediation, operational stability, vendor architecture evaluations, and adoption of AI-assisted engineering practices with strong validation, security, and responsible-use standards.
Top Skills:
Ai-Assisted Software Development ToolsApplication Resiliency ToolsAutomated TestingCi/CdCloud-Native Technologies
4 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills:
AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills:
AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Reposted 21 Days AgoSaved
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead architecture and implementation of reliability improvements across CrowdStrike's cloud-native platform. Build shared libraries and services, drive observability and SLO practices, perform performance and cost optimization, run resilience engineering and chaos experiments, automate infrastructure-as-code, mentor engineers, and embed with product teams to deliver scalable, highly reliable distributed systems at organizational scale.
Top Skills:
AIAlertingAWSCassandraElasticsearchGCPGoInfrastructure-As-CodeJavaKafkaKotlinKubernetesNode.jsObservability (TracingOciOpensearchProfilingProtobufPythonScalaSlos)
Reposted 23 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Reposted 25 Days AgoSaved
Cloud • Software
Responsible for maintaining FedRAMP-compliant infrastructure, collaborating with software engineers, and ensuring system availability and security. Duties include infrastructure design, automation, monitoring, and incident response.
Top Skills:
AWSGoKubernetesPuppetPythonTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and operate cloud-native, production-grade Kubernetes platforms and tooling to improve reliability, operability, and developer experience. Develop IaC and CI/CD automation, use observability to debug distributed systems, document self-service workflows, and influence engineering teams on scalable architectural patterns. Scope and deliver platform projects with focus on security, cost efficiency, and service stability.
Top Skills:
AksApmAWSAzureC#Ci/CdContainer OrchestrationEksGoInfrastructure As CodeJavaKubernetesLoggingMetricsPulumiPythonTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
The Senior Site Reliability Engineer I will enhance Axon's observability platform, work on distributed tracing, log aggregation, and metrics infrastructure, and develop internal tools while collaborating with engineering teams.
Top Skills:
ArgocdCdkCortexGoGrafanaHelmJaegerJavaLokiOpentelemetryPrometheusPythonTerraform
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Operate and improve Microsoft Defender services in US government cloud environments: provide 24x7 on-call support, run live-site incident response and postmortems, automate deployments and tooling, validate security/compliance during onboarding, collaborate with engineering teams, and apply software engineering best practices to increase reliability and observability at scale.
Top Skills:
C#Ci/CdCloudDistributed SystemsGoJavaMicrosoft DefenderObservabilityPython
Reposted 20 Days AgoSaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 21 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Artificial Intelligence • Natural Language Processing • Generative AI
Own production safeguards infrastructure for Claude model launches and safety classifier deployments. Configure and verify safeguards across first-party, AWS Bedrock, and GCP Vertex platforms; lead canary rollouts, post-deployment validation, incident response, and rollback decisions. Build automation, continuous validation, repeatable deployment pipelines, and a provenance registry to reduce operational toil and configuration drift. Participate in on-call rotations and launch readiness processes.
Top Skills:
AWSGCPLlm Inference SystemsMl InfrastructurePythonRustTransformer-Based Models
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead reliability strategy and SRE best practices for Substrate services in regulated environments. Serve as a senior on-call engineer, lead incident response and post-incident remediation, architect large-scale automation, observability, and self-healing systems, influence cross-organizational design for reliability, security, and compliance, and mentor senior engineers while representing SRE perspectives to leadership.
Top Skills:
AutomationDod)Exchange OnlineGcc HighMicrosoft 365Microsoft Government Cloud (Gcc ModerateMicrosoft SubstrateObservabilitySelf-HealingSlo Frameworks
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills:
BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Artificial Intelligence • Marketing Tech
Own and scale Cognitiv’s AWS infrastructure while evaluating architecture, networking, security, scalability, and service management. Lead improvements in deployments, monitoring, disaster recovery, and infrastructure-as-code practices. Support co-located Equinix datacenter deployments and hybrid cloud operations alongside a datacenter-focused SRE. Collaborate with engineering and product teams, provide multi-datacenter coverage, and help establish long-term service management best practices.
Top Skills:
AnsibleAWSBashDatadogEc2EquinixKubernetesPrometheusPythonTerraform
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills:
Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application, network, and infrastructure reliability, security, performance, and capacity; automate cloud deployments; monitor services and SLAs; analyze logs and events; troubleshoot infrastructure issues; and provide operational recommendations for large-scale cloud platforms.
Top Skills:
Amazon Ec2Amazon Web Services (Aws)AnsibleAws CloudwatchBashCi/CdContainerizationDockerGitGitlabHelmInfrastructure As Code (Iac)JenkinsKubernetesLog AnalysisMavenAzureNagiosOrchestrationPythonSplunkSvnVersion ControlVmware Vsphere
Cloud • Security • Software • Cybersecurity
Deploys and operates scalable, highly available cloud systems; improves application and network security, stability, speed, and capacity; automates cloud deployments; monitors and troubleshoots services to meet SLAs; analyzes logs and events; manages large-scale cloud infrastructure; and resolves infrastructure issues.
Top Skills:
Amazon Ec2Amazon EventbridgeAmazon Route 53Amazon S3Amazon Web ServicesAnsibleAws LambdaBashCi/CdDockerElastic Load BalancingJenkinsJinjaKubernetesPackerPulumiPython
Insurance
Designs and operates reliable hybrid application platforms, leading CI/CD, infrastructure automation, cloud resource management, monitoring, security scanning, container orchestration, and database self-service tooling. The role owns the enterprise CI/CD technology stack and hosting strategy across on-premises, hybrid, and cloud environments. It requires extensive collaboration with engineering, infrastructure, security, application, QA, and governance teams, along with 24/7 mission-critical support and technical leadership.
Top Skills:
AnsibleApache CamelAWSCi/CdCloudFormationConfluenceDatabase SystemsDockerGitGithub ActionsGitlab CiGradleHelmInfrastructure As CodeJenkinsJIRAKubernetesLinuxMavenMonitoring And AlertingNetworkingPythonSonatypeTerraformVulnerability Scanning
Reposted 29 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Big Data • Cloud • Software • Database
Seeking a Site Reliability Engineer with expertise in networking and distributed systems for building secure multi-cloud infrastructure. Responsibilities include maintaining network architecture and ensuring reliable service-to-service communication, involving a 24/7 on-call rotation.
Top Skills:
AWSAzureBgpDnsGCPIpv6KubernetesLoad BalancingMtlsService MeshTcp/IpTlsVpcsVpns
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Designs, operates, and improves large-scale Microsoft 365 and Purview services. Responsibilities include developing automation scripts, building telemetry pipelines and monitoring tools, troubleshooting production issues, optimizing code and services, participating in engineering reviews, and responding to incidents through on-call rotations. The role focuses on cloud engineering, service reliability, observability, resiliency, security, and operational excellence for enterprise and government customers.
Top Skills:
Ai-Powered Operational ToolingAutomationCC#C++Cloud EngineeringDistributed SystemsJavaJavaScriptMicrosoft 365Microsoft PurviewMonitoring ToolsObservabilityPythonTelemetry Pipelines
Cloud • Security
Own reliability, availability, performance, and capacity for production SaaS services across Azure, AWS, and a FedRAMP High environment. Build observability, SLOs, monitoring, infrastructure automation, and remediation workflows; lead incident response, postmortems, support escalations, and disaster recovery efforts. Manage Terraform, CI/CD, Kubernetes, WAF, networking, and observability costs while improving on-call operations and collaborating with Security, Product, Support, and Development.
Top Skills:
AksAWSAzureAzure App ServiceAzure DevopsAzure Front DoorAzure MonitorAzure Service BusAzure SqlAzure StorageAzure WafCloudflareCloudwatch Logs InsightsConsulDatadogElk StackFedrampImpervaIso 27001JenkinsJira Service ManagementJSONKubernetesMicrosoft Entra IdNist 800-53OidcPagerdutyPciPowershellPythonRedisS3SaltstackSAMLSoc 2TerraformYaml
Cloud • Information Technology • Security • Virtual Reality • Cybersecurity
Support reliability, security, scalability, and performance of the TAK CI/CD pipeline and critical infrastructure. Monitor system health, respond to incidents, patch systems, optimize performance and costs, improve DevOps practices, automate operations, and maintain technical documentation. Requires systems and network administration experience, Linux and cloud expertise, containerization, CI/CD knowledge, Zero Trust implementation experience, Security+ certification, and an active T3 investigation.
Top Skills:
AWSCi/Cd PipelinesContainerizationLinuxZero Trust
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Seattle, WA Companies Hiring SRE Engineers
See AllPopular Seattle, WA Engineering Job Searches
Engineering Jobs in Seattle, WA
Software Engineer Jobs in Seattle, WA
Android Developer Jobs in Seattle, WA
C# Jobs in Seattle, WA
C++ Jobs in Seattle, WA
DevOps Jobs in Seattle, WA
Front End Developer Jobs in Seattle, WA
Golang Jobs in Seattle, WA
Hardware Engineer Jobs in Seattle, WA
iOS Developer Jobs in Seattle, WA
Java Developer Jobs in Seattle, WA
Javascript Jobs in Seattle, WA
Linux Jobs in Seattle, WA
Engineering Manager Jobs in Seattle, WA
.NET Developer Jobs in Seattle, WA
PHP Developer Jobs in Seattle, WA
Python Jobs in Seattle, WA
QA Jobs in Seattle, WA
Ruby Jobs in Seattle, WA
Salesforce Developer Jobs in Seattle, WA
Scala Jobs in Seattle, WA
Automation Engineer Jobs in Seattle, WA
AWS Engineer Jobs in Seattle, WA
Backend Engineer Jobs in Seattle, WA
Cloud Engineer Jobs in Seattle, WA
Controls Engineer Jobs in Seattle, WA
CTO Jobs in Seattle, WA
Design Engineer Jobs in Seattle, WA
DevOps Engineer Jobs in Seattle, WA
Director of Engineering Jobs in Seattle, WA
Electrical Engineering Jobs in Seattle, WA
Embedded Software Engineer Jobs in Seattle, WA
Field Engineer Jobs in Seattle, WA
Full-Stack Engineer Jobs in Seattle, WA
Infrastructure Engineer Jobs in Seattle, WA
Manufacturing Engineer Jobs in Seattle, WA
Mechanical Design Engineer Jobs in Seattle, WA
Mechanical Engineering Jobs in Seattle, WA
Network Engineer Jobs in Seattle, WA
Platform Engineer Jobs in Seattle, WA
Principal Engineer Jobs in Seattle, WA
Process Engineer Jobs in Seattle, WA
Product Engineer Jobs in Seattle, WA
Project Engineer Jobs in Seattle, WA
QA Engineer Jobs in Seattle, WA
Robotics Engineer Jobs in Seattle, WA
Security Engineer Jobs in Seattle, WA
Software Architect Jobs in Seattle, WA
Software Development Manager Jobs in Seattle, WA
Solutions Architect Jobs in Seattle, WA
Solutions Engineer Jobs in Seattle, WA
SRE Engineer Jobs in Seattle, WA
Staff Engineer Jobs in Seattle, WA
Staff Software Engineer Jobs in Seattle, WA
Systems Engineer Jobs in Seattle, WA
Web Developer Jobs in Seattle, WA
All Filters
Total selected ()
No Results
No Results


.png)

























