Lead production reliability efforts for a microservices iGaming platform: own observability, support releases and rollbacks, run incident response and RCAs, maintain runbooks, validate deployment readiness, and coordinate cross-functional teams to ensure scalable, secure production operations.
We're building a multi-brand iGaming platform that helps operators run their business in regulated markets worldwide. It brings together three main products — Player Account Management (PAM), Sportsbook, and Casino — all built on a modern microservices architecture using Scala, with a React/TypeScript back office.
We're looking for a Senior SRE to join our team and help us build and maintain reliable, observable, and scalable infrastructure. You'll work closely with developers, own reliability practices, and contribute to the team's DevOps culture.
Requirements
Must Have: Technical Skills
- Linux — confident troubleshooting in terminal, logs, processes, networking basics, resource usage.
- Kubernetes / GKE — strong hands-on understanding of workloads, pods, services, ingress/gateway, probes, RBAC, resources, autoscaling, and troubleshooting.
- GCP — practical experience with cloud infrastructure, especially around GKE, IAM, networking, load balancing, artifact registry, and production diagnostics.
- Docker / Containers — images, registries, runtime debugging, container lifecycle.
- Helm — understands Helm releases, values, deployment state, and rollback.
- GitOps / FluxCD — able to understand and troubleshoot GitOps deployment flow, drift, image automation, and Git-based rollback.
- CI/CD understanding — can investigate failed pipelines and deployment issues; does not need to be the main pipeline builder if DevOps owns that.
- Observability — Prometheus, Alertmanager, Grafana; understands metrics, logs, traces, dashboards, alerting, and production monitoring.
- Networking fundamentals — DNS, load balancers, ingress, gateways, TLS, routing, firewall/security rules.
- SLI / SLO / SLA — can define and apply reliability targets, not just explain the terms.
- Incident response — experience with production incidents, rollback, service recovery, RCA/postmortem.
- Databases / messaging operational basics — PostgreSQL, Couchbase, Kafka, Elasticsearch/ELK or similar; enough to monitor health, migrations, indexes, topics, and failure signals.
- Security / compliance basics — secrets, IAM/RBAC, audit trail, release/change evidence.
Must Have: Reliability / Operational Skills
- Can assess if a release is technically safe to proceed.
- Can define what should be monitored during and after deployment.
- Can verify production health after release.
- Can prepare or validate rollback plans.
- Can lead or support post-release watch.
- Can improve runbooks and incident response procedures.
- Can identify gaps in dashboards, alerts, deployment checks, and release evidence.
- Can work with DevOps without duplicating their ownership.
Must Have: Soft Skills
- Ownership and accountability — follows production issues through to closure.
- Calm under pressure — can operate clearly during incidents, failed deployments, and rollbacks.
- Clear communication — gives concise updates: impact, current action, ETA, risk, decision needed.
- Structured troubleshooting — breaks issues down by app, infra, network, database, config, release diff, external dependency.
- Risk awareness — can say when a release is risky and explain why.
- Collaboration — works well with Dev, QA, Product, Support, Release Manager, and DevOps.
- Documentation discipline — maintains runbooks, rollback steps, incident timelines, and operational evidence.
- Blameless mindset — focuses on root cause and prevention.
- Proactivity — finds missing alerts, dashboards, runbooks, and process gaps before incidents.
- Prioritization — separates real production impact from alert noise.
Nice To Have
- Terraform / IaC experience.
- Advanced GCP infrastructure design.
- ArgoCD or other GitOps tools.
- Advanced database or performance tuning.
- Service mesh experience.
- On-call rotation experience with PagerDuty/Opsgenie or similar.
- Experience facilitating RCA/postmortem sessions.
- Betting/gaming domain experience.
- Good written English for incident updates, release notes, and audit evidence.
- Basic understanding of AI-related concepts: agents, skills, MCP.
- Own production reliability practices together with DevOps and engineering teams.
- Monitor production health and improve observability.
- Support releases, hotfixes, rollback, and post-release watch.
- Validate deployment readiness and rollback readiness.
- Participate in incident response and RCA/postmortem.
- Maintain runbooks, dashboards, alerting rules, and operational documentation.
- Ensure release evidence is traceable and audit-ready.
- Coordinate with Dev, QA, Product, Support, Release Manager, and DevOps during production events.
Similar Jobs
Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Design, implement, and maintain regression and smoke test automation for card and banking platforms. Build scalable MVP regression packs, run automation suites, analyze results, report outcomes, and collaborate with manual testers, developers, and stakeholders to prioritize automation across integrations and environments.
Top Skills:
Api TestingCi/CdCypressIntegration TestingPlaywrightRestassuredSelenium
Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Execute manual functional and end-to-end testing of card lifecycle processes (issuance, authorization, clearing, settlement). Validate business processes (fees, interest, installments), document evidence, log and retest defects, and support SIT, UAT, regression and programme delivery in banking/payments environments. Fluent Polish required.
Top Skills:
AgileErpHp AlmJIRAMastercardVisa
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Lead design, delivery and maintenance of GMP/GDP audit strategy for biologics, aseptic, small molecule and medical device areas. Plan and execute complex audits, coach auditors, analyze regulatory intelligence, drive corrective actions, support inspection readiness, and partner with PGS/PharmSci and site stakeholders to ensure compliance and continuous improvement.
Top Skills:
Aseptic ManufacturingDigital HealthEu DirectiveFdaGdpGmpIchIsoMedical Device RegulationsPic/SQuality Management SystemSoftware As A Medical Device (Samd)Tga
What you need to know about the Seattle Tech Scene
Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.
Key Facts About Seattle Tech
- Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Amazon, Microsoft, Meta, Google
- Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
- Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Madrona, Fuse, Tola, Maveron
- Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute


