Metasys Logo

Metasys

Site Reliability Engineer Internship

Reposted Yesterday
Remote
Hiring Remotely in United States
Internship
Remote
Hiring Remotely in United States
Internship
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
The summary above was generated by AI
Overview: Reliability and Operational Excellence

The Site Reliability Engineer (SRE) is responsible for the ultimate stability, performance, and scalability of our entire integrated supply chain e-commerce platform. You will apply software engineering principles to operations, ensuring the high availability and resilience of the customer-facing e-commerce storefront, internal SaaS tools (WMS, OMS), and specialized AI agent services.

Internship Details

Duration: 3 months
Start Date: Immediate
Location: Remote
Stipend: None initially. Based on your first-quarter performance, you may be offered a paid full-time opportunity, or even be absorbed directly by the client as an FTE.

Key Responsibilities & Core Projects

You will be the champion of uptime, performance, and automated operations for systems handling the critical MES → WMS → OMS flow.

  • Availability & SLO Management: Define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for core business processes and all application layers. Manage the platform's overall Service Level Agreement (SLA).

  • Observability & Alerting: Architect, maintain, and optimize the comprehensive observability stack (Prometheus, Grafana, Loki, Tempo, OpenTelemetry). Develop high-fidelity alerting and ensure distributed tracing across the NestJS modular monolith and associated data stores (PostgreSQL, Redis).

  • Incident Response & Review: Own the incident response workflow, ensuring rapid triage, mitigation, and root cause analysis. Conduct thorough post-incident reviews to drive continuous improvement and eliminate recurring toil.

  • Scalability & Capacity Planning: Optimize auto-scaling policies for all services running on Docker containers. Conduct capacity planning based on business projections, especially for peak e-commerce and manufacturing load.

  • Disaster Recovery (DR): Design, implement, and regularly test Disaster Recovery procedures, including backup and restoration workflows for PostgreSQL 15 using tools like pgBackRest.

  • Automation: Eliminate operational toil through automation, managing infrastructure-as-code (Terraform) and CI/CD pipelines (Makefile).

Required Technologies & Tools

Candidates must possess deep experience in cloud operations, observability, and infrastructure automation:

  • Observability Stack: Prometheus, Grafana, Loki, Tempo, OpenTelemetry (mandatory).

  • Infrastructure & Platform: Terraform, Docker, Traefik, Oracle Cloud Free VMs (or equivalent public cloud).

  • Data & Resilience: PostgreSQL (Deep knowledge), Redis, pgBackRest.

  • Automation: Strong scripting skills (Python/Bash) and experience with CI/CD tools and Makefile.

  • Methodology: Expert knowledge of SRE principles, toil reduction, and error budgeting.

AI Agent Focus

You will be responsible for the operational reliability of the emerging AI layer.

  • Agent Reliability: Implement specialized monitoring and logging for the AI agent services, ensuring LLM integrations and multi-agent systems (built with frameworks like LangChain) meet defined performance and availability SLOs.

  • Resource Optimization: Efficiently manage resource allocation for computationally intensive AI workloads to maintain platform stability and cost-efficiency.

Success Metrics & Career Path

Performance will be measured by:

  • Uptime/Availability: Achieving defined SLAs/SLOs across the platform.

  • MTTR: Reduction in Mean Time To Recover from production incidents.

  • Toil Reduction: Measured percentage reduction in manual, repetitive operational tasks through automation.

Mentorship Structure: Reports to the Head of Technology/CTO, working collaboratively with DevSecOps and development teams to ensure software is designed for reliability.

Similar Jobs

A Minute Ago
Remote or Hybrid
80K-90K Annually
Senior level
80K-90K Annually
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Supports FP&A for NBCUniversal’s Operations & Technology division through budgeting, forecasting, financial modeling, headcount tracking, OCF consolidation, reporting, and variance analysis. Maintains financial data accuracy, improves FP&A processes and systems, supports reporting hierarchy solutions, and partners with finance, IT, and operational teams on decision-making and ad hoc analyses.
Top Skills: BpcEssbaseExcelMicrosoft Office SuitePowerPointSAP
A Minute Ago
Remote or Hybrid
190K-230K Annually
Expert/Leader
190K-230K Annually
Expert/Leader
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Leads the architecture and deployment of AI-powered agentic workflows and automation across DreamWorks’ animation production pipeline. Designs frameworks, integrates digital content creation tools, manages telemetry, evaluates AI models, and delivers reliable production-ready tools. Partners with artists, engineers, platform teams, and security teams to reduce manual work, establish standards, monitor performance, drive adoption, and measure productivity impact. The role also includes prototyping, documentation, training, and mentoring across technical and creative departments.
Top Skills: Autodesk Flow Production TrackingContainerizationJIRALanggraphLlmsMcpOrchestrationPythonReactSpring AiTypescriptUsd
A Minute Ago
Remote or Hybrid
110K-145K Annually
Senior level
110K-145K Annually
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Owns the lifecycle of GenAI-powered products spanning machine learning systems, media workflows, metadata, and enterprise integrations. Defines requirements, roadmaps, user stories, and product priorities; partners with Engineering, Data Science, Program Management, vendors, and business stakeholders; evaluates model quality, cost, latency, and performance; validates and improves products post-launch; and mentors junior Product Managers.
Top Skills: AgileAirtableConfluenceData ScienceGenaiJIRALarge Language Models (Llms)LucidchartMachine LearningExcelMicrosoft PowerpointMicrosoft WordScrumSharepoint

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account