Baseten Logo

Baseten

Site Reliability Engineer (SRE)

Job Posted 21 Days Ago Reposted 21 Days Ago
In-Office or Remote
3 Locations
150K-250K Annually
Mid level
In-Office or Remote
3 Locations
150K-250K Annually
Mid level
As a Site Reliability Engineer, you will build and maintain scalable infrastructure, automate processes, and collaborate with cross-functional teams while mentoring others and owning projects end-to-end.
The summary above was generated by AI

ABOUT BASETEN

Baseten provides the infrastructure, tooling, and expertise needed to bring great AI products to market - fast. Backed by top investors including IVP, Spark Capital, Greylock, and Conviction, we’re trusted by leading AI-driven innovators like Writer, Abridge, Bland, Patreon, Descript, Retool, and Zed to deliver industry-leading performance, security, and reliability for their mission-critical workloads. With our recent $75M Series C funding, we’re growing fast to make AI accessible across all products.

THE ROLE

As a Site Reliability Engineer, you'll envision and build robust systems and processes that ensure our infrastructure is scalable, reliable, and efficient. This can range from automating deployments and monitoring systems to optimizing performance and managing incidents.

We all work closely with our users, learning from their past struggles in operationalizing ML, onboarding them onto our platform, and turning our learnings into ideas for improving Baseten.

EXAMPLE INITIATIVES

You'll get to work on these types of projects as part of our Infrastructure team:

  • Multi-cloud capacity management

  • Inference on B200 GPUs

  • Multi-node inference

  • Fractional H100 GPUs for efficient model serving

RESPONSIBILITIES

  • Build and maintain scalable infrastructure to support the deployment and operation of machine learning models.

  • Establish standards and best practices for reliability and performance across the infrastructure.

  • Automate processes when relevant, particularly for managing CI/CD pipelines.

  • Own products and projects end-to-end, functioning as both an engineer and a project manager, with a focus on user empathy, project specification, and end-to-end execution.

  • Collaborate with cross-functional teams to understand project requirements and translate them into technical solutions.

  • Mentor junior team members and contribute to knowledge sharing within the organization.

  • Navigate ambiguity and exercise good judgment on tradeoffs and tools needed to solve problems, avoiding unnecessary complexity.

  • Demonstrate pride, ownership, and accountability for your work, expecting the same from your teammates.

REQUIREMENTS

  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field.

  • 3+ years of work professional work experience in a fast-paced, high-growth environment.

  • Extensive experience with Kubernetes.

  • Experience in building and maintaining scalable infrastructure.

  • Experience with infrastructure-as-code tools (e.g., Terraform, CloudFormation, Pulumi) and CI/CD tooling (e.g., GitHub Actions, GitLab CI, Circle CI, Jenkins).

  • Relevant OSS observability experience (Prometheus, ELK stack, Grafana stack, Opentelemetry) is a plus.

  • Ability to own projects end-to-end, from project specification to execution.

  • No prior machine learning experience required, but should be open to learning about it.

BENEFITS

  • Competitive compensation package (Flexible PTO, 401k, covered healthcare premiums).

  • This is a unique opportunity to be part of a rapidly growing startup in one of the most exciting engineering fields of our era.

  • An inclusive and supportive work culture that fosters learning and growth.

  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.


At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

Top Skills

Circle Ci
CloudFormation
Elk Stack
Github Actions
Gitlab Ci
Grafana
Jenkins
Kubernetes
Prometheus
Pulumi
Terraform

Similar Jobs

3 Days Ago
Remote or Hybrid
San Diego, CA, USA
111K-172K Annually
Junior
111K-172K Annually
Junior
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
As a Site Reliability Engineer, you'll enhance the reliability and performance of ServiceNow's infrastructure, troubleshooting issues, driving automation, and employing DevOps practices.
Top Skills: AutomationAWSAzureCloud TechnologiesDevOpsJavaScriptLinuxMySQLPythonRuby
4 Days Ago
Remote
United States
161K-180K Annually
Senior level
161K-180K Annually
Senior level
Consumer Web • Digital Media • Information Technology • News + Entertainment • Social Media
The Senior Site Reliability Engineer will optimize system performance, enhance infrastructure resilience, and lead improvements for both physical and cloud systems, collaborating with other engineering teams.
Top Skills: AnsibleBashCC++DjangoDockerFlaskGoJavaKubernetesLaravelLinuxMvcOrmPythonRustTerraform
4 Days Ago
Remote or Hybrid
San Diego, CA, USA
156K-273K Annually
Senior level
156K-273K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Lead Site Reliability Engineering efforts within the DevSecOps team, focusing on operational excellence, security, reliability, and cost optimization for security services, while mentoring a team of SREs.
Top Skills: AIAnsibleAWSAzureBashDockerElkGCPGrafanaKubernetesPrometheusPythonTerraform

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute
By clicking Apply you agree to share your profile information with the hiring company.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account