Lightning AI
Jobs at Lightning AI
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Recently posted jobs
Artificial Intelligence • Machine Learning • Software
Build and operate backend services, control planes, and automation for Lightning AI’s managed Kubernetes and Slurm infrastructure. Develop distributed systems for cluster provisioning, lifecycle management, scheduling, observability, reliability, and scalability across large-scale GPU environments. Diagnose complex production issues and collaborate with infrastructure, AI, and platform engineering teams. Contribute to architecture, technical design, mentoring, engineering practices, and on-call operations.
Artificial Intelligence • Machine Learning • Software
Designs and implements microservices-based platform services, RESTful APIs, backend systems, infrastructure automation, and cloud integrations. Responsibilities include architecture planning, monitoring and alerting, production troubleshooting, disaster recovery, performance optimization, code reviews, mentoring, and sprint planning. The role requires extensive experience with backend and distributed systems, cloud platforms, Kubernetes, microservices, programming languages, infrastructure as code, networking, security, and scalability.
Artificial Intelligence • Machine Learning • Software
Designs, deploys, automates, and operates large-scale NVIDIA InfiniBand fabrics for GPU clusters. Manages UFM, switches, firmware, congestion control, routing, QoS, observability, troubleshooting, capacity planning, and incident response. Collaborates with AI platform, GPU infrastructure, storage, and systems teams to build reliable AI Factory environments.
Artificial Intelligence • Machine Learning • Software
Develop and scale Lightning AI’s backend platform using Go, spanning APIs, infrastructure, billing, security, and integrations. Own features end to end, design reliable and scalable systems, improve architecture and performance, maintain software quality and continuous delivery, reduce technical debt, and mentor engineers. Collaborate with engineering, product, and design teams in a fast-changing SaaS environment.
Artificial Intelligence • Machine Learning • Software
Build and deploy production AI systems for customers, translating business objectives into scalable technical solutions. Responsibilities include architecture, proof-of-concepts, software development, deployment, monitoring, debugging, inference optimization, and distributed systems operation. The engineer partners with customer engineering teams, collaborates with product and engineering, improves reusable platform capabilities, and owns technical engagements from discovery through production scaling.
Artificial Intelligence • Machine Learning • Software
Build and operate production software, APIs, tooling, and automation for large-scale GPU, bare-metal, and HPC infrastructure. Responsibilities include provisioning, configuration, monitoring, lifecycle management, observability, hardware integration, reliability improvements, and infrastructure capacity deployment. The role partners with networking, data center, platform, and infrastructure teams to design scalable systems and define technical direction.
Artificial Intelligence • Machine Learning • Software
Designs, implements, and maintains secure network infrastructure for Lightning AI’s production and AI environments. Responsibilities include firewall, VPN, IDS/IPS, segmentation, vulnerability assessment, penetration testing, SIEM monitoring, incident response, compliance documentation, access controls, and security automation. The role partners with infrastructure and engineering teams to mitigate threats and optimize security controls.
Artificial Intelligence • Machine Learning • Software
This is a general talent community opportunity rather than a specific open position. Lightning AI invites candidates to submit their information for consideration as future roles become available across U.S. and London hubs, with occasional remote opportunities. The company develops tools and infrastructure for building, training, and deploying AI systems and values ownership, urgency, communication, teamwork, continuous improvement, and long-term thinking.
Artificial Intelligence • Machine Learning • Software
Design, optimize, and deploy large language model training and post-training pipelines. Improve model quality through fine-tuning, reinforcement learning, preference optimization, evaluation, and experimentation. Build PyTorch-based infrastructure, optimize distributed multi-GPU training, diagnose performance and convergence issues, and develop production-ready AI systems. Collaborate with researchers, infrastructure engineers, platform teams, and customers while contributing to open-source projects and reusable training capabilities.
Artificial Intelligence • Machine Learning • Software
Build, validate, and operate large-scale bare-metal GPU infrastructure for AI/ML and HPC workloads. Responsibilities include managing Linux systems, image pipelines, test clusters, provisioning, firmware and driver validation, GPU diagnostics, performance analysis with NVIDIA DCGM, automation, virtualization, and hardware management interfaces. The role requires troubleshooting across hardware and software layers while collaborating with infrastructure, hardware, data center, platform, and ML teams.
Artificial Intelligence • Machine Learning • Software
Operate, scale, and optimize distributed storage infrastructure supporting large-scale AI/ML and HPC workloads. Build Python automation, manage Linux bare-metal systems, troubleshoot storage, hardware, networking, and operating system issues, and improve performance, reliability, monitoring, capacity planning, and lifecycle management. Collaborate with infrastructure, networking, platform, and data center teams on storage deployments and scaling strategies.
Artificial Intelligence • Machine Learning • Software
Operate and scale large GPU infrastructure platforms, including Linux systems, bare-metal environments, provisioning workflows, observability, and reliability automation. Responsibilities include platform deployment, incident response, break/fix operations, customer provisioning, on-call participation, infrastructure troubleshooting, and collaboration across engineering, networking, customer success, and software teams. The role also develops automation to reduce manual work and improve operational efficiency.
Artificial Intelligence • Machine Learning • Software
Leads complex, cross-functional infrastructure programs spanning compute, networking, storage, observability, hardware, and datacenter operations. Owns program planning, milestones, dependencies, risks, and execution across multiple workstreams. Partners with engineering, operations, security, platform, and leadership teams to deliver network expansions, datacenter growth, hardware rollouts, GPU platform deployments, and infrastructure readiness. Provides leadership updates and improves execution frameworks, tooling, and program visibility.
Artificial Intelligence • Machine Learning • Software
Develop and post-train deep learning models while building systems, tooling, and workflows for training, evaluation, debugging, and deployment. Contribute to open-source projects, distributed AI infrastructure, backend services, and developer platforms. Collaborate with research, product, infrastructure, and customer teams to solve complex technical problems, prototype ideas, and productionize successful experiments.
Artificial Intelligence • Machine Learning • Software
Develop and scale the Lightning AI platform across frontend, CLI, APIs, and backend systems. Build features using React, Python, or Go; improve stability, performance, architecture, automation, and continuous delivery. Collaborate with engineering, product, and design leaders, own end-to-end feature development, reduce technical debt, and mentor engineers on system design and problem-solving.
Artificial Intelligence • Machine Learning • Software
Support ML engineers running large-scale training and inference workloads across Kubernetes, cloud, and GPU infrastructure. Diagnose distributed PyTorch, CUDA, networking, storage, scheduling, performance, and reliability issues. Analyze observability data, guide customers through complex infrastructure problems, support production incidents, and improve platform reliability through automation, tooling, documentation, runbooks, and operational processes. The role partners closely with infrastructure, networking, and platform engineering teams in a hybrid office environment.
Artificial Intelligence • Machine Learning • Software
Develop and scale the Lightning AI platform’s frontend and UI infrastructure using React and Redux. Build end-to-end features, improve stability and performance, evaluate technical architecture, automate software delivery, reduce technical debt, and mentor engineers. Collaborate with engineering, product, and design teams in a rapidly changing SaaS environment while maintaining high standards for code quality and system design.
Artificial Intelligence • Machine Learning • Software
Design and scale reliable spine-leaf data center networks supporting GPU clusters and AI workloads. Build Ethernet fabrics using EVPN/VXLAN and BGP, optimize traffic, support RoCE/RDMA, and manage backbone, WAN, DCI, and edge connectivity. Develop network automation and Infrastructure-as-Code solutions, improve observability and telemetry, troubleshoot distributed performance issues, and collaborate with compute, storage, platform, and operations teams. The role requires hands-on Cumulus NOS expertise and experience with large-scale, high-performance infrastructure.
