Senior Kubernetes Operations Engineer

Remote (EU)

Lambda

Published 1 month ago

Hey, this job isn't fresh anymore! 👉 Find fresh remote jobs here

Lambda's GPU cloud is used by deep learning engineers at Stanford, Berkeley, and Carnegie Mellon. Lambda's on-prem systems power research and engineering at Intel, Microsoft, Kaiser Permanente, major universities, and the Department of Defense.

If you'd like to build the world's best deep learning cloud, join us.

What You’ll Do

Remotely install, upgrade, operate and maintain bare-metal Kubernetes clusters (up to thousands of nodes each)
Handle cluster degradation, recovery and resizing using our fleet management tooling
Perform out-of-hours on-call response for critical incidents as part of a well-balanced on-call rotation
Work on improving our tooling, automation, and processes, for both daily operations, alerting, and incident response
Dive into systems at a low level to solve unique cluster problems and write up your findings
Assist customers with high-level Kubernetes questions and integration with applications, storage and authentication
Assist with initial cluster build-outs and validation to help identify failed hardware before customer delivery
Work closely with our HPC Ops and Datacenter Ops teams on issues that require lower-level expertise or cross-functional solutions
Mentor and assist less-experienced team members
Have a voice in our product direction and help us think about how to minimize operational costs and complexity

You

Are an experienced operations engineer, SRE, sysadmin or similar with a deep knowledge of running Linux clusters and systems
Are very familiar with running on bare-metal (including knowledge of BMCs, kernel drivers, PXE, RAID, VLANs, hypervisors)
Have a good understanding of containers, virtualisation, and the mechanisms underpinning them
Have a good understanding of daily operation, bug-fixing and maintenance of Kubernetes
Have experience in an on-call environment and with incident response
Can perform incident post-mortems and develop procedures and tooling to prevent root causes from reoccurring
Have an excellent ability to learn on-the-fly and adapt to solve problems
Are able to work either independently with limited direction, or as part of a team
Are able to work with customers during incidents either via tickets, live messaging, or as part of a larger call.

Nice to Have

Deep Kubernetes experience
Experience with user-level restrictions and hardening (e.g. AppArmor)
Experience with network engineering
Experience with HPC clusters, environments & tooling
Experience with large-scale AI/ML training clusters
Experience with machine learning/AI frameworks
A passion …

This job isn't fresh anymore!

Search Fresh Jobs

Job Profile

Benefits/Perks

401(k) plan with company match Equity Compensation Flexible paid time off Health, dental, and vision coverage

Tasks

Mentor team members

Skills

AI Automation Containers Deep Learning HPC Incident Response Kubernetes Linux Machine Learning ML Network Engineering Storage Virtualization

Experience

5 years

Remote Jobs in North America Remote Jobs in Europe Remote Jobs in South America Remote Jobs in Asia/Pacific Remote Jobs in Africa Remote Jobs in Middle East Full Time Remote Jobs Part Time Remote Jobs Internship Remote Jobs Contract Remote Jobs Temporary Remote Jobs Freelance Remote Jobs Mid-Level Remote Jobs Senior-Level Remote Jobs Entry-Level Remote Jobs Exec-Level Remote Jobs Lead-Level Remote Jobs

Remote Jobs with CAD > 180K in Salary Remote Jobs with EUR > 200K in Salary Remote Jobs with EUR > 220K in Salary Remote Jobs with GBP > 300K in Salary Remote Jobs with GBP > 260K in Salary Remote Jobs with GBP > 280K in Salary Remote Jobs with CAD > 200K in Salary Remote Jobs with EUR > 240K in Salary Remote Jobs with CAD > 220K in Salary Remote Jobs with EUR > 260K in Salary Remote Jobs with CAD > 240K in Salary Remote Jobs with CAD > 260K in Salary