Senior Kubernetes Operations Engineer
Remote (UK)
Hey, this job isn't fresh anymore! 👉 Find fresh remote jobs here
Lambda's GPU cloud is used by deep learning engineers at Stanford, Berkeley, and Carnegie Mellon. Lambda's on-prem systems power research and engineering at Intel, Microsoft, Kaiser Permanente, major universities, and the Department of Defense.
If you'd like to build the world's best deep learning cloud, join us.
What You’ll Do
- Remotely install, upgrade, operate and maintain bare-metal Kubernetes clusters (up to thousands of nodes each)
- Handle cluster degradation, recovery and resizing using our fleet management tooling
- Perform out-of-hours on-call response for critical incidents as part of a well-balanced on-call rotation
- Work on improving our tooling, automation, and processes, for both daily operations, alerting, and incident response
- Dive into systems at a low level to solve unique cluster problems and write up your findings
- Assist customers with high-level Kubernetes questions and integration with applications, storage and authentication
- Assist with initial cluster build-outs and validation to help identify failed hardware before customer delivery
- Work closely with our HPC Ops and Datacenter Ops teams on issues that require lower-level expertise or cross-functional solutions
- Mentor and assist less-experienced team members
- Have a voice in our product direction and help us think about how to minimize operational costs and complexity
You
- Are an experienced operations engineer, SRE, sysadmin or similar with a deep knowledge of running Linux clusters and systems
- Are very familiar with running on bare-metal (including knowledge of BMCs, kernel drivers, PXE, RAID, VLANs, hypervisors)
- Have a good understanding of containers, virtualisation, and the mechanisms underpinning them
- Have a good understanding of daily operation, bug-fixing and maintenance of Kubernetes
- Have experience in an on-call environment and with incident response
- Can perform incident post-mortems and develop procedures and tooling to prevent root causes from reoccurring
- Have an excellent ability to learn on-the-fly and adapt to solve problems
- Are able to work either independently with limited direction, or as part of a team
- Are able to work with customers during incidents either via tickets, live messaging, or as part of a larger call.
Nice to Have
- Deep Kubernetes experience
- Experience with user-level restrictions and hardening (e.g. AppArmor)
- Experience with network engineering
- Experience with HPC clusters, environments & tooling
- Experience with large-scale AI/ML training clusters
- Experience with machine learning/AI frameworks
- A passion …
This job isn't fresh anymore!
Search Fresh JobsJob Profile
Regions
Countries
401(k) plan with company match Flexible paid time off Health, dental, and vision coverage
Tasks- Mentor team members
Automation Containers Deep Learning Incident Response Kubernetes Virtualization
Timezones
Remote Jobs in North America
Remote Jobs in Europe
Remote Jobs in Asia/Pacific
Remote Jobs in South America
Remote Jobs in Africa
Remote Jobs in Middle East
Full Time Remote Jobs
Part Time Remote Jobs
Internship Remote Jobs
Contract Remote Jobs
Temporary Remote Jobs
Freelance Remote Jobs
Mid-Level Remote Jobs
Senior-Level Remote Jobs
Entry-Level Remote Jobs
Exec-Level Remote Jobs
Lead-Level Remote Jobs
Remote Contract Jobs
Remote Program Manager Jobs
Remote Engineer I Jobs
Remote Marketing Manager Jobs
Remote Spanish Jobs
Remote Technician Jobs
Remote Data Scientist Jobs
Remote Advisor Jobs
Remote Machine Learning Jobs
Remote Sales Manager Jobs
Remote Mobile Jobs
Remote Therapist Jobs
Remote Customer Success Jobs
Remote Inside Sales Jobs
Remote Finance Jobs
Remote Writer Jobs
Remote Customer Service Jobs
Remote Data Engineer Jobs
Remote Associate Director Jobs
Remote Associate Dir Jobs
Remote Jobs with PHP > 280K in Salary
Remote Jobs with PHP > 160K in Salary
Remote Jobs with PHP > 260K in Salary
Remote Jobs with PHP > 180K in Salary
Remote Jobs with PHP > 100K in Salary
Remote Jobs with GBP > 140K in Salary
Remote Jobs with CAD > 180K in Salary
Remote Jobs with GBP > 160K in Salary
Remote Jobs with EUR > 140K in Salary
Remote Jobs with GBP > 180K in Salary
Remote Jobs with CAD > 200K in Salary
Remote Jobs with ₱ > 40K in Salary
Remote Jobs with PLN > 100K in Salary
Remote Jobs with PLN > 80K in Salary
Remote Jobs with PLN > 40K in Salary
Remote Jobs with PLN > 60K in Salary
Remote Jobs with PLN > 120K in Salary
Remote Jobs with PLN > 200K in Salary
Remote Jobs with PLN > 160K in Salary
Remote Jobs with PLN > 180K in Salary