FreshRemote.Work

HPC Support Engineering Manager

Remote, USA

Lambda's GPU cloud is used by deep learning engineers at Stanford, Berkeley, and Carnegie Mellon. Lambda's on-prem systems power research and engineering at Intel, Microsoft, Kaiser Permanente, major universities, and the Department of Defense.

If you'd like to build the world's best deep learning cloud, join us.

About the role 

At Lambda Labs, we are seeking a driven and experienced Manager of High Power Computer (HPC) Support Operations. As the HPC Support Operations leader at Lambda, you will play a pivotal role in providing design feedback on HPC solutions as well as ensure the highest level of customer satisfaction by responding to and resolving technical issues.  In addition to leading, developing, and mentoring a team of  HPC Support Engineers, you will also engage product, engineering, and sales teams to provide input into solution and product development.

Must be flexible in working nights and weekends as needed, as well as maintaining an on-call schedule.  This position reports to the Director of Customer Support.

What You'll Do

  • Ensure escalations are handled appropriately and consistently across the team.
  • Collaborate with the Director of Customer Support to support the development and coaching of the HPC Support team, ensuring the team continues in their technical growth and consistently delivers outstanding customer experiences.
  • Stay updated on the latest HPC and Nvidia technologies and provide recommendations based on thorough research and knowledge.
  • Support the HPC Support team by selecting, participating in, and leading training sessions, team meetings, and addressing roadblocks.
  • Ensure that departmental policies, procedures, and documentation accurately reflect best practices, making necessary changes or modifications as needed.
  • Provide thought leadership in the evolution of product development based on experiences from customer deployments and field installations.
  • Review, develop, and distribute support metrics to track team performance and customer satisfaction, constantly seeking opportunities for improvement in support processes and practices.
  • Assist in developing workflows and procedures for the team based on industry-standard frameworks.
  • Lead your team to develop tools that assist in the troubleshooting and resolution of technical issues encountered.
  • Manage team on-call schedule and duties.
  • Conduct performance reviews for members of the HPC Support team.
  • Lead by example, actively engaging in resolving customer cases while maintaining the necessary technical knowledge to function effectively as a team member.

You

  • Proven experience in a technical leadership role, preferably within HPC or AI industry
  • Strong knowledge of GPU InfiniBand HPC clusters, including hardware, software, and networking components
  • Advanced knowledge of Linux administration and troubleshooting
  • Have excellent leadership and team management skills with the ability to motivate and develop a high-performing team
  • Exceptional customer service and communication skills, with the ability to interact effectively with internal and external customers and stakeholders
  • Strong problem solving and analytical skills, with a proactive approach to identifying and resolving technical issues
  • You are action-oriented, humble, have a strong willingness to learn, and serve the team members you lead

Nice to Have

  • Advanced degree in related field
  • Certifications in HPC, network, or related technologies
  • Experience working with AI startups and large enterprises

Salary Range Information 

Based on market data and other factors, the salary range for this position is $170K - $210K/yr.  However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • We offer generous cash & equity compensation
  • Investors include Gradient Ventures, Google’s AI-focused venture fund
  • We are experiencing extremely high demand for our systems, with quarter over quarter, year over year profitability
  • Our research papers have been accepted into top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
  • We have a wildly talented team of 300, and growing fast
  • Health, dental, and vision coverage for you and your dependents
  • Commuter/Work from home stipends for select roles
  • 401k Plan with 2% company match
  • Flexible Paid Time Off Plan that we all actually use

A Final Note:

You do not need to match all of the listed expectations to apply for this position. We are committed to building a team with a variety of backgrounds, experiences, and skills.

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Apply

Job Profile

Regions

North America

Countries

United States

Restrictions

Must work nights and weekends On-call schedule required

Benefits/Perks

Equity Compensation Flexible paid time off Health, dental, and vision coverage Vision coverage

Tasks
  • Collaborate with product and engineering teams
  • Develop training sessions
  • Documentation
  • Improve support processes
  • Lead and mentor support team
  • Manage HPC support operations
  • Resolve technical issues
  • Track team performance
  • Troubleshooting
Skills

AI Analytical Customer service Customer Support Deep Learning Documentation GPU HPC InfiniBand Linux Linux administration Machine Learning Networking Problem-solving Product Development Support Metrics Team Management Technical Leadership Training Troubleshooting Workflow Development

Education

Advanced degree Bachelor's degree

Certifications

HPC Certification Network Certification Related Technology Certification

Timezones

America/Anchorage America/Chicago America/Denver America/Los_Angeles America/New_York Pacific/Honolulu UTC-10 UTC-5 UTC-6 UTC-7 UTC-8 UTC-9