Base Career helps you apply smarter for this job.
Key skills for this role
This is a systems operations support role focused on deep-diving into escalated infrastructure issues that go beyond frontline triage. You’ll be the engineering resource our L1 support team relies on when tickets become complex, investigating and resolving issues across the full infrastructure stack—including hardware, BIOS and firmware, networking, Ubuntu, Docker, NVIDIA CUDA and GPUs, and KVM-based virtual machines.
You’ll own complex escalations end-to-end: gathering evidence, reproducing issues, identifying the root cause, proposing solutions, and working with the appropriate teams to bring each issue to resolution. The best engineers in this role don’t just resolve individual tickets—they identify recurring patterns, improve operational tooling, and build runbooks that prevent future incidents. You’ll collaborate directly with the engineering and host support teams on systemic infrastructure issues.
Strong Linux systems knowledge, technical depth, and support experience are the primary requirements. You should be comfortable working autonomously in Ubuntu environments, troubleshooting hardware, networking, containers, virtual machines, and GPU workloads, and clearly communicating your findings and proposed solutions to both technical and non-technical audiences.
Vast.ai users or hosts strongly preferred.
Vast.ai 's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing — reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation.
We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state-of-the-art AI systems while collaborating with a globally distributed team.
This is a systems operations support role focused on deep-diving into escalated infrastructure issues that go beyond frontline triage. You’ll be the engineering resource our L1 support team relies on when tickets become complex, investigating and resolving issues across the full infrastructure stack—including hardware, BIOS and firmware, networking, Ubuntu, Docker, NVIDIA CUDA and GPUs, and KVM-based virtual machines.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
San Francisco, USA
San Francisco, USA
San Francisco, USA
Los Angeles, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
Los Angeles, USA
San Francisco, USA
You’ll own complex escalations end-to-end: gathering evidence, reproducing issues, identifying the root cause, proposing solutions, and working with the appropriate teams to bring each issue to resolution. The best engineers in this role don’t just resolve individual tickets—they identify recurring patterns, improve operational tooling, and build runbooks that prevent future incidents. You’ll collaborate directly with the engineering and host support teams on systemic infrastructure issues.
Strong Linux systems knowledge, technical depth, and support experience are the primary requirements. You should be comfortable working autonomously in Ubuntu environments, troubleshooting hardware, networking, containers, virtual machines, and GPU workloads, and clearly communicating your findings and proposed solutions to both technical and non-technical audiences.
Vast.ai users or hosts strongly preferred.
Experienced with Linux, especially Ubuntu, and comfortable troubleshooting from the command line
Someone who enjoys debugging difficult problems and fixing broken systems
Methodical and focused on finding root causes, not just temporary fixes
Able to manage complex tickets independently
A clear written communicator with an interest in AI infrastructure and GPU computing
Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU-accelerated containers
Monitoring and observability experience (Prometheus, Grafana)
Relevant certifications: RHCSA, CompTIA Linux+, or similar
Knowledge of the Vast.ai platform as a client or infrastructure supplier
After you submit your application, our technical team will review your experience and qualifications. Selected candidates will proceed through the following stages:
15 minutes — Initial Screening (Virtual): A brief conversation about your background, availability, and interest in the role
45 minutes — Experience Interview (Virtual): An introduction to Vast.ai and a deeper discussion of your technical and support experience
2 hours — Meet and Greet and Technical Assessment (On-site): Meet the team and complete an LLM-assisted Linux systems operations assessment
Verified company details for this employer are not available yet.
USD 90000-160000 yearly / year
Full-time
Mid
Onsite
Apply faster on company sites with our extension.