Base Career helps you apply smarter for this job.
Key skills for this role
We are seeking a Software Engineer, Infrastructure & Platform to build the systems and infrastructure that power advanced AI evaluations, including evaluations focused on autonomous model behavior, agentic systems, and loss-of-control risks.
This is a hands-on engineering role at the intersection of backend systems, infrastructure, and AI. You will build secure and reproducible environments where frontier models can interact with tools, execute code, complete complex tasks, and operate across realistic multi-step workflows, including machine learning research and engineering.
The ideal candidate has strong backend and infrastructure fundamentals, enjoys debugging complex distributed systems, and is excited to apply those skills to difficult problems in AI safety and evaluation.
We are seeking a Software Engineer, Infrastructure & Platform to build the systems and infrastructure that power advanced AI evaluations, including evaluations focused on autonomous model behavior, agentic systems, and loss-of-control risks.
This is a hands-on engineering role at the intersection of backend systems, infrastructure, and AI. You will build secure and reproducible environments where frontier models can interact with tools, execute code, complete complex tasks, and operate across realistic multi-step workflows, including machine learning research and engineering.
The ideal candidate has strong backend and infrastructure fundamentals, enjoys debugging complex distributed systems, and is excited to apply those skills to difficult problems in AI safety and evaluation.
Design and build sandboxed evaluation environments where AI models can safely execute code, use tools, interact with services, and complete complex tasks.
Build backend services and infrastructure supporting large-scale, repeatable AI and agentic evaluations.
Develop agent scaffolding and evaluation harnesses, including tool-use loops, context management, retries, state management, token budgets, and multi-agent or subagent workflows.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
New York City, USA
, USA
New York City, USA
, USA
, USA
, USA
, USA
, USA
, USA
Build systems for provisioning and orchestrating isolated environments using technologies such as Docker, Kubernetes, VMs, and cloud infrastructure.
Design secure approaches to networking, permissions, secrets, credentials, and resource isolation for model-driven environments.
Develop APIs, internal tools, and automation that allow researchers, engineers, and subject-matter experts to create and run evaluations efficiently.
Improve the reliability and reproducibility of evaluations through logging, observability, snapshotting, debugging tools, and automated testing.
Build systems capable of running thousands of evaluation tasks reliably and capturing the artifacts and telemetry needed to understand model behavior.
Partner with analysts, red teamers, and domain experts to translate complex evaluation ideas into robust technical systems.
Investigate failures across the evaluation stack and distinguish between model limitations and infrastructure, harness, or environment failures.
3–5+ years of professional software engineering experience, particularly in backend, infrastructure, platform, SRE, or distributed systems engineering.
Strong programming skills in Python and experience building production-quality software.
Experience designing and operating backend services, APIs, or distributed systems.
Hands-on experience with Docker, Kubernetes, virtual machines, or other container/orchestration technologies.
Experience working with AWS, GCP, or similar cloud infrastructure.
Strong understanding of Linux systems, networking, authentication, permissions, and infrastructure security.
Experience with infrastructure-as-code or automation tools such as Terraform.
Strong debugging skills and comfort diagnosing failures across application, infrastructure, and networking layers, especially in agentic loops.
Ability to build systems that are reproducible, observable, scalable, and secure.
Comfort working on ambiguous technical problems where the architecture and requirements may evolve quickly.
Interest in AI systems, agentic workflows, AI security, or model evaluations. Prior professional AI experience is helpful but not required.
Experience building developer platforms, CI/CD systems, test infrastructure, sandboxes, or ephemeral compute environments.
Experience with agent frameworks, LLM APIs, tool-calling systems, or AI evaluation infrastructure.
Experience designing secure execution environments for untrusted or semi-trusted code.
Background in SRE, platform engineering, cloud infrastructure, cybersecurity, or developer tooling.
Experience with distributed task execution, queues, workflow orchestration, or large-scale automated testing.
Familiarity with AI safety, adversarial testing, model evaluations, or autonomous-agent systems.
Familiarity with agentic AI fundamentals, including common harnesses, Model Context Protocol, agent benchmarks, and security risks to AI agents.
Applied AI security and threat-intelligence company serving frontier AI labs, Fortune 10 companies, and technology platforms.
Visit company websiteJobs and hiring trendsUSD 110000-160000 yearly / year
Full-time
Mid · 3+ years experience
Remote
Apply faster on company sites with our extension.