Base Career helps you apply smarter for this job.
Key skills for this role
As an Infrastructure Engineer at Reducto, you will influence every aspect of our infrastructure from the ground up. You will architect and scale resilient systems for AI and ML workloads, automate cloud infrastructure, and implement monitoring and incident response practices that set the standard for reliability. This role requires technical leadership, hands-on systems engineering, and strong collaboration with our founders and product teams as we build a company around reliability, rapid iteration, and high-impact product delivery.
Designing, building, and maintaining highly available, scalable infrastructure to support intensive AI/ML workloads and real-time model deployments.
Implementing robust monitoring, alerting, and observability systems to ensure system health, performance, and uptime across cloud and on-prem environments.
Debugging, optimizing, and automating infrastructure for fast iteration and rapid deployment cycles, focusing on both reliability and developer velocity.
Proactively identifying, investigating, and resolving incidents to minimize downtime and maintain world-class service levels for enterprise customers.
Collaborating closely with engineers, ML specialists, and founders to shape product, infrastructure, and security strategies.
Are your own worst critic—have an extremely high bar for quality and always aim for robust solutions rather than quick fixes.
Have 5+ years of hands-on experience in building or supporting production-grade infrastructure and reliability processes for high-throughput systems.
Are comfortable with Python or similar languages, and exceptional at working across cloud platforms, container orchestration (e.g., Kubernetes), networking, and storage technologies.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
San Francisco, USA
About Reducto Reducto is the agentic document platform for leading AI teams who demand enterprise performance at scale. We provide a comprehensive toolkit for working with documents the way a human would, combining custo
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
San Francisco, USA
New York City, USA
Build your own tools on the fly to diagnose, experiment, and address reliability problems—whether it's an internal dashboard or an automated remediation workflow.
Bring a quantitative, hands-on approach to system operations, automation, and continuous improvement.
Have prior experience founding a company or building products/infrastructure in early-stage, high-growth environments.
Are excited about automating incident management processes with LLMs/AI.
Are driven, ambitious, and deeply care about both technical excellence and collaborative problem-solving.
Keep up with the latest trends in cloud, observability, and SRE best practices.
Are passionate about open-source and have contributed tools or automation to reliability communities.
Have built or optimized monitoring, incident response, or high-performance computing systems for demanding AI/ML, fintech, or enterprise clients.
This is an in person role at our office in SF. We’re an early stage company which means that the role requires working hard and moving quickly. Please only apply if that excites you.
Reducto is an agentic document platform helping enterprise AI teams automate document workflows.
Visit company websiteJobs and hiring trendsUSD 150000-300000 yearly / year
Full-time
Senior · 5+ years experience
Onsite
Apply faster on company sites with our extension.