Base Career helps you apply smarter for this job.
Key skills for this role
Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform.
We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters.
We’re a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we’re building the next generation of enterprise AI and we’re looking for exceptional people to help us scale.
You're a solid engineer with 2-4 years of experience building backend systems and platform infrastructure. You write clean, well-abstracted code with proper design patterns and comprehensive test coverage. You're comfortable working on both the Compute Platform (multi-cloud orchestration, resource management) and Inference Platform (model serving, autoscaling) under the guidance of senior engineers and platform leads.
You have strong proficiency in Golang and understand how to build maintainable, production-grade distributed systems. You take pride in code quality, enjoy collaborating on low-level designs, and are eager to learn from experienced engineers while contributing meaningfully to critical infrastructure components.
You're product-minded you understand how your technical decisions impact developers using AION's platform and think about the end-to-end user experience. You're a team player comfortable wearing multiple hats one day you're building product features, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
St. Louis, USA
St. Louis, USA
London, GBR
London, GBR
London, GBR
London, GBR
London, GBR
, GBR
Build and maintain platform services across AION's Compute and Inference platforms, working closely with senior engineers and platform leads
Implement features for multi-cloud orchestration, resource scheduling, model deployment pipelines, and autoscaling systems
Write well-maintained, production-grade code with proper abstractions, design patterns, and comprehensive test coverage
Contribute to low-level design (LLD) including service APIs, database schema design, data models, and component interactions
Collaborate with senior engineers on high-level design discussions, providing implementation perspectives and feasibility inputs
Develop RESTful APIs and gRPC services for platform control planes, resource management, and inference serving
Design and implement database schemas for storing platform state, resource metadata, billing data, and observability metrics
Work with distributed storage systems, message queues (Kafka, RabbitMQ), and databases (PostgreSQL, Redis) to build reliable platform components
Build event-driven architectures for asynchronous processing, job scheduling, and platform automation
Implement monitoring, logging, and alerting for platform services to ensure production reliability
Write comprehensive unit tests, integration tests, and end-to-end tests to ensure code reliability
Participate in code reviews, providing constructive feedback and learning from senior engineers' perspectives
Refactor existing code to improve maintainability, performance, and scalability
Document design decisions, API specifications, and operational runbooks for platform services
Debug production issues and contribute to incident response and post-mortems
Having expertise in one or more of these specializations is highly desired:
HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration
Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks
Systems-Level Programming: Experience with low-level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm-d, system agents)
ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine-tuning workflows, and Jupyter notebook orchestration systems
Enterprise Deployment: Experience platformizing and packaging software for on-premises deployments or customer VPC installations with emphasis on security, compliance, and operational simplicity
Founder-level ownership and bias for action.
Strong strategic thinking and ability to connect technical decisions to business impact.
Excellent communication and mentoring skills.
Thrives in ambiguity, fast-paced environments, and early-stage startup culture.
Work directly with high-pedigree founders shaping technical and product strategy.
Build infrastructure powering the future of AI compute globally.
Significant ownership and impact with equity reflective of your contributions.
Competitive compensation, flexible work options, and wellness benefits.
Decentralized AI cloud platform for high-performance computing.
Visit company websiteJobs and hiring trendsFull-time
Mid · 2+ years experience
Hybrid
Apply faster on company sites with our extension.