Base Career helps you apply smarter for this job.
Key skills for this role
Own and improve the reliability, availability and performance of critical production services across the Runware platform
Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation
Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
, GBR
, GBR
, GBR
, GBR
, GBR
, GBR
, GBR
, GBR
Experience operating high-throughput or low-latency APIs and distributed systems
Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
Experience with RabbitMQ or other distributed messaging and queueing systems
Experience operating MySQL, Redis, ClickHouse or similar production data systems
Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
Experience building automated scaling, capacity management or self-healing systems
Generative AI API provider helping developers and businesses create image and visual media.
Visit company websiteJobs and hiring trendsFull-time
Senior
Remote
Apply faster on company sites with our extension.