Senior Site Reliability Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Full Job Posting
Do you enjoy collaborating with teams to solve complex challenges?
Do you enjoy solving large scale distributed content delivery challenges?
Join our critical AI Hardware SRE Team!
The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure.
You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings.
Partner with the best
This position focuses on enhancing system reliability, scalability, and performance across high-density hardware and software infrastructure in regional data centers.
Responsibilities
include defining KPIs, proactive monitoring, automation, and urgent issue resolution.
Collaboration with teams ensures best practices, reduced downtime, and optimized systems.
The role supports seamless operations, business-critical applications, and improved user experiences through efficient, data-driven solutions.
As a Senior Site Reliability Engineer, you will be responsible for:
- Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning.
- Integrating automated workflows across diverse corporate ticketing systems to enhance resolution times for hardware and network break-fix incidents.
- Leveraging advanced AI tools and LLM-based development approaches to enhance technical execution, script creation, and comprehensive system evaluation.
- Working on cutting-edge private cloud and compute technologies to improve the availability, latency, and overall systemic health of high-density hardware environments.
- Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection tailored for bare-metal and virtualized environments.
- Participating in 24x7x365 on-call rotations, spearheading real-time incident management, and managing high-severity service disruption protocols via automated PagerDuty and Slack workflows.
- Partnering directly with third-party infrastructure vendors and coordinating on-site field technicians to facilitate uptime activities
Do what you love
To be successful in this role you will:
- Have 5+ years of relevant experience and a Bachelor's degree in Computer Science or related field
- Demonstrate exceptional proficiency in tooling and coding using languages like Python to build scalable operational tools, API integrations, and automation frameworks.
- Show hands-on experience with modern observability stacks and timeseries engines, like Prometheus, Grafana, OpenTelemetry, and Loki.
- Possess a working understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks.
- Demonstrate expertise as a primary designer for new service rollouts, establishing operational readiness criteria, telemetry baselines, and alerting thresholds.
- Demonstrate extensive experience building technical runbooks, leading complex incident response bridges, and driving comprehensive, blameless post-mortems.
- Demonstrate a proven ability to fully own ambiguous technical challenges, coordinate cross-functional teams, and drive toward production-grade solutions effectively.
About Us
At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes.
We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple: Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge.
And we're the ones who solve it.
When millions of people hit play or pay, Akamai ensures it just works.
Benefits
at Akamai: We support your health, well-being, finances, and life beyond work.
See our benefits.
FlexBase adapts to your job's needs
Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience.
It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.
Connect with us on social and see what life at Akamai is like!
About Cloud Data Services
Cloud and edge computing platform for secure digital experiences.
Visit company websiteApply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at Cloud Data Services
Network Infrastructure Engineer
, USA
Would you enjoy improving stability and safety of one of the largest global networks? Would you enjoy doing hands-on network operations work on a global scale to improve our operational efficiency? Join our critical Netw
Senior II Software Engineer - Linux Performance
, USA
Do you have a passion for making systems run blazingly fast? Does it bother you when they don't? Join our innovative team Our Linux Performance team is a specialized group that looks to improve the performance and effici
Associate Financial Analyst
, USA
Do you want to begin an exciting career in finance with mentorship from industry leaders? Are you looking to explore different facets of finance in an exciting learning environment? Become a leader through our 3-year rot
Finance Intern
, USA
Do you want to enhance your corporate finance experience with a cutting-edge technology company? Are you looking to join a motivated, curious team of rotation program mentors and industry leaders? Work with our corporate
Software Engineer
, IND
Are you a fast learner who is fascinated by technology? Shape the end-to-end experience for millions on our Next Generation portal. Join our world-class Delivery Experience Team Akamai's Control Center is our face to our
Senior Site Reliability Engineer
, IND
Do you like collaborating across teams to solve complex problems? Do you enjoy solving large scale distributed content delivery challenges? Join our highly skilled Compute Site Reliability team Our team designs, develops
Software Engineer II
, IND
Do you enjoy collaborating with teams to solve complex technical challenges? Are you passionate about cutting-edge technologies and solving system problems? Join our highly skilled Software Engineering team The team crea
Senior II Public Sector Business Development Manager
, USA
Are you passionate about building strategic partnerships that drive government innovation? Do you thrive on creating channel relationships that transform federal technology landscapes? Join our dynamic Public Sector Grow
Network Infrastructure Engineer
, USA
Senior II Software Engineer - Linux Performance
, USA
Associate Financial Analyst
, USA
Finance Intern
, USA
Software Engineer
, IND
Senior Site Reliability Engineer
, IND
Software Engineer II
, IND
Senior II Public Sector Business Development Manager
, USA