Base Career helps you apply smarter for this job.
Key skills for this role
Current Tech Stack & Infrastructure Cloud Infrastructure: AWS (ElastiCache, EKS, RDS Aurora MySQL) Orchestration: Kubernetes on EKS with Karpenter for node management Infrastructure as Code: Terraform (multiple repositories, experiencing drift and collaboration challenges) Monitoring & Alerting: Datadog for monitoring, alerting, incident management, and runbooks CI/CD: GitHub Actions with Atlantis for infrastructure PRs, Bitbucket Pipelines for Helm deployments, Rundeck for scripted operations GitOps: ArgoCD for Kubernetes workloads Data Infrastructure: Transitioning to segmented data stores with ClickHouse, MySQL Aurora, and Pulsar for event streaming APM: Limited Datadog APM usage due to cost ($45 per host/month, ~10 hosts) Additional Monitoring: Started Prometheus clusters within EKS for more verbose metrics alongside Datadog Scale & Workload Handling hundreds of thousands of events per second via API (mix of synchronous and asynchronous processing) EC2 instances: M5/M6 xlarge instances, scaling between 15-100 instances per environment depending on load Currently transitioning workloads from EC2 to Kubernetes (K8s workload still smaller than core application) Real-time workloads requiring minimal downtime (minutes not hours for maintenance windows) Key Operational Challenges Infrastructure as Code: Terraform has become unmanageable due to drift reconciliation and multi-person collaboration issues Legacy Environments: Some environments set up entirely manually, never in Terraform, with major operational challenges to migrate while keeping them online Terraform Migration: Proven migration patterns in test environments, but moving to production challenging due to real-time workload requirements Alert Management: Receiving too many alerts, need prioritization and structured approach to reduce noise and recategorize/adjust thresholds Alert Distribution: Historically all alerts went to one tech ops team instead of being distributed to five different dev teams; working to shift left closer to developers Service Catalog: Still defining service catalog and ownership model Runbooks: Exist in Datadog but need more polish and structure Database Scaling: Operational challenges with database scaling being addressed through data store segmentation Monitoring Costs: Datadog is expensive; moved from CloudWatch ~8 years ago due to cost Manual Changes: Infrastructure changes often implemented manually first, then imported to Terraform and rolled out to other environments Top 5 Required Skill Sets 1. Strong Terraform and CI/CD process expertise 2.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
breaking down alerts to identify critical vs. non-critical and handling frequent alarms (Datadog-specific experience preferred) 3.
SLA/SLI definition experience (team currently being asked to do this without prior experience) 4.
AWS infrastructure expertise, DevOps-heavy background 5.
Monitoring and alerting experience (Datadog preferred, though Prometheus/Grafana experience also valuable)
Full-stack software engineering and digital transformation solutions provider.
Visit company websiteJobs and hiring trendsFull-time
Senior · 7+ years experience
Onsite
Apply faster on company sites with our extension.