Base Career helps you apply smarter for this job.
Key skills for this role
Own and extend Terraform modules across AWS — ECS/Fargate or EKS, RDS, VPC networking, IAM, ALB/NLB, S3, Route 53, CloudWatch.
Manage Terraform state safely; write and review plans that reviewers can trust.
Eliminate manually created resources by bringing them under code (via terraform import, refactors, and module extraction).
Keep environments (dev, QA, production) consistent and reproducible.
Build and maintain GitHub Actions pipelines for build, test, containerization, and deployment.
Migrate legacy pipelines (Jenkins, CircleCI, GitLab CI) onto a single, maintainable platform.
Design safe deployment paths — staged rollouts, health-gated releases, fast and reliable rollback. Keep pipelines fast: caching, parallelism, test splitting, and honest gating.
Own the production change window on your shift: deploys, migrations, cutovers, and maintenance.
Run and improve observability — dashboards, SLOs, alerting that fires on real user impact rather than noise.
Debug containerized applications in production: logs, metrics, traces, resource limits, networking.
Participate in on-call rotation and incident response; drive blameless post-incident reviews. Improve cost efficiency and resource utilization across the AWS footprint.
Operate PostgreSQL in production — read query plans, identify slow queries, understand connections, locks, and replication.
Plan and execute schema migrations against live systems with minimal disruption.
Operate supporting data services: Redis/ElastiCache, message brokers, object storage.
Write Python and shell automation to remove repetitive operational work.
Build tooling that makes the wider team faster — self-service scripts, runbooks, guardrails.
Harden secrets handling, access control, and infrastructure security posture.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Bellevue, USA
Gurugram, IND
Gurugram, IND
Gurugram, IND
Bellevue, USA
Bellevue, USA
, USA
, USA
Gurugram, IND
Write clear, complete handover notes at the end of every shift. Where shifts do not overlap, your writing is how the rest of the team learns what happened.
Maintain runbooks and infrastructure documentation as systems change.
Coordinate with the wider engineering team on planned work and follow-ups.
Use modern AI tools to accelerate infrastructure work, scripting, and troubleshooting.
Apply AI-assisted practices while maintaining strong engineering, security, and review standards.
6+ years in DevOps, SRE, Platform, or Infrastructure engineering with genuine production ownership.
Strong AWS experience — container orchestration (ECS/Fargate or EKS), RDS, VPC networking, IAM, load balancing, S3, CloudWatch.
Hands-on Terraform: writing modules, managing state, reviewing plans.
Practical CI/CD experience, ideally GitHub Actions. GitLab CI, CircleCI, or Jenkins backgrounds are fine if you can migrate.
Docker, and comfort debugging containerized applications in production.
Solid Linux fundamentals and shell scripting; Python for automation.
Working PostgreSQL knowledge — query plans, slow queries, connections, locks, replication. Experience deploying and operating Django/Python and Node.js/React applications. Hands-on use of an observability platform in anger (New Relic, Datadog, Grafana, Prometheus, or similar).
Clear written English. Your handover notes are how the rest of the team learns what happened. Openness to either a normal or a night shift, and willingness to participate in an on-call rotation.
AWS (ECS / Fargate / EKS, EC2, Lambda)
VPC, Subnets, Security Groups, NAT, Peering
ALB / NLB, Route 53, CloudFront
IAM, Roles, Policies, Least-Privilege Design
RDS, ElastiCache, S3
Terraform / Infrastructure as Code
GitHub Actions
Jenkins / CircleCI / GitLab CI
Docker, Container Registries
Blue-Green & Staged Deployments, Rollback Strategies
Python, Bash / Shell Scripting
Git and Trunk-Based Workflows
New Relic / Datadog / Grafana / Prometheus
CloudWatch Metrics, Logs, Alarms
OpenTelemetry
SLOs, Error Budgets, Alert Design
Incident Management & Post-Incident Review
PostgreSQL Administration & Tuning
Schema Migrations on Live Systems
Redis / ElastiCache
Kafka / Redpanda / Amazon MSK, RabbitMQ, NATS
Secrets Management (AWS Secrets Manager, SOPS, Vault)
Network and Access Hardening
Vulnerability and Patch Management
Audit Logging
Celery, RabbitMQ, NATS, or Kafka in production.
Redis / ElastiCache operations.
Secrets management (AWS Secrets Manager, SOPS, Vault).
Experience with sharded or multi-tenant database architectures.
Healthcare or other compliance-sensitive environments (HIPAA, SOC 2).
Telephony / VoIP infrastructure exposure.
Kubernetes.
Prior experience as the sole engineer on shift, or on a night/off-hours rotation.
A builder’s instinct — you would rather codify a fix than repeat it.
Comfort with autonomy and sound judgment on when to escalate.
Careful, methodical change management on production systems.
Strong written communication; your handover notes are a first-class deliverable.
Curiosity about how systems actually behave, not just how they are supposed to. Ownership, follow-through, and continuous learning.
Real ownership of infrastructure, not a ticket queue.
Flexibility on shift, and a protected change window for high-impact work.
Modern stack: AWS, Terraform, GitHub Actions, Docker, PostgreSQL, Kafka-compatible messaging.
Direct impact on the reliability of a platform used by thousands of healthcare practices.
Dental software company providing an all-in-one operations, analytics, communications, payments, and marketing platform for dental practices.
Visit company websiteJobs and hiring trendsFull-time
Senior · 6+ years experience
Onsite
Apply faster on company sites with our extension.