Base Career helps you apply smarter for this job.
Key skills for this role
Lumbra is building Nebula, an agentic harness running as a set of microservices on managed Kubernetes, backed by managed databases, caching, and workflow orchestration, all provisioned with OpenTofu and deployed with Helm via CI/CD. We currently run on GCP but are not wed to any single provider. We're looking for an infrastructure engineer to own the reliability, scalability, and developer experience of the harness across dev, demo, and production environments.
Lumbra is building Nebula, an agentic harness running as a set of microservices on managed Kubernetes, backed by managed databases, caching, and workflow orchestration, all provisioned with OpenTofu and deployed with Helm via CI/CD. We currently run on GCP but are not wed to any single provider. We're looking for an infrastructure engineer to own the reliability, scalability, and developer experience of the harness across dev, demo, and production environments.
Author and maintain Infrastructure as Code (OpenTofu/Terraform) modules for cloud resources including networking, managed Kubernetes clusters, databases, caching, and container registries. Strong IaC skills and experience with GCP (or equivalent) are essential.
Design and manage Kubernetes cluster configurations including node pool autoscaling, workload identity, private connectivity for database access, and network policies. You need deep Kubernetes knowledge, not just manifest authoring.
Build and optimize Helm charts for a shared service template consumed by multiple services, managing environment-specific overrides across dev, demo, staging, and production. Experience with Helm inheritance patterns and chart libraries is important.
Own the CI/CD pipeline architecture : multi-stage builds, conditional triggers based on file-change detection, and deployment orchestration. You should be comfortable authoring and debugging complex pipeline configurations.
Implement and maintain the observability stack (metrics, traces, logs) across all services using Grafana, Prometheus, and OpenTelemetry. Experience instrumenting distributed systems and building actionable dashboards is needed.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Arlington, USA
Arlington, USA
Arlington, USA
Arlington, USA
Arlington, USA
Arlington, USA
Arlington, USA
Arlington, USA
Manage secrets lifecycle and credential rotation with automated syncing to Kubernetes, plus identity provider configuration. Understanding of zero-trust patterns and secrets management at scale is essential.
Configure and maintain production networking including load balancing, TLS termination, DNS, and authentication proxies. Solid networking fundamentals are a must.
Optimize the container build pipeline for speed and security: multi-stage builds, layer caching, image hardening, and size reduction for faster, safer deployments.
Continuously profile and optimize platform performance : query latency, pod startup times, resource utilization, and network throughput. You care about measurable improvements and treat sluggish infrastructure as a bug, not a tradeoff.
Maintain developer experience tooling including local development environments, task automation, and environment bootstrapping that lets engineers go from clone to running system quickly.
Building AI orchestration frameworks for secure intelligence environments.
Visit company websiteJobs and hiring trendsFull-time
Mid
Onsite
Apply faster on company sites with our extension.