Site Reliability Engineer III
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
The position is Site Reliability Engineer III supporting Java-based production systems running on Kubernetes in cloud environments.
The role focuses on application-level reliability, JVM troubleshooting, production incident management, and collaboration with development teams.
The role is not focused on infrastructure provisioning or CloudOps.
Key Skills for This Role
Full Job Posting
About Candescent
Candescent develops digital banking, account opening, and branch solutions for financial institutions.
Its platform uses human-centered design, data, automation, cloud innovation, and an API-first architecture.
Position Overview
The position is Site Reliability Engineer III supporting Java-based production systems running on Kubernetes in cloud environments.
The role focuses on application-level reliability, JVM troubleshooting, production incident management, and collaboration with development teams.
The role is not focused on infrastructure provisioning or CloudOps.
Experience and Required Skills
- The role requires 6–9 years of experience.
- Candidates need strong hands-on experience supporting Java applications in production.
- Candidates need deep knowledge of JVM memory management, garbage collection, OOM analysis, thread dumps, and performance analysis.
- Candidates need experience with incident response, production troubleshooting, Kubernetes application operations, and observability.
- Candidates need knowledge of SLIs, SLOs, reliability-driven operations, deployment strategies, scripting, application architecture, and service dependencies.
- Strong collaboration, communication, accountability, and judgment during high-pressure incidents are required.
Key Responsibilities
- Operate Java applications on Kubernetes and GKE and troubleshoot issues using logs, metrics, traces, heap dumps, and thread dumps.
- Participate in incident response, root cause analysis, and blameless postmortems.
- Analyze JVM behavior and recommend performance improvements.
- Define SLIs, SLOs, alerts, and dashboards.
- Support deployments, rollbacks, and runtime configuration changes.
- Identify reliability, performance, and scalability gaps in application behavior.
- Automate operational tasks and improve runbooks, operational readiness, and on-call effectiveness.
- Promote shift-left reliability practices with development teams.
Cloud and Platform Exposure
- Experience with applications deployed on Kubernetes in GCP and familiarity with GKE from an application operations perspective are required.
- Understanding of networking, IAM, storage, and compute constructs relevant to application behavior is required.
Good-to-Have Skills
- CI/CD pipeline exposure, including GitHub Actions or Jenkins, is beneficial.
- Familiarity with GitOps practices is beneficial.
- Experience supporting cloud migrations or modernization initiatives is beneficial.
- Exposure to platform or infrastructure concepts supporting application workloads is beneficial.
Values
The team values ownership, reliability-first thinking, curiosity, collaboration, automation, continuous improvement, and clear incident communication.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career