Lead the design and implementation of chaos, failure, and performance test frameworks targeting distributed backend systems and infrastructure.
Build and maintain CI/CD pipelines in Jenkins using Groovy-based DSL and scripted/declarative pipeline definitions.
Own infrastructure-level test environments using Kubernetes — spinning up, tearing down, and managing test clusters for reliability experiments.
Execute and automate chaos engineering scenarios: network partitions, node failures, pod evictions, resource exhaustion, and latency injection.
Design and run performance and load tests to identify bottlenecks, regressions, and capacity limits across services.
Develop failure testing strategies to validate system behavior under degraded conditions — partial failures, cascading failures, and data corruption scenarios.
Define quality metrics and SLOs for infrastructure reliability; report on test coverage and failure patterns to engineering leadership.
Mentor and guide junior QA engineers; champion a reliability-first quality culture across the engineering org.
Required Qualifications
Strong programming skills in Python for test automation, tooling, and scripting.
Proficiency in Groovy, particularly for Jenkins pipeline development (scripted and declarative).