β’Β Β Β Β Lead
the design and implementation of chaos, failure, and performance test
frameworks targeting distributed backend systems and infrastructure.
β’Β Β Β Β Build
and maintain CI/CD pipelines in Jenkins using Groovy-based DSL and
scripted/declarative pipeline definitions.
β’Β Β Β Β Own
infrastructure-level test environments using Kubernetes β spinning up, tearing
down, and managing test clusters for reliability experiments.
β’Β Β Β Β Execute
and automate chaos engineering scenarios: network partitions, node failures,
pod evictions, resource exhaustion, and latency injection.
β’Β Β Β Β Design
and run performance and load tests to identify bottlenecks, regressions, and
capacity limits across services.
β’Β Β Β Β Develop
failure testing strategies to validate system behavior under degraded
conditions β partial failures, cascading failures, and data corruption
scenarios.
β’Β Β Β Β Define
quality metrics and SLOs for infrastructure reliability; report on test
coverage and failure patterns to engineering leadership.
β’Β Β Β Β Mentor
and guide junior QA engineers; champion a reliability-first quality culture
across the engineering org.
β’Β Β Β Β Strong
programming skills in Python for test automation, tooling, and scripting.
β’Β Β Β Β Proficiency
in Groovy, particularly for Jenkins pipeline development (scripted and
declarative).
β’Β Β Β Β Extensive
hands-on Jenkins experience β pipeline authoring, plugin management, shared
libraries, and CI/CD architecture.
β’Β Β Β Β Deep
Kubernetes knowledge β deploying workloads, managing namespaces, configuring
resource limits, and operating clusters for test environments.
β’Β Β Β Β Proven
experience in chaos engineering: fault injection, resilience testing, and tools
such as Chaos Monkey, Litmus, or Gremlin.
β’Β Β Β Β Hands-on
performance testing experience: load generation, profiling, bottleneck
analysis, and tooling (e.g., Locust, k6, Gatling, JMeter).
β’Β Β Β Β Experience
designing failure test scenarios for distributed systems β covering partial
outages, dependency failures, and data-path degradation.
β’Β Β Β Β Strong
understanding of Linux, networking, and infrastructure fundamentals.
β’Β Β Β Β Experience
with service meshes (Istio, Linkerd) for traffic shaping and fault injection.
β’Β Β Β Β Familiarity
with observability tooling: Prometheus, Grafana, Jaeger, or OpenTelemetry for
validating test outcomes.
β’Β Β Β Β Background
in SRE practices: SLOs, error budgets, incident post-mortems.
β’Β Β Β Β Experience
with cloud infrastructure testing on AWS, GCP, or Azure at the platform level.
β’Β Β Β Β Exposure
to eBPF-based tools for low-level system observability during failure tests.