Site Reliability Engineer II [Performance - Chaos]
As a Site Reliability chaos Engineer, your role is to provide reliability engineering services through chaos engineering, and performance engineering techniques. Using monitoring, fault-injection, and performance tools, you will deliver detailed feedback to product owners and development teams. You will collaborate with cross-functional teams to design, build, automate, and maintain scalable, resilient infrastructure. Your responsibilities will include ensuring high availability, monitoring system performance, executing chaos experiments, and aiding support staff with resolving incidents. This role requires a strong background in scripting, cloud platforms, and a passion for optimizing operational efficiency and system resilience. You will use Site Reliability Engineering chaos practices to deliver a seamless user experience.
Responsibilities:
- Design and develop custom, automated fault-injection playbooks to test system recovery mechanisms under pressure.
- Execute end-to-end chaos experiments (such as network latency, container crashes, and region failovers) across non-production environments.
- Analyze infrastructure bottlenecks and application dependencies exposed during simulated system failures.
- Draft comprehensive post-mortem reports and present findings to software development teams to guide code-hardening efforts.
- Construct custom dashboards in Grafana and Prometheus to visualize real-time system degradation during chaos testing.
- Implement automated chaos tests directly into active CI/CD deployment pipelines.
- Perform extensive performance and load testing using JMeter to baseline system stability before injecting faults.
- Configure advanced Dynatrace alerting policies to ensure simulated outages are immediately and accurately detected by monitoring tools.
- Simulate full Disaster Recovery (DR) and business continuity scenarios to verify the reliability of backup and failover procedures.
- Partner with software engineers to refactor application code for better fault tolerance and self-healing capabilities.
- Collaborate with the broader operations team to define system boundaries and minimize risk when executing testing windows.
- Build clean, modular automation scripts in Python, Shell, or JavaScript to provision test environments and schedule recurring chaos runs.
“Quest is a very patient centric company; we’re looking to raise the quality of healthcare through diagnostic and digital insights. You will get lots of exposure to different people and geographies.”
- Megha Kandagal, Analyst, Data Quality
Submit your resume
Submit your updated resume to us via email at HTASIndiaCareers@questdiagnostics.com. Our team will process your request and contact you about appropriate vacancies.
- Software Engineer I - CMS Hyderabad, India 07/24/2026
- Lead Software Engineer - Site Reliability Engineering Hyderabad, India 07/24/2026
- Site Reliability Engineer II [Performance - Chaos] Hyderabad, India 07/24/2026
No jobs have been viewed recently.
No jobs have been saved.