Reducing the time between incident detection, investigation, and remediation is a critical priority for organizations running production workloads on AWS. When an issue arises, on-call engineers often need to quickly diagnose the problem across application components, identify the root cause, and apply the fix, often in the middle of the night.