No matter how well you design a DR strategy, "does it actually work?" is a separate matter. A 5-minute RTO on paper becomes 30 minutes when real failure hits, and a system you thought was "safe because Multi-AZ" collapses the moment one AZ dies. The uncomfortable truth of software reliability is "untested recovery doesn't work." That's where Chaos Engineering comes in — a methodology that intentionally injects failures into healthy systems to expose weaknesses before real disasters strike.
AWS has productized this into two services