Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

40% Positive

Analyzed from 199 words in the discussion.

Trending Topics

#system#breakage#bit#drills#years#things#reliable#every#automation#problem

Discussion (7 Comments)Read Original on HackerNews

andaiabout 2 hours ago
Had a strange thought last year, that if a system runs too smoothly, eventually the knowledge of dealing with breakage will disappear, and when it inevitably breaks, everyone will be unprepared.

So a little bit of breakage is like a healthy "exercise."

Presumably, a functional system is supposed to have some kind of drills to fill in for that. But I have seen ones that don't!

toast043 minutes ago
Yeah, I've lived that one. When your system is falling over all the time, you get good at quickly bringing it back. When you have the first failure in 2 years and you didn't do drills or etc, it takes a lot longer to get things going again.
ggambettaabout 1 hour ago
Hence DiRT or whatever they call it these days
gatioabout 1 hour ago
Very true. If things are too reliable, systems can come to depend on them always being so so reliable... So it can actually pay off to inject transient issues deliberately.

https://netflix.github.io/chaosmonkey/ and similar can help.

That said ... most instability is introduced with normal changes, so every engineer can be a chaos monkey. ;-)

jknoepfler25 minutes ago
This is known as the "Paradox of Automation." It's a real problem for many industries.

The OG paper on this is called "The Ironies of Automation" (Bainbridge, 1983). It's a clunky but fascinating read. I give it a skim every couple years and usually come away with a bit of a fresh take on the problem.

_defabout 2 hours ago
Oh, the legendary backup restore drill? I'm sure it will happen someday ...
sixtyjabout 1 hour ago
Nice that it is cc-commons.

Unfortunately site would need some refurbishment for mobile… slide show doesn’t show, and page is floating.