This is often referred to as the “Mean Time Between Failures (MTBF)” in the context of Site Reliability Engineering. It’s a somewhat counterintuitive concept that highlights the fact that failures or …
At a prior role, I managed just under 20 AWS accounts. Their uses varied from production workloads, to dedicated CI/CD environments, sandboxed areas for our engineers to experiment in, log aggregation, the list goes on.
I’m in the middle of decommissioning a service at the moment and I have to do the typical process of performing final snapshots and cleaning up extant resources. It’s incredibly tedious, so let’s walk …
If you’ve spent any time in an ops-related position, you’ve had to add these records when using custom domain names while integrating with some 3rd-party service. Especially, if you’ve ever wanted to …