Introduction
Cloud-based reliability engineering
()
1. How to Design, Build, Operate, and Stress-Test Highly Reliable Systems
Learning objectives
()
Defining resilience, reliability, engineering, and engineering excellence
()
Ensuring engineering excellence in the cloud: Why your business can’t succeed without it
()
Understanding how to design and build resilient and reliable systems
()
Understanding how to test the resilience of your applications
()
Responding to and mitigating potential issues
()
Understanding how to leverage artificial intelligence (AI) and large language models (LLM)
()
Lesson 1 review and an exercise
()
2. Defining Engineering Strategies for Building Resilient, Available, and Scalable Systems
Learning objectives
()
Understand the foundational concepts of reliability, such as fault tolerance, high availability, scalability, and recovery
()
Choosing between alternatives for uptime and architecture design
()
Implementing service level objectives (SLO) and service level indicators (SLI) as performance measurements
()
Exploring immutable infrastructure, containerization, and event-driven architecture
()
Validating application and infrastructure resilience with chaos engineering and other modern techniques
()
Lesson 2 review and an exercise
()
3. The Power of Artificial Intelligence, Value Streams, and Cloud Reliability Engineering (CRE)
Learning objectives
()
Understanding foundational AI components
()
Applying ML and GenAI to CRE
()
Incorporating value streams and the CRE strategy
()
Fostering a culture of innovation: leadership, ownership, and fast decision-making
()
Lesson 3 review and an exercise
()
4. Leveraging Observability, Monitoring, and Reliability Metrics
Learning objectives
()
Defining observability and monitoring
()
Deploying a 10-step process to create effective monitoring
()
Surveying monitoring and alerting tools from leading cloud providers
()
Recognizing and proactively mitigating known service disruptions
()
Establishing objectives and key results (OKRs)
()
Lesson 4 review and an exercise
()
5. CRE Tooling and Chaos Engineering
Learning objectives
()
Distributing load with autoscaling and load balancing
()
Enabling automatic failovers for high availability
()
Implementing continuous deployments with rollback strategies
()
Leveraging chaos engineering for resilience testing
()
Lesson 5 review and an exercise
()
6. Incident Response for Fast Recovery
Learning objectives
()
Understanding incident response foundational concepts
()
Implementing a structured approach to incident response and CRE tools
()
Understanding incident handling in CRE
()
Defining time to detect (TTD) and time to recover (TTR)
()
Understanding playbooks and runbooks
()
Lesson 6 review and an exercise
()
7. Operational Excellence and Change Management
Learning objectives
()
Defining operational excellence in CRE
()
Identifying processes, people, and tools for operational excellence
()
Establishing key performance indicators
()
Understanding root cause analysis (RCA) and correction of error (CoE) form
()
Identifying tools for operational excellence assessments
()
Lesson 7 review and an exercise
()
Conclusion
Summary and next steps
()