Victor Certuche is an entrepreneur, educator and transformation leader with 30+ years in business and technology who helps organizations turn complex challenges into real results.
Site Reliability Engineering (SRE) Foundation Overview
Explores the discipline of applying software engineering practices to infrastructure and operations problems to build scalable, reliable software services.
What we cover
Module 1: SRE Principles & Practices Origins of SRE: Google’s operational paradigm and foundational literature. SRE vs. DevOps: SRE as a concrete implementation of DevOps pillars (silo reduction, accepting failure, gradual change, tooling/automation, and measurement). Core Principles: Treating operations as a software problem, managing to Service Levels, eliminating toil, automating repetitive work, reducing the cost of failure, and fostering shared ownership between Dev and Ops.
Module 2: Service Level Objectives (SLOs) & Error Budgets Service Level Terminology: Distinguishing between SLAs (contractual), SLOs (operational targets), and SLIs (quantitative metrics). The Error Budget Concept: Balancing release velocity with service stability; why 100% reliability is the wrong target. Error Budget Policies: Defining actions when budgets are exhausted (e.g., deployment freezes, prioritizing stability backlogs). The VALET Framework: Setting SLO dimensions across Volume, Availability, Latency, Errors, and Tickets.
Module 3: Reducing Toil Defining Toil: Work that is manual, repetitive, automatable, tactical, devoid of enduring value, and scales linearly with service size. The Cost of Toil: Individual burnout, career stagnation, delivery bottlenecks, and "Engineering Bankruptcy". The 50% Rule: Capping operational toil at 50% of an engineer's time to preserve capacity for engineering projects. Toil Elimination Strategies: Internal and external automation, self-service portals, and service enhancements.
Module 4: Monitoring & Service Level Indicators (SLIs) Defining SLIs: Calculating success ratios (Good EventsTotal Events) across relevant time windows. Monitoring vs. Observability: Active health checks (monitoring) vs. inferring internal system states from external outputs (observability). Telemetry Pillars: Distributed tracing, structured event logging, metric collection, and application instrumentation. Modern Tooling Stack: Prometheus, Grafana, Logstash, Splunk, and Catchpoint.
Module 5: SRE Tools & Automation Automation as a Force Multiplier: Consistency, speed, and time savings. SRE-Led Service Automation: Shifting operations left into CI/CD pipelines; Infrastructure as Code (Terraform, CloudFormation) and Configuration as Code (Ansible, Puppet, Docker). Production Validation: In-production functional/non-functional testing, versioned and signed artifacts (Nexus, Artifactory), and multi-criteria alerting. Automated Scalability & Resiliency: Auto-scaling groups, Kubernetes orchestration, and runtime self-protection. Automation Maturity Hierarchy: Progressing from manual intervention to self-healing platforms.
Module 6: Anti-Fragility & Learning from Failure Failure Paradigms: Treating failure as an opportunity to build resilience. Core Reliability Metrics: MTTD (Mean Time to Detect), MTTR (Mean Time to Recover), MTRS (Mean Time to Restore Service), and RPO (Recovery Point Objective). Organizational Culture: The Westrum Organizational Typology (Pathological vs. Bureaucratic vs. Generative cultures). Resilience Engineering: Conducting operational fire drills, implementing Chaos Engineering (e.g., Netflix Simian Army, Chaos Monkey), and minimizing outage blast radiuses.
Module 7: Organizational Impact of SRE Business Drivers: Protecting revenue, preserving brand reputation, and managing scale. SRE Team Topologies: Consulting, Embedded, Platform, and Sliced team models. Sustainable Operations: Sustainable on-call rotations, the 25% on-call rule, and automating incident management. Blameless Post-Mortems: Defining post-mortem triggers, timeline reconstruction, root cause analysis without blame, and actionable remediation items.
Module 8: SRE, Other Frameworks, and Emerging Trends Framework Coexistence: Integrating SRE with Agile, DevOps, Continuous Delivery, and ITIL 4 Service Management. Emerging Reliability Specializations: Network Reliability Engineering (NRE) Database Reliability Engineering (DBRE) Customer Reliability Engineering (CRE) Heritage Reliability Engineering (HRE) for legacy systems The Future of Observability: Consolidating APM, metrics, and logs into unified observability platforms.
Every session is customized to your organization's needs. The final agenda is confirmed with your facilitator before delivery.
Explores the discipline of applying software engineering practices to infrastructure and operations problems to build scalable, reliable software services.
Confirm the course level, prerequisites and organizational goals with the facilitator before booking.
Instructor-led training with Victor Certuche
Course scope and delivery plan confirmed after a needs analysis
Live online: instructor-led, fully interactive, no travel
In-person: at your offices or a venue of your choice
Indicative duration: 2 days
Final duration, delivery arrangements and pricing are confirmed in a proposal after a needs analysis.
Pricing depends on delivery mode, group size and how much customization you want. Tell us the shape and we will quote it.