Site Reliability Engineering (SRE) Training Program
Duration: 5 days / 40 hours
Delivery Method: Classroom-based, Virtual Instructor Led Training
· Day 1: SRE Principles, Metrics & Incident Management
· Day 2: Observability + Containers
o Topic: Monitoring and Observability Deep Dive
§ Description: Covers full stack metrics (infra + app); instrumentation and business KPIs.
Ø Mode: Lecture
o Topic: Prometheus & Grafana Lab
§ Description: Collect metrics, build dashboards, and set up custom alerts.
Ø Mode: Hands-on Lab
o Topic: Distributed Tracing & Log Aggregation
§ Description: Trace microservices across requests; use ELK or Datadog for correlation.
Ø Mode: Live Demo
o Topic: Containerization with Docker
§ Description: Build and run containerized apps; image layering and registry management.
Ø Mode: Hands-on Lab
· Day 3: Kubernetes + Application Reliability
o Topic: Kubernetes for SREs
§ Description: Pods, services, rolling updates, auto-healing, readiness/liveness probes.
Ø Mode: Hands-on Lab
o Topic: Advanced Kubernetes: Helm, HPA, Node Affinity
§ Description: Automate deployments, manage scaling and placement for reliability.
Ø Mode: Hands-on Lab
o Topic: Progressive Delivery (Blue/Green, Canary)
§ Description: Control release risk through staged deployment strategies.
Ø Mode: Tool Demo + Use Case Design
o Topic: Application Reliability Patterns
§ Description: Timeouts, retries, bulkheads, fallbacks; demo using a fault-injected app.
Ø Mode: Workshop + Code Review
· Day 4: Automation + Chaos Engineering
o Topic: Automation with Python & Bash
§ Description: Write scripts for health checks, failover triggers, and Slack integrations.
Ø Mode: Hands-on Lab
o Topic: Chaos Engineering Theory
§ Description: Understanding controlled failures, hypotheses, and safety checks.
Ø Mode: Lecture + Group Exercise
o Topic: Chaos Tools: Chaos Mesh / Gremlin lab
§ Description: Simulate pod failures, latency injection, and auto-recovery.
Ø Mode: Hands-on Lab
o Topic: Lab: Resilience Testing a Microservice
§ Description: Inject chaos and monitor system behavior with dashboards and alerts.
Ø Mode: Hands-on Lab
· Day 5: DB, Capstone, and Final Review
o Topic: Database Management and Failover
§ Description: Scaling, tuning, read replicas, automated failover strategies.
Ø Mode: Lecture + SQL Lab
o Topic: Release Engineering & Deployment Governance
§ Description: How to enforce policies, approvals, and rollback strategies for safe deploys.
Ø Mode: Case Study
o Topic: Capstone: Design a Resilient Cloud-Native System
§ Description: Teams architect a production-ready system with SLIs/SLOs, observability, CI/CD, and resilience patterns.
Ø Mode: Workshop
o Topic: Team Presentations + Expert Review
§ Description: Feedback from instructors on system design, trade-offs, and tooling.
Ø Mode: Presentation + Group Feedback
OPTIONAL ADD-ONS OR CUSTOMIZATION
· Executive Briefing Track (1-day): For CIOs and VPs — focused on reliability ROI, staffing models, OKRs, and governance.
· Tool-Specific Deep Dives: Jenkins Advanced, Terraform Enterprise, Kubernetes Security, AWS Fault Injection Simulator.
· Post-Training Assessment & Certification: Knowledge check, capstone grading, and certificate of completion.
REGISTER NOW
Copyright © 2026 Axentra Global Inc.