+63 995 394 7258 / +63 917 132 5623 | marketing@axentra-global.com Blog Register

Site Reliability Engineering (SRE) Training Program


Duration: 5 days / 40 hours

Delivery Method: Classroom-based, Virtual Instructor Led Training

COURSE OUTLINE


·      Day 1: SRE Principles, Metrics & Incident Management

·      Day 2: Observability + Containers

o  Topic: Monitoring and Observability Deep Dive

§ Description: Covers full stack metrics (infra + app); instrumentation and business KPIs.

Ø Mode: Lecture

o  Topic: Prometheus & Grafana Lab

§ Description: Collect metrics, build dashboards, and set up custom alerts.

Ø Mode: Hands-on Lab

o  Topic: Distributed Tracing & Log Aggregation

§ Description: Trace microservices across requests; use ELK or Datadog for correlation.

Ø Mode: Live Demo

o  Topic: Containerization with Docker

§ Description: Build and run containerized apps; image layering and registry management.

Ø Mode: Hands-on Lab

 

·      Day 3: Kubernetes + Application Reliability

o  Topic: Kubernetes for SREs

§ Description: Pods, services, rolling updates, auto-healing, readiness/liveness probes.

Ø Mode: Hands-on Lab

o  Topic: Advanced Kubernetes: Helm, HPA, Node Affinity

§ Description: Automate deployments, manage scaling and placement for reliability.

Ø Mode: Hands-on Lab

o  Topic: Progressive Delivery (Blue/Green, Canary)

§ Description: Control release risk through staged deployment strategies.

Ø Mode: Tool Demo + Use Case Design

 

o  Topic: Application Reliability Patterns

§ Description: Timeouts, retries, bulkheads, fallbacks; demo using a fault-injected app.

Ø Mode: Workshop + Code Review

 

·      Day 4: Automation + Chaos Engineering

o  Topic: Automation with Python & Bash

§ Description: Write scripts for health checks, failover triggers, and Slack integrations.

Ø Mode: Hands-on Lab

o  Topic: Chaos Engineering Theory

§ Description: Understanding controlled failures, hypotheses, and safety checks.

Ø Mode: Lecture + Group Exercise

o  Topic: Chaos Tools: Chaos Mesh / Gremlin lab

§ Description: Simulate pod failures, latency injection, and auto-recovery.

Ø Mode: Hands-on Lab

o  Topic: Lab: Resilience Testing a Microservice

§ Description: Inject chaos and monitor system behavior with dashboards and alerts.

Ø Mode: Hands-on Lab

 

·      Day 5: DB, Capstone, and Final Review

o  Topic: Database Management and Failover

§ Description: Scaling, tuning, read replicas, automated failover strategies.

Ø Mode: Lecture + SQL Lab

o  Topic: Release Engineering & Deployment Governance

§ Description: How to enforce policies, approvals, and rollback strategies for safe deploys.

Ø Mode: Case Study

o  Topic: Capstone: Design a Resilient Cloud-Native System

§ Description: Teams architect a production-ready system with SLIs/SLOs, observability, CI/CD, and resilience patterns.

Ø Mode: Workshop

o  Topic: Team Presentations + Expert Review

§ Description: Feedback from instructors on system design, trade-offs, and tooling.

Ø Mode: Presentation + Group Feedback

 

OPTIONAL ADD-ONS OR CUSTOMIZATION

·      Executive Briefing Track (1-day): For CIOs and VPs — focused on reliability ROI, staffing models, OKRs, and governance.

·      Tool-Specific Deep Dives: Jenkins Advanced, Terraform Enterprise, Kubernetes Security, AWS Fault Injection Simulator.

·      Post-Training Assessment & Certification: Knowledge check, capstone grading, and certificate of completion.


REGISTER NOW

Learning Experience Survey

Learning Experience Survey

Learning Experience Survey

Learning Experience Survey

Learning Experience Survey

Learning Experience Survey

Learning Experience Survey

Learning Experience Survey