Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of AWS Cloud Operations
- Defining operational roles and responsibilities in a cloud environment
- Understanding AWS account structures, organizations, and multi-account strategies
- Core operational services: CloudWatch, CloudTrail, and AWS Config
Infrastructure as Code and Provisioning
- Core principles of IaC and the benefits of immutable infrastructure
- Provisioning resources using Terraform and AWS CloudFormation
- Managing state, modules, and the promotion of environments
CI/CD and Deployment Strategies
- Designing CI/CD pipelines optimized for cloud-native applications
- Implementing Blue/Green, canary, and rolling deployment methods
- Automating rollbacks, health checks, and release validation processes
Monitoring, Observability, and Alerting
- Handling metrics, logs, and traces: shipping, storing, and analyzing data
- Leveraging CloudWatch, X-Ray, and third-party observability tools
- Establishing SLOs/SLIs, alerting policies, and on-call procedures
Security Operations and Identity Management
- Applying IAM best practices, least privilege principles, and cross-account access
- Managing secrets with KMS and secure parameter stores
- Operational security: patching strategies, vulnerability scanning, and maintaining audit trails
Resilience, Backup, and Disaster Recovery
- Architecting for fault tolerance and high availability
- Developing backup strategies, automating snapshots, and defining restore procedures
- Planning for disaster recovery and creating effective runbooks
Cost Optimization and Governance
- Enhancing cost visibility through billing, tagging, and cost allocation strategies
- Rightsizing resources, utilizing reserved instances/savings plans, and enforcing budget controls
- Governance: implementing policies, guardrails, and compliance automation
Containers, Serverless, and Runtime Operations
- Operational best practices for ECS, EKS, and Lambda
- Managing service discovery, autoscaling, and resource limitations
- Logging, tracing, and debugging containerized workloads
Incident Response, Playbooks, and Chaos Engineering
- Runbook-driven incident response and conducting postmortems
- Automating remediation and implementing self-healing patterns
- Introduction to chaos engineering for validating system resilience
Hands-on Workshop: Operating a Sample Workload
- Deploying a sample application using IaC and a CI/CD pipeline
- Implementing monitoring, alerts, and automated remediation scripts
- Simulating incidents and practicing runbook-based response procedures
Summary and Future Steps
Requirements
- A fundamental grasp of cloud computing concepts and networking principles
- Proficiency with the Linux command line and scripting languages
- Practical experience with source control (Git) and foundational CI/CD concepts
Target Audience
- Cloud operations engineers
- Site Reliability Engineers (SREs) and platform engineers
- DevOps engineers and technical team leads
21 Hours
Testimonials (1)
I've find out new interesting things about Lambda and Serverless