Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI for Operations
- Evolution of IT automation: moving from static runbooks to reasoning agents
- Anatomy of an agent: understanding the reasoning loop, tool usage, memory, and planning
- Decision-making framework: determining when to automate and when to involve humans
Agent Frameworks and Architectures
- Single-agent patterns: exploring ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent architectures: supervisor, hierarchical, and swarm models
- Framework comparison: LangGraph, CrewAI, AutoGen, and custom agent implementations
- Building your first operational agent: querying monitoring, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
- Agent-driven log querying: integrating Elasticsearch, Loki, and Splunk
- Infrastructure tool usage: executing kubectl, Terraform, and Ansible commands via agent actions
- Designing secure tool interfaces with parameter validation and idempotency
Incident Response Automation
- Automated incident triage: classifying severity and routing alerts
- Generating root cause hypotheses and gathering supporting evidence
- Automated remediation: executing restart, scale, rollback, and failover actions
- Creating an incident runbook agent with progressive levels of autonomy
Safety, Guardrails, and Human-in-the-Loop
- Action classification: distinguishing between read-only, low-risk, high-risk, and destructive operations
- Establishing approval gates and escalation policies for critical actions
- Guardrail patterns: implementing action allowlists, blast radius limits, and rollback guarantees
- Maintaining audit trails and decision provenance for compliance purposes
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialist agents: triage, diagnosis, and remediation roles
- Managing inter-agent communication and shared context
- Resolving conflicts when agents propose contradictory actions
- Simulating end-to-end major incidents with multi-agent responses
Observability and Evaluation
- Tracing agent reasoning chains for debugging and auditing
- Evaluating agent decision quality: measuring precision, recall, and time-to-resolution
- Implementing feedback loops to learn from operator overrides and outcomes
- Tracking costs and managing token economics for operational agents
Production Deployment and Operations
- Deploying agents as services: utilizing APIs, webhooks, and scheduled jobs
- Rolling out gradual autonomy: transitioning from shadow mode to full auto-remediation
- Handling agent failures: operational runbooks for when the AI system breaks
- Building the business case and measuring ROI for autonomous operations
Requirements
- Practical experience with IT operations, DevOps, or SRE practices.
- Proficiency in Python scripting and REST API interactions.
- A foundational understanding of Large Language Model (LLM) capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation strategies.
- Platform engineers focused on building self-healing infrastructure.
- IT operations leaders evaluating agentic AI for enhanced incident management.
14 Hours