Get in Touch

Course Outline

Foundations of Agentic Systems in Production

  • Agentic architectures: covering loops, tools, memory, and orchestration layers
  • The agent lifecycle: spanning development, deployment, and continuous operation
  • Key challenges in managing production-scale agents

Infrastructure and Deployment Models

  • Deploying agents across containerized and cloud-based environments
  • Scaling strategies: comparing horizontal vs vertical scaling, concurrency, and throttling
  • Orchestrating multi-agent systems and balancing workloads

Monitoring and Observability

  • Critical metrics: tracking latency, success rates, memory usage, and agent call depth
  • Tracing agent activity and mapping call graphs
  • Implementing observability using Prometheus, OpenTelemetry, and Grafana

Logging, Auditing, and Compliance

  • Implementing centralized logging and structured event collection
  • Ensuring compliance and auditability within agentic workflows
  • Creating audit trails and replay mechanisms to facilitate debugging

Performance Tuning and Resource Optimization

  • Minimizing inference overhead and refining agent orchestration cycles
  • Utilizing model caching and lightweight embeddings to accelerate retrieval
  • Conducting load testing and stress simulations for AI pipelines

Cost Control and Governance

  • Analyzing agent cost drivers: API calls, memory, compute resources, and external integrations
  • Monitoring agent-level expenditures and implementing chargeback models
  • Establishing automation policies to prevent agent sprawl and reduce idle resource consumption

CI/CD and Rollout Strategies for Agents

  • Embedding agent pipelines into CI/CD workflows
  • Executing testing, versioning, and rollback strategies for iterative agent updates
  • Implementing progressive rollouts and secure deployment mechanisms

Failure Recovery and Reliability Engineering

  • Designing systems for fault tolerance and graceful degradation
  • Applying retry, timeout, and circuit breaker patterns to enhance agent reliability
  • Establishing incident response and post-mortem frameworks for AI operations

Capstone Project

  • Developing and deploying an agentic AI system equipped with comprehensive monitoring and cost tracking
  • Simulating load conditions to measure performance and optimize resource usage
  • Presenting the final architecture and monitoring dashboards to peers

Summary and Next Steps

Requirements

  • A solid grasp of MLOps and production machine learning systems
  • Hands-on experience with containerized deployments using Docker/Kubernetes
  • Proficiency with cloud cost optimization and observability tooling

Target Audience

  • MLOps engineers
  • Site Reliability Engineers (SREs)
  • Engineering managers responsible for AI infrastructure
 21 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories