Get in Touch
 Duration 14 hours

Course Outline

Introduction to AIOps

  • Defining AIOps and understanding its strategic importance.
  • Comparing traditional monitoring with AIOps-driven observability.
  • Examining AIOps architecture and its key components.

Collecting and Normalizing Operational Data

  • Identifying types of observability data: metrics, logs, and traces.
  • Ingesting data from diverse sources, including servers, containers, and cloud environments.
  • Utilizing agents and exporters such as Prometheus, Beats, and Fluentd.

Data Correlation and Anomaly Detection

  • Applying time series correlation and statistical methods.
  • Employing ML models for effective anomaly detection.
  • Identifying incidents across distributed systems.

Alerting and Noise Reduction

  • Designing intelligent alert rules and appropriate thresholds.
  • Implementing suppression, deduplication, and alert grouping techniques.
  • Integrating with tools like Alertmanager, Slack, PagerDuty, or Opsgenie.

Root Cause Analysis and Visualization

  • Visualizing metrics and detecting trends using dashboards.
  • Exploring events and timelines for thorough Root Cause Analysis (RCA).
  • Tracing issues across layers with distributed tracing tools.

Automation and Remediation

  • Triggering automated scripts or workflows based on incident data.
  • Integrating with ITSM systems such as ServiceNow or Jira.
  • Exploring use cases like self-healing, automatic scaling, and traffic rerouting.

Open Source and Commercial AIOps Platforms

  • Overview of key tools: Prometheus, Grafana, ELK, Moogsoft, and Dynatrace.
  • Establishing evaluation criteria for selecting the right AIOps platform.
  • Demonstrating and practicing with a selected technology stack.

Summary and Next Steps

Requirements

  • A solid grasp of IT operations and system monitoring concepts.
  • Practical experience with monitoring tools or dashboards.
  • Familiarity with standard log and metric formats.

Target Audience

  • Operations teams overseeing infrastructure and applications.
  • Site Reliability Engineers (SREs).
  • IT monitoring and observability teams.

Number of participants


Price per participant

Upcoming Courses

Related Categories