AI-Driven Observability: From Logs to LLM-Powered Insights Training Course
Traditional observability relies heavily on dashboards, threshold-based alerts, and manual log analysis. AI-driven observability revolutionizes this approach by enabling natural language querying of telemetry data, leveraging LLMs for root cause analysis, utilizing foundation models for anomaly detection, and generating automated incident summaries that grasp contextual nuances.
This instructor-led, live training (available online or onsite) is designed for observability and SRE engineers looking to integrate Large Language Models (LLMs) and artificial intelligence into their monitoring, alerting, and incident analysis workflows.
By the end of this training, participants will be able to:
- Create natural language interfaces for querying Prometheus, Elasticsearch, and SQL-based observability repositories.
- Implement pipelines for LLM-powered log analysis and anomaly detection.
- Generate automated incident summaries and draft postmortems from raw telemetry data.
- Design AI-assisted root cause analysis workflows utilizing evidence chaining.
- Integrate foundation models for time-series anomaly detection and forecasting.
- Deploy an AI-enhanced on-call experience featuring smart alert enrichment.
Format of the Course
- Interactive lectures and discussions.
- Extensive exercises and practical practice.
- Hands-on implementation within a live-lab environment.
Course Customization Options
- To request customized training, please contact us to arrange.
Course Outline
The AI Observability Landscape
- From dashboards to conversations: the shift toward AI-augmented observability
- LLM capabilities relevant to observability: summarization, reasoning, pattern matching
- Architecture patterns: embedding AI into existing observability stacks
Natural Language Telemetry Querying
- Text-to-PromQL: translating natural language into monitoring queries
- NL querying for Elasticsearch, OpenSearch, and Loki log stores
- SQL generation from natural language for structured telemetry
- Building a query assistant agent with tool use and context awareness
LLM-Powered Log Analysis
- Automated log parsing and structuring with LLMs
- Anomaly detection in log streams using embedding similarity
- Log clustering and pattern discovery at scale
- Generating human-readable explanations from raw log sequences
Intelligent Alerting and Incident Enrichment
- Alert correlation and deduplication with semantic understanding
- Automated incident context gathering from runbooks, past incidents, and docs
- Smart alert routing based on content understanding and team expertise
- Reducing alert fatigue with AI-driven noise reduction
AI-Assisted Root Cause Analysis
- Hypothesis generation from multi-source telemetry correlation
- Evidence chaining: connecting symptoms across metrics, logs, and traces
- Guided troubleshooting with interactive AI diagnosis sessions
- Building a root cause analysis agent with progressive investigation
Automated Incident Response and Communication
- Generating incident summaries and status updates from telemetry
- Automated postmortem drafting with timeline reconstruction
- Stakeholder communication tailored to technical and executive audiences
- Runbook suggestion and automated remediation recommendations
ML for Observability
- Time-series forecasting for capacity planning and anomaly prediction
- Foundation models for zero-shot anomaly detection on metrics
- Embedding-based service dependency mapping and topology discovery
- Training and deploying lightweight ML models alongside observability pipelines
Production Deployment and Ethics
- Latency and cost considerations for real-time AI observability
- Data privacy: ensuring LLMs do not leak sensitive telemetry
- Human oversight: when AI diagnosis needs operator validation
- Measuring impact: MTTD, MTTR, and on-call experience metrics
Requirements
- Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry.
- Familiarity with log management and metrics concepts.
- Basic Python scripting skills for data processing.
Audience
- SRE and observability engineers adopting AI-enhanced tooling.
- Platform engineers building next-generation monitoring pipelines.
- DevOps leads evaluating LLM integration into incident workflows.
Open Training Courses require 5+ participants.
AI-Driven Observability: From Logs to LLM-Powered Insights Training Course - Booking
AI-Driven Observability: From Logs to LLM-Powered Insights Training Course - Enquiry
AI-Driven Observability: From Logs to LLM-Powered Insights - Consultancy Enquiry
Upcoming Courses
Related Courses
Agentic Development with Gemini 3 and Google Antigravity
21 HoursGoogle Antigravity serves as an agentic development environment built to create autonomous agents that leverage Gemini 3's multimodal capabilities for planning, reasoning, coding, and execution.
Delivered as an instructor-led, live training session—available either online or onsite—this program targets advanced technical professionals seeking to design, build, and deploy autonomous agents utilizing Gemini 3 within the Antigravity environment.
By the conclusion of this training, participants will be equipped to:
- Construct autonomous workflows that harness Gemini 3 for reasoning, planning, and execution.
- Create agents in Antigravity capable of analyzing tasks, generating code, and interacting with various tools.
- Seamlessly integrate Gemini-driven agents with enterprise systems and APIs.
- Enhance agent behavior, safety, and reliability across complex environments.
Course Format
- Expert-led demonstrations paired with interactive discussions.
- Practical, hands-on experimentation focused on autonomous agent development.
- Real-world implementation using Antigravity, Gemini 3, and complementary cloud tools.
Course Customization Options
- If your team requires specific domain-specific agent behaviors or custom integrations, please reach out to us to tailor the program to your needs.
Advanced Antigravity: Feedback Loops, Learning & Long-Term Agent Memory
14 HoursGoogle Antigravity serves as a sophisticated framework dedicated to experimenting with long-lived agents and their emergent interactive behaviors.
This live, instructor-led training—available both online and onsite—targets advanced-level professionals who aim to design, analyze, and optimize agents capable of retaining memories, enhancing their capabilities through feedback, and evolving over extended operational periods.
Upon completing this course, participants will acquire the skills to:
- Design long-term memory structures for agent persistence.
- Implement effective feedback loops to shape agent behavior.
- Evaluate learning trajectories and model drift.
- Integrate memory mechanisms into complex multi-agent ecosystems.
Format of the Course
- Expert-led discussions paired with technical demonstrations.
- Hands-on exploration through structured design challenges.
- Application of concepts to simulated agent environments.
Course Customization Options
- If your organization requires tailored content or case-specific examples, please contact us to customize this training.
Advanced Mastra Integrations: APIs, Tools, Enterprise Data & External Systems
21 HoursMastra is a framework that supports deep integration between AI agents, APIs, enterprise applications, and external data systems.
This instructor-led, live training (online or onsite) is aimed at intermediate-level engineers who wish to build reliable, secure, and scalable integrations between Mastra agents and the broader enterprise ecosystem.
Upon completion of this training, participants will be prepared to:
- Implement API-driven integrations between Mastra agents and external services.
- Connect enterprise data systems and tools to automated agent workflows.
- Apply secure data exchange and authentication best practices.
- Design integration layers that are scalable, maintainable, and production ready.
Course Format
- Interactive lectures and discussions.
- Hands-on integration engineering and API exercises.
- Live-lab implementation using real-world enterprise scenarios.
Course Customisation Options
- Custom API scenarios, enterprise system mappings, or data-integration workshops are available upon request.
Interactive AI Agents: AgentCore Memory, Code Interpreter & Browser Tool in Action
14 HoursAgentCore facilitates the creation of interactive, dynamic, and context-aware AI agent experiences by providing memory persistence, a secure code interpreter, and an integrated browser tool.
This instructor-led live training, available online or onsite, is designed for intermediate to advanced technical professionals seeking to develop and deploy AI agents with capabilities in long-term context retention, real-time computation, and direct web UI interaction.
Upon completion of this training, participants will be equipped to:
- Implement AgentCore memory to establish stateful, context-aware workflows.
- Utilize the secure code interpreter to perform dynamic calculations and data transformations.
- Integrate the browser tool for immediate data acquisition and user interface engagement.
- Design interactive agents tailored for analytics, customer support, and research applications.
Course Delivery Format
- Interactive lectures combined with group discussions.
- Practical lab exercises focused on AgentCore memory and tool utilization.
- Analysis of case studies covering analytics, automation, and customer support scenarios.
Customization Options
- Contact us to request and arrange a customized training curriculum for this course.
Accelerating AI Agent Deployment with AgentCore Runtime & Gateway
14 HoursAgentCore Runtime and Gateway constitute a pair of AWS services designed to package, deploy, and securely expose AI agents, featuring streamlined integrations with external systems.
This instructor-led live training (available online or onsite) targets intermediate-level engineering teams aiming to transition from agent prototypes to production. Participants will master the AgentCore Runtime for deployment and the Gateway for secure connectivity and API integration.
By the conclusion of this training, participants will be capable of:
- Establishing AgentCore Runtime environments and packaging agents for deployment.
- Exposing agents via the Gateway using authenticated, rate-limited endpoints.
- Integrating external tools and APIs into agent workflows through stable contracts.
- Implementing observability, logging, and usage monitoring for production operations.
Course Format
- Interactive lectures and discussions.
- Hands-on labs focused on Runtime deployments and Gateway integrations.
- Practical exercises emphasizing reliability, security, and rollout strategies.
Course Customization Options
- To request customized training for this course, please contact us to make arrangements.
Antigravity for Developers: Building Agent-First Applications
21 HoursAntigravity is a development platform purpose-built for constructing AI-powered, agent-centric applications.
This instructor-led live training, available online or on-site, targets intermediate developers looking to build real-world solutions utilizing autonomous AI agents within the Antigravity ecosystem.
Upon completion, participants will be prepared to:
- Create applications powered by autonomous and coordinated AI agents.
- Leverage the Antigravity IDE, editor, terminal, and browser for comprehensive development.
- Oversee multi-agent workflows via the Agent Manager.
- Embed agent capabilities into production-ready software architectures.
Course Delivery Format
- Interactive presentations featuring detailed demonstrations.
- Substantial hands-on practice with guided exercises.
- Practical implementation tasks within the live Antigravity environment.
Customization Opportunities
- To tailor the curriculum to your specific development stack, please reach out to us for a customized training arrangement.
Getting Started with Antigravity: An Introduction to Agent-First IDEs
14 HoursGoogle Antigravity serves as an agent-centric development environment engineered to optimize engineering processes via intelligent automation.
This live, instructor-led training session, available either online or onsite, is tailored for beginners looking to grasp the core principles of Antigravity and explore how agent-powered coding environments boost productivity.
By the end of this program, attendees will be equipped to:
- Install and set up Google Antigravity.
- Explore and comprehend both the Editor and Manager views.
- Collaborate with agents to streamline basic development assignments.
- Leverage Antigravity for the creation, refinement, and administration of project files.
Training Delivery Method
- Instructor-led instruction accompanied by live demonstrations.
- Supervised practical exercises emphasizing direct interaction with agents.
- Hands-on investigation of primary Antigravity functionalities within a secure lab setup.
Customization Possibilities
- If you need a specific variation of this training, reach out to us to design a bespoke program.
Antigravity for Web Automation & Browser-Based Tasks
21 HoursGoogle Antigravity serves as a robust platform for developing agents that interact seamlessly with web applications, browser environments, and complex multi-surface workflows.
Designed for intermediate-level professionals, this instructor-led live training—available either online or on-site—equips you with the skills to build, automate, and test browser-based workflows using Google Antigravity.
By the end of this program, participants will be equipped to:
- Develop agents that engage with web applications within a browser surface.
- Streamline end-to-end workflows across various browser contexts.
- Effectively validate and troubleshoot agent behavior in UI-driven environments.
- Execute cross-surface automation strategies leveraging Antigravity.
Course Delivery Format
- Guided instruction complemented by live demonstrations.
- Hands-on practical activities and scenario-based exercises.
- Implementation of agent workflows within an interactive lab environment.
Customization Opportunities
- Contact us to tailor the course content to your specific training objectives and requirements.
Building Fully Managed AI Agents with AgentCore: From Concept to Production
14 HoursAgentCore streamlines the creation, optimization, and oversight of fully managed AI agents by offering a cohesive suite of services designed for large-scale implementation.
This live, instructor-led training (available online or onsite) is designed for practitioners with beginner to intermediate experience who seek practical skills in developing production-ready AI agents using AgentCore.
Upon completion, participants will be equipped to:
- Comprehend the fundamental capabilities of AgentCore in AI agent development.
- Architect and set up basic AI agents leveraging managed services.
- Incorporate workflows to boost agent capabilities.
- Launch and oversee AI agents in live production settings.
Course Delivery Format
- Engaging lectures and interactive discussions.
- Practical labs utilizing AgentCore services.
- Structured exercises guiding you from concept to rollout.
Customization Possibilities
- For tailored training options, please reach out to schedule a discussion.
AI Agent Development with Mastra
14 HoursThis instructor-led, live training session (available online or onsite) is designed for intermediate-level software developers and engineering teams aiming to build scalable, observable AI systems with Mastra.
Upon completion, participants will be equipped to:
- Understand Mastra’s architecture and how it connects with LLMs and external APIs.
- Design and create AI agents and workflows using TypeScript.
- Utilize Mastra’s observability and memory tools to monitor and boost agent performance.
- Launch production-ready AI applications by harnessing Mastra’s framework capabilities.
Mastra Debugging, Evaluation & Quality Assurance for AI Agents
21 HoursMastra is a framework that provides structured tools for evaluating, debugging, and assuring the reliability of AI agents operating across complex workflows.
This instructor-led, live training (online or onsite) is aimed at intermediate-level practitioners who wish to rigorously test agent behaviour, improve reliability, and implement measurable evaluation processes.
By the end of this training, participants will confidently:
- Apply debugging techniques to identify and correct agent behaviour issues.
- Evaluate agents using structured metrics, benchmarks, and quality scores.
- Implement tooling and workflows that track reliability, drift, and hallucinations.
- Design QA strategies that ensure consistent and predictable agent performance.
Course Format
- Interactive lecture and discussion.
- Hands-on debugging and evaluation exercises.
- Live-lab analysis of agent behaviours using observability tools.
Course Customisation Options
- Customised reliability testing scenarios and industry-specific QA methods can be arranged upon request.
Mastra Ops & Production Engineering: Deploying and Scaling AI Agents
21 HoursMastra is an operational framework designed to streamline the deployment, scaling, and lifecycle management of AI agents in production environments.
This instructor-led, live training (online or onsite) is aimed at intermediate-level to advanced-level technical professionals who need to operationalize AI agents reliably and efficiently across production systems.
Upon completion of this training, attendees will be equipped to:
- Deploy Mastra-based AI agents into controlled, production-grade environments.
- Scale agents horizontally and vertically using platform-native primitives.
- Implement observability pipelines to track agent behaviour and performance.
- Optimize runtime configurations to reduce latency, costs, and operational risks.
Format of the Course
- Interactive lecture and discussion.
- Hands-on exercises focused on real deployment scenarios.
- Live-lab implementation using containerized and orchestrated environments.
Course Customization Options
- Customization of topics, hands-on labs, or industry-specific scenarios is available upon request.
Mastra Workflow Automation & Multi-Agent Orchestration
21 HoursMastra is a framework designed to facilitate sophisticated workflow automation and coordination among multiple AI agents within distributed systems.
This instructor-led training, available both online and on-site, is tailored for intermediate-level professionals aiming to design, orchestrate, and manage multi-agent workflows at scale.
Upon completion of this training, participants will acquire the following skills:
- Design complex workflows leveraging Mastra’s orchestration capabilities.
- Coordinate multiple agents executing parallel or dependent tasks.
- Implement monitoring and debugging tools for workflow execution.
- Optimize orchestration logic to enhance reliability, throughput, and automation efficiency.
Course Format
- Interactive lectures and discussions.
- Hands-on exercises in workflow design and automation.
- Practical implementation within a containerized live-lab environment.
Course Customization Options
- Customized automation scenarios, enterprise integrations, or workflow patterns can be provided upon request.
Managing Agent Workflows in Google Antigravity: Orchestration, Planning and Artifacts
14 HoursGoogle Antigravity serves as a specialized development platform centered around AI agents, designed to coordinate, oversee, and manage automated coding and workflow processes.
This live, instructor-led training program, available both online and onsite, is tailored for intermediate-level professionals seeking to refine their ability to design, oversee, and enhance multi-agent operations within the Google Antigravity ecosystem.
By the end of this session, participants will be equipped with the following capabilities:
- Setting up agent responsibilities and orchestration pipelines through the Manager interface.
- Creating and analyzing Antigravity artifacts, such as task lists, strategic plans, system logs, and browser session recordings.
- Applying verification methods that ensure agent behavior remains transparent and subject to audit.
- Enhancing multi-agent cooperation to effectively handle complex development and operational demands.
Course Structure
- Structured presentations paired with practical live demonstrations.
- Scenario-driven exercises that address real-world workflow complexities.
- Direct experimentation within an active Antigravity workspace environment.
Customization Possibilities
- Should you need a version of this course adjusted to specific organizational needs, please reach out to discuss potential modifications.
Testing & Verifying Agent-Driven Code: Quality Assurance in Antigravity
14 HoursAntigravity serves as a sophisticated framework designed to manage advanced agent-driven development processes.
This live, instructor-led training session, available either online or on-site, is specifically tailored for intermediate to advanced professionals seeking to rigorously verify, validate, and secure the code outputs generated by AI agents operating in Antigravity environments.
By the end of this program, participants will be equipped to:
- Evaluate the precision and security integrity of code artifacts produced by agents.
- Utilize systematic methods to confirm the correct execution of agent-assigned tasks.
- Efficiently analyze browser session recordings and trace the activity history of agents.
- Implement quality assurance and security protocols to guarantee the stability of agent workflows.
Course Delivery Format
- Technical briefings and interactive discussions guided by the instructor.
- Practical exercises dedicated to the verification of live agent workflows.
- Direct testing and validation activities conducted within a controlled laboratory setting.
Customization Opportunities
- Scenarios, workflow structures, and testing examples can be tailored to specific requirements upon request.