EdX

Logging, Monitoring and Observability in Google Cloud (edX)

Offered by Google Cloud,
Logging, Monitoring and Observability in Google Cloud (edX)

This class is intended for the following participants: - Cloud architects -Administrators -SysOps personnel Cloud developers -DevOps personnel. Learn how to monitor, troubleshoot, and improve your infrastructure and application performance. Guided by the principles of Site Reliability Engineering (SRE), this course features a combination of lectures, demos, hands-on labs, and real-world case studies. In this course, you'll gain experience with full-stack monitoring, real-time log management and analysis, debugging code in production, and profiling CPU and memory usage.

Class Deals by MOOC List - Click here and see EdX's Active Discounts, Deals, and Promo Codes.

This course is part of the following programs:

What you'll learn

  • Plan and implement a well-architected logging and monitoring infrastructure.
  • Define service level indicators (SLIs) and service level objectives (SLOs).
  • Create effective monitoring dashboards and alerts.
  • Monitor, troubleshoot, and improve Google Cloud infrastructure.

Syllabus

  1. Introduction

Welcome to Logging, Monitoring and Observability in Google Cloud! Use the resources below to become familiar with the topics this course will cover, learn how to access course materials and how to send feedback.

  1. Introduction to Monitoring in Google Cloud

In this module, we will take some time to do a high-level overview of the various products which comprise Google Cloud’s logging, monitoring, and observability suite.

  1. Avoiding Customer Pain

In this module, we discuss several Site Reliability Engineering (SRE) concepts and how we can use them to help avoid customer pain. In this context, a customer is any consumer of a cloud-based system.

  1. Alerting Policies

Alerting gives timely awareness to problems in your cloud applications so you can resolve the problems quickly. In this module, you will learn how to develop alerting strategies, define alerting policies, add notification channels, identify types of alerts and common uses for each, construct and alert on resource groups, and manage alerting policies programmatically.

  1. Monitoring Critical Systems

Monitoring is all about keeping track of exactly what's happening with the resources we've spun up inside of Google's Cloud. In this module, we'll take a look at options and best practices as they relate to monitoring project architectures. We'll differentiate the core Cloud IAM roles needed to decide who can do what as it relates to monitoring. Just like architecture, this is another crucial early step. We will examine some of the Google created default dashboards, and see how to use them appropriately. We will create charts and use them to build custom dashboards to show resource consumption and application load. And, finally, we will define uptime checks to track liveliness and latency.

  1. Configuring Google Cloud Services for Observability

In the next part of our Metrics discussion, let’s take a little time to examine the art of Configuring Google Cloud Services for Observability. In this module, we're going to spend a little time learning how to integrate logging and monitoring agents into Compute Engine VMs and images using Agents, enable and utilize Kubernetes Monitoring, extend and clarify Kubernetes monitoring with Prometheus, and expose custom metrics through code, and with the help of OpenCensus.

  1. Advanced Logging and Analysis

In this module, we will examine some of Google Cloud's advanced logging and analysis capabilities. Specifically, in this module you will learn to identify and choose among resource tagging approaches, define log sinks, create monitoring metrics based on log entries, link application errors to Logging and other operation tools using Error Reporting, and export logs to BigQuery for long term storage and SQL based analysis.

  1. Monitoring Network Security and Audit Logs

In this module, we will examine two key topics: Monitoring as it relates to the VPC network, and how to use Google's Cloud Audit logs. You will learn to collect and analyze VPC Flow, Firewall Rule, and Cloud NAT logs, enable Packet Mirroring, explain the capabilities of the Network Intelligence Center, and use Cloud Audit logs to answer the question, “Who, did what, and when?” We will also cover best practices for Audit Logging.

  1. Managing Incidents

Up to this point in our course, we've mostly focused on ways to inspect and monitor the status of our systems running in Google Cloud. But no matter how solid your planning, design, architecture, and preventive maintenance strategies are, things will go wrong. When they do go wrong, how you manage those incidents will have a huge impact on user perception. In this module, you will learn how to handle incidents using a systematic process.

  1. Investigating Application Performance Issues

When deploying applications to Google Cloud, the Application Performance Management products (Cloud Trace, Cloud Debugger, and Cloud Profiler) provide a suite of tools to give insight into how your code and services are functioning, and to help troubleshoot where needed.

  1. Optimizing the Costs of Monitoring

In our final module we discuss optimizing the costs for Google Cloud’s operations suite. Specifically, you will learn to analyze resource utilization costs for operations related components within Google Cloud, and implement best practices for controlling the cost of operations within Google Cloud.

  1. Course Resources

PDF links to all modules

Go to Class
MOOC List is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Related Courses

ML Pipelines on Google Cloud (Coursera) Coursera
Google Cloud

ML Pipelines on Google Cloud (Coursera)

In this course, you will be learning from ML Engineers and Trainers who work with the state-of-the-art development of ML pipelines here at Google Cloud. The first few modules will cover about TensorFlow Extended (or TFX), which is Google’s production machine learning platform based on TensorFlow for management of ML pipelines and metadata. You will learn about pipeline components and pipeline orchestration with TFX. You will also learn how you can automate your pipeline through continuous integration and continuous deployment, and how to manage ML metadata.

Aug 17th 2026
4 Weeks
Dataplex by Google Cloud (Coursera) Coursera
Board Infinity

Dataplex by Google Cloud (Coursera)

Welcome to "Dataplex By Google Cloud " a comprehensive course designed to provide a thorough understanding of Google Cloud Dataplex, a platform for managing, monitoring, and analyzing data across various data systems in Google Cloud. Spanning two modules, the course begins with the fundamentals of Dataplex, including its setup, configuration, and basic functionalities.

Aug 17th 2026
2 Weeks
Google Cloud Customer Care Fundamentals (Coursera) Coursera
Google Cloud

Google Cloud Customer Care Fundamentals (Coursera)

This course will teach you how to get the most out of Google Cloud Support. You will learn about the different support services provided by Google Cloud Customer care, how to create and manage support cases, how to view known issues affecting Google Cloud services, and how to communicate effectively with Support Engineers. You will also learn about the different case priorities and Service Level Objectives (SLOs), increase your understanding around case status, and how to escalate a support case if necessary.

Aug 17th 2026
4 Weeks
Optimizing Your Google Cloud Costs en Español (Coursera) Coursera
Google Cloud

Optimizing Your Google Cloud Costs en Español (Coursera)

Optimizing Your Google Cloud Platform (GCP) Costs es el segundo curso de una serie de dos partes sobre los conceptos básicos de la administración de costos y facturación de GCP. Este curso está destinado principalmente a personas que se desempeñan en funciones relacionadas con finanzas o TI, y que están a cargo de optimizar la infraestructura de nube de sus organizaciones.

Aug 17th 2026
3 Weeks
SRE Capstone (edX) EdX
IBM

SRE Capstone (edX)

The SRE Capstone offers interactive study guides and flash cards that will help you prepare for the Professional SRE - Cloud V2 certification exam. Also included are hands-on lab exercises that allow you to put the knowledge you gained from the SRE Fundamentals and Security and SRE Infrastructure, Resiliency and Deployment Automation courses into action.

Self Paced
Self-Paced