The page has been translated by Gen AI.

ServiceWatch metric

Simple AI Training sends metrics to ServiceWatch. The metrics provided by default monitoring are data collected at 5‑minute intervals.

Reference
To view metrics in ServiceWatch, see the ServiceWatch guide.

Basic Metrics

The following are the basic metrics for the Simple AI Training namespace. The indicators whose names are displayed in bold below are the key indicators selected from the basic indicators provided by Simple AI Training. The main metrics are used to build service dashboards that are automatically created for each service in ServiceWatch. Each metric guides users via the user guide on which statistical value is meaningful when querying the metric, and among the meaningful statistics, the values displayed in bold are the primary statistics.

In the service dashboard or monitoring tab, you can view key metrics through primary statistical values. Or you can also view the key metrics on the monitoring tab of the Simple AI Training detail page. In ServiceWatch’s metrics menu, you can also view utilization by GPU device.

Performance Item (Metric Name)Detailed descriptionunitmeaningful statistics
CPU UsageAverage number of CPU cores used by the Training Job Pod in the last 5 minutesCores
  • Total
  • Average
  • Maximum
  • Minimum
GPU UtilizationGPU utilization in the Training JobPercent
  • Average
  • Maximum
  • Minimum
Memory UsageMemory currently used in the Training Job PodBytes
  • Total
  • Average
  • Maximum
  • Minimum
Table. Simple AI Training basic metrics
Server type
How-to guides