ServiceWatch metric
Simple AI Training sends metrics to ServiceWatch. The metrics provided by default monitoring are data collected at 5‑minute intervals.
Basic Metrics
The following are the basic metrics for the Simple AI Training namespace. The indicators whose names are displayed in bold below are the key indicators selected from the basic indicators provided by Simple AI Training. The main metrics are used to build service dashboards that are automatically created for each service in ServiceWatch. Each metric guides users via the user guide on which statistical value is meaningful when querying the metric, and among the meaningful statistics, the values displayed in bold are the primary statistics.
In the service dashboard or monitoring tab, you can view key metrics through primary statistical values. Or you can also view the key metrics on the monitoring tab of the Simple AI Training detail page. In ServiceWatch’s metrics menu, you can also view utilization by GPU device.
| Performance Item (Metric Name) | Detailed description | unit | meaningful statistics |
|---|---|---|---|
| CPU Usage | Average number of CPU cores used by the Training Job Pod in the last 5 minutes | Cores |
|
| GPU Utilization | GPU utilization in the Training Job | Percent |
|
| Memory Usage | Memory currently used in the Training Job Pod | Bytes |
|