The page has been translated by Gen AI.

ServiceWatch metric

Simple AI Inference sends metrics to ServiceWatch. The metrics provided by default monitoring are data collected at 5‑minute intervals.

Reference
For checking metrics in ServiceWatch, see the ServiceWatch guide.

Basic Metrics

The following are the basic metrics for the Simple AI Inference namespace. The indicators whose names are displayed in bold below are the key indicators selected among the default indicators provided by Simple AI Inference. The key metrics are used to build service dashboards that are automatically created for each service in ServiceWatch. Each metric provides guidance in the user guide on which statistical values are meaningful when querying that metric, and among the meaningful statistics, the values shown in bold are the primary statistics.

In the service dashboard or monitoring tab, you can view key metrics through primary statistical values. Or you can also view the key metrics on the monitoring tab of the Simple AI Inference detail page. You can also view the usage rate per GPU device in the ServiceWatch metrics menu.

Performance item (metric name)Detailed descriptionunitmeaningful statistics
Model Total TokensModel token usage (total)Count
  • Total
Model Request Server ErrorNumber of model request failures (server error)Count
  • total
Model Input TokensModel token usage (input)Count
  • Total
Model Request ThrottledModel request limit count (request quota exceeded)Count
  • Total
Model Request Client ErrorModel request failure count (client error)Count
  • Total
Model Output TokensModel token usage (output)Count
  • total
Model Cached TokensModel token usage (cache)Count
  • Total
Model Request Prompt RejectedNumber of model request rejections (prompt review)Count
  • Total
Model Request SuccessNumber of successful model requestsCount
  • Total
Table. Simple AI Inference Basic Metrics
Overview
How-to Guides