How-to Guides
Create Simple AI Inference
To use Simple AI Inference, you must first create an Inference. To create an inference, follow these steps.
Click the All Services > AI-ML > Simple AI Inference menu. 1. Go to the Service Home page of Simple AI Inference.
Service Home on the page, click the Create Simple AI Inference button. 2. Navigate to the Create Serverless Inference page.
Serverless Inference creation page, enter the information required to create the service and select detailed options.
- In the Service Information Input area, select the options required to create the service.
Category RequiredDetailed description Inference service name Required Enter Serverless Inference service name - Enter using lowercase English letters and numbers, 3 ~ 25 characters
Endpoint Required Select external access for Simple AI Inference - Private: Use only private endpoint access control
- Private&Public: Use both private and public endpoint access control
Private endpoint access control Selection Add resources within Samsung Cloud Platform and allow access only to those resources - Private Access Allowed Resource: Select the resource to grant access to
- Click the Add button to select the resource to grant access to
- Select the resource to delete from the resource list, then click the Delete button to remove it
- If no resources are added, access is granted to all resources on subnets within the same region
- Can be modified after applying for a Serverless endpoint
Public endpoint access control Selection Set whether to use public endpoint access control - Enabled if set, you can add IPs or resources that are allowed access
- Public Access Allowed IP: After entering the IP range to allow access in CIDR format or as an IP address, you can add it by clicking the Add button
- Up to 100 entries can be added
- If not used, access is allowed for all IPs
- Can be modified after applying for a Serverless endpoint
Table. Serverless Inference Service Information Input ItemsCautionIf you do not use public endpoint access control or set it to the entire IP range (Any, 0.0.0.0/0), the registry can be exposed to security attacks such as external scanning and hacking. - In the Additional Information Input area, enter or select the required information.
Category Required statusDetailed description tag Selection Add Tag - Up to 50 per resource can be added
- After clicking the Add Tag button, enter or select Key, Value values
Table. Serverless Inference additional information input fields
- In the Service Information Input area, select the options required to create the service.
Summary Verify the detailed information and estimated charges generated in the panel, then click the Create button.
When the popup notifying creation opens, click the Confirm button. 5. The creation request has been completed.
- When creation is complete, check the created items on the Serverless Inference List page.
Check usage by LLM model
On the Service Home page of Simple AI Inference, you can view the list of LLMs and token usage per model.
- All Services > AI-ML > Simple AI Inference Click the menu. 1. Go to the Service Home page of Simple AI Inference.
- Check the per-model usage of LLMs in the LLM Model Usage list on the dashboard of Service Home.
Category Detailed description Model name LLM name - clicking the name moves to the Report tab on the model’s detail page
Model type LLM type - information for each model, see Provided model
Token usage (1 Week) Token usage for the past week as of today Table. Simple AI Inference LLM model usage items
View Serverless Inference details
Follow these steps to view detailed information about Serverless Inference.
All Services > AI-ML > Simple AI Inference Click the menu. 1. Go to the Service Home page of Simple AI Inference.
On the Service Home page, click the Serverless Inference menu. 2. Serverless Inference List Go to the page.
Item Explanation Create Service Serverless Inference can be created - When the button is clicked, navigate to the Serverless Inference creation page
- For creation method, see Create Simple AI Inference
Inference service name Serverless Inference name Model ID Model ID value - When the Model ID is clicked, navigate to the detailed page of that model
- For detailed information about the model, see View model detailed information
Model name Model Name - When clicking the Model ID, navigate to the model’s detail page
- For detailed information about the model, see View model detailed information
Planned model termination date Model’s scheduled end-of-service date Latency Average response time Throughput The average number of tokens the model generates per second Uptime System uptime ratio that allows the system to operate normally without service interruption and handle user requests - Green: 95% or higher
- Yellow: 80% or higher ~ less than 95%
- Red: less than 80%
Service cancellation Serverless Inference can be terminated - When the button is clicked, navigate to the Serverless Inference termination page
- For termination instructions, see Terminate Inference
Table. Serverless Inference list informationReferenceClicking Model ID or Model name takes you to the Model Catalog’s model detail page, where you can view the model’s detailed information.Serverless Inference List page, click the Inference service name to view detailed information. 3. Serverless Inference Details Go to the page.
- Serverless Inference Detailed page consists of Details, Report, Tags, Job History tabs.
Detailed Information
Serverless Inference List page lets you view detailed information of the selected resource and modify the information if necessary.
| Category | Detailed description |
|---|---|
| service | Service Name |
| Resource Type | Resource Type |
| SRN | Unique resource ID in Samsung Cloud Platform |
| Resource name | Resource Name |
| Resource ID | Unique resource ID in the service |
| Constructor | User who created the service |
| Creation Date/Time | Service creation date and time |
| Modifier | User who edited the service information |
| Modification date and time | Date and time the service information was modified |
| Endpoint | External access methods for Simple AI Inference
|
| Private endpoint | Private endpoint value
|
| Public endpoint | Public endpoint value
|
| Private endpoint access control | Information about resources with private access allowed
|
| Public endpoint access control | Publicly accessible IP and resource information
|
Report
On the Serverless Inference List page, you can view the daily LLM call count and token usage for the selected resource.
| Category | Detailed description |
|---|---|
| Search filter | Select items to view in the report
|
| Number of calls | Display the number of calls as a graph for the selected period |
| Total call count | Provide the number of calls per model during the query period. |
| Token usage | Display Input and Output token usage as a graph over the selected period |
| Total token count | Display the total token usage during the query period, separated into Input and Output. |
| Average number of tokens per request | Display the average number of tokens used for LLM calls during the query period, separated into Input and Output. |
Tag
On the Serverless Inference List page, you can view the tag information of the selected resource, and you can add, modify, or delete it.
| Category | Detailed description |
|---|---|
| Tag List | Tag list
|
Job History
You can view the operation history of the selected resource on the Serverless Inference List page.
| Category | Detailed description |
|---|---|
| Task History List | Resource change history
|
View model detailed information
You can view the models provided by Simple AI Inference and their detailed information. To view the model details, follow these steps.
- All Services > AI-ML > Simple AI Inference Click the menu. 1. Go to the Service Home page of Simple AI Inference.
- On the Service Home page, click the Model Catalog menu. 2. Navigate to the Model Catalog page.
- On the Model Catalog page, click the model whose detailed information you want to view. 3. Model Catalog Navigate to the detailed page.
Item Explanation License Click the button to view the model’s license information. Overview Basic description of the model Sales criteria Model developer Category Scope of model usage latest version Provided version Release date Model release year and date Model ID Model ID information Maximum token Maximum token size Output modelities Model output method Input modelities Model input method language Model language types Deployment type Model deployment method Token Limits Token limit value Reqeust Limits request limit Table. Simple AI Inference Provided Model Details
Managing API Keys
You must create and register an API key to use Simple AI Inference in Severless Inference.
Create API Key
To generate an API key, follow these steps.
Click the All Services > AI-ML > Simple AI Inference menu. 1. Go to the Service Home page of Simple AI Inference.
On the Service Home page, click the API Key menu. 2. Navigate to the API key page.
On the API key page, click the Create key button. 3. Create API Key Go to the detail page.
On the API Key Creation page, after entering the information required to generate an API key, click the Create button.
Category Required statusDetailed description Inference type Required Select inference type Expiration period Required Enter the expiration period of the API key - permanent checking the item allows use without any time restriction
Usage Selection Enter the purpose of using the API key within 128 characters Table. Serverless Inference Service Information Input ItemsCautionIf you do not use public endpoint access control or set it to the entire IP range (Any, 0.0.0.0/0), the registry can be exposed to security attacks such as external scanning and hacking.When the popup informing you to create an API key opens, click the Confirm button.
- When an API key is created, it is downloaded once at the time of creation.
Check API Key
To check the API key, follow these steps.
- All Services > AI-ML > Simple AI Inference Click the menu. 1. Go to the Service Home page of Simple AI Inference.
- On the Service Home page, click the API Key menu. 2. Go to the API key page.
Item Explanation Authentication key Authentication key information Inference type Inference type with a registered authentication key Creation timestamp Authentication key generation time Expiration date and time Authentication key expiration time Delete Delete the selected authentication key - It becomes active when you select the authentication key to delete from the key list
More Change the usage status of the selected authentication key - Disable When selected, the authentication key is not deleted, only its functionality is blocked
Key generation Create API key - When the button is clicked, go to the Create API Key page
- Refer to Create API Key for how to create an API key
Table. Simple AI Inference provided model detailed information
Terminate Inference
Terminate Serverless Inference
To cancel Serverless Inference, follow the steps below.
- Click the All Services > AI-ML > Simple AI Inference menu. 1. Go to the Service Home page of Simple AI Inference.
- On the Service Home page, click the Serverless Inference menu. 2. Serverless Inference List Go to the page.
- On the Serverless Inference List page, click the Cancel Service button of the Serverless Inference you want to delete.
- When the pop-up notifying service termination opens, enter the service name and click the Confirm button.