Service Overview
Simple AI Inference is a serverless service that provides various global foundation models as APIs, offering public or private environments so that LLMs can be used on internal Samsung Cloud Platform resources or externally. By using Simple AI Inference, you can use multiple LLM models through the same API and improve productivity in AI application service development. It also supports compatibility with OpenAI and the LangChain SDK, enabling easy integration with existing development environments and frameworks.
Features
- Convenient LLM Model Usage: As a fully managed serverless service, you can use multiple LLM models through the same API.
- Efficient Cost Management: Costs are charged based on the actual usage of input (Input) and output (Output) tokens.
- Stable Service Provision: We provide stable services through traffic control (TPM/RTM).
- Enterprise security provided: Data is securely protected in a rigorous security environment and is not used for external model training.
Service architecture diagram
Provided Features
Simple AI Inference provides the following features.
Check convenient LLM model
- You can easily view the features and primary use cases of LLM models provided through the LLM model catalog.
- You can view and test the provided LLM model directly on the console screen using PlayGround.ReferencePlayGround is scheduled to be offered after September 2026.
Shared Use of LLM Model Account: If you request a model to use in Simple AI Inference, all users within the same Account can use it.
Serverless Service Provision : Users can request the desired model via an API and use it immediately without managing resources, and they pay only for what they use.
Public/Private endpoint provision: Depending on the user’s inference usage pattern, you can choose to use either a Public or Private endpoint.
Stable Service Provision: We provide a stable service environment through traffic control (TPM/RPM).
Provided model
The LLM models provided by Simple AI Inference are as follows.
| Model name | Application | Input type | TPM | RPM | Context Size | Image input limit count |
|---|---|---|---|---|---|---|
| Qwen3.6-27B | Text, Agent | Text, Image | 1,000,000 | 100 | 262,144 | 8 |
| gemma-4-31B-it | Text, Agent | Text, Image | 1,000,000 | 100 | 262,144 | 8 |
| gpt-oss-120b | Text | Text | 1,000,000 | 100 | 131,072 | - |
| Llama-Guard-4-12B | Security | Text, Image | 1,000,000 | 250 | 307,200 | 8 |
| Qwen3-VL-Embedding-8B | embedding | Text, Image | 1,000,000 | 250 | 262,144 | 8 |
| Qwen3-VL-Reranker-8B | reranker | Text, Image | 1,000,000 | 250 | 262,144 | 8 |
Provision status by region
The regions that provide Simple AI Inference service are as follows.
| Region | Provision status |
|---|---|
| Korea West (kr-west1) | Provide |
| Korea East (kr-east1) | Not provided |
| South Korea 1 (kr-south1) | Not provided |
| South Korea South 2 (kr-south2) | Not provided |
| South Korea South 3 (kr-south3) | Not provided |
Preceding Service
There are no services that need to be pre-configured before creating this service.
