The page has been translated by Gen AI.

How-to guides

Users can create the service by entering the required information for the Multi-node GPU Cluster service and selecting detailed options through the Samsung Cloud Platform Console.

Multi-node GPU Cluster Getting Started

You can create and use a Multi-node GPU Cluster service from the Samsung Cloud Platform Console.

This service consists of a GPU Node and a Cluster Fabric service.

Create GPU Node

Follow the steps below to create a Multi-node GPU Cluster.

  1. All Services > Compute > Multi-node GPU Cluster Click the menu. 1. Go to the Service Home page of the Multi-node GPU Cluster.

  2. On the Service Home page, click the Create GPU Node button. 2. Navigate to the GPU Node Creation page.

  3. On the GPU Node creation page, enter the information required to create the service, and select detailed options.

    • Select the required information in the Image and Version Selection area.
      Category
      required status
      Detailed description
      imageRequiredSelect the type of image provided
      • Ubuntu
      Image versionRequiredSelect version of the selected image
      • Provide version list of the provided server image
      Table. GPU Node image and version selection options
    • Enter or select the required information in the Service Information Input area.
      Category
      required status
      Detailed description
      Number of serversRequiredNumber of GPU Node servers to create simultaneously
      • Only numeric input is allowed, and the minimum number of servers to create is 2.
      • When initially configuring, you must create at least 2, and expansions can be done one at a time.
      Service Type > Server TypeRequiredGPU Node server type
      • Select the desired CPU, Memory, GPU, and Disk specifications
      Service Type > Planned ComputeRequiredPlanned Compute가 설정된 자원 현황
      • In Use: Number of resources with Planned Compute that are currently in use
      • Configured: Number of resources with Planned Compute set
      • Coverage Preview: Amount applied per resource by Planned Compute
      Table. GPU Node Service Information Input Items
    • In the Required Information Input area, enter or select the required information.
      Category
      required status
      Detailed description
      Administrator accountRequiredSet the administrator account and password to be used when connecting to the server
      • Ubuntu OS is provided with a fixed root account
      Server name PrefixRequiredEnter a prefix to distinguish each GPU Node generated when the selected number of servers is 2 or more
      • Automatically generated in the form of user input value (prefix) + ‘-###
      • Must start with a lowercase English letter and be entered using lowercase letters, numbers, and special characters (-) within 3 to 11 characters
      • Must not end with a special character (-)
      Network SettingsRequiredSet the network where the GPU Node will be installed
      • VPC Name: Select a pre‑created VPC
      • General Subnet Name: Select a pre‑created standard Subnet
        • IP can be Auto‑generated or Manual Input; if Manual Input is selected, the user enters the IP directly
      • NAT: Available only when there is a single server and the VPC is attached to an Internet Gateway. Checking the option allows selection of a NAT IP. (Initially, only configurations with two or more servers can be created, so modify on the resource detail page)
      • NAT IP: Select a NAT IP
        • If no NAT IP is available to select, click the Create New button to generate a Public IP
        • Refresh button to view and select the created Public IP
        • Creating a Public IP incurs charges according to the Public IP pricing policy
      Table. GPU Node required information input items
    • In the Cluster Selection area, create or select a Cluster Fabric.
      Category
      required status
      Detailed description
      Cluster FabricRequiredConfiguration of GPU Node server groups that can apply GPU Direct RDMA together
      • Optimal GPU performance and speed can be secured only within the same Cluster Fabric
      • When creating a new Cluster Fabric, *New Input > select Node pool, then enter the name of the Cluster Fabric to be created
      • To add to an existing Cluster Fabric, Existing Input > select Node pool, then choose the previously created Cluster Fabric
      Table. GPU Node Cluster Fabric options
    • In the Additional Information Input area, enter or select the required information.
      Category
      required status
      Detailed description
      LockSelectionUsing a lock prevents actions caused by mistakes, such as terminating, starting, or stopping the server.
      Init ScriptSelectionScript to run at server startup
      • The Init Script must be selected differently depending on the image type
        • For Linux: Choose Shell Script or cloud-init
      tagSelectionAdd Tag
      • You can add up to 50 per resource
      • After clicking the Add Tag button, enter or select Key, Value values
      Table. GPU Node additional information input fields
  4. Summary Check the detailed information and estimated billing amount generated in the panel, and click the Create button.

  5. When the popup notifying creation opens, click the Confirm button.

    • When creation is complete, check the created resources on the GPU Node List page.
Caution
  • When creating a service, the GPU MIG/ECC settings are reset. * However, to apply the correct settings, perform an initial reboot, verify that the settings have been applied, and then use it.
  • For detailed information on resetting GPU MIG/ECC settings, refer to the GPU MIG/ECC 설정 초기화 점검 가이드.

Check GPU Node detailed information

The Multi-node GPU Cluster service can view and modify the full list of GPU Node resources and detailed information.

GPU Node Details page consists of Details, Tags, Job History tabs.

To view detailed information about the GPU Node, follow these steps.

  1. Click the All Services > Compute > Multi-node GPU Cluster > GPU Node menu. 1. Go to the Service Home page of the Multi-node GPU Cluster.

  2. On the Service Home page, click the GPU Node menu. 2. Go to the GPU Node list page.

    • Resource items other than the required columns can be added through the Settings button.
      Category
      required status
      Detailed description
      Resource IDSelectionUser-created GPU Node ID
      Cluster Fabric nameRequiredUser-created Cluster Fabric name
      Server nameRequiredUser-created GPU Node name
      Server typeRequiredServer type of GPU Node
      • The user can view the number of cores, memory capacity, and GPU type and count of the created resources
      imageRequiredUser-generated GPU Node image version
      IPRequiredIP of the GPU node created by the user
      statusRequiredStatus of the GPU Node created by the user
      Creation date and timeSelectionGPU Node creation timestamp
      Table. GPU Node resource list items
  3. GPU Node List page, click the resource to view detailed information. 3. Go to the GPU Node Details page.

    • GPU Server Details At the top of the page, status information and descriptions of additional features are displayed.
      CategoryDetailed description
      GPU Node statusStatus of the GPU Node created by the user
      • Creating: State while the server is being created
      • Running:: State when creation is complete and the server is available for use
      • Editing:: State while the IP is being changed
      • Unknown: Error state
      • Starting: State while the server is starting
      • Stopping: State while the server is stopping
      • Stopped: State when the server has stopped
      • Terminating: State while terminating
      • Terminated: State when termination is complete
      Server controlButton to change server status
      • Start: Start a stopped server
      • Stop: Stop a running server
      Service cancellationCancel service button
      Table. GPU Node status information and additional features

Detailed Information

On the GPU Node List page’s Details Tab, you can view the detailed information of the selected resource and edit the information if needed.

CategoryDetailed description
serviceservice name
Resource TypeResource Type
SRNUnique resource ID in Samsung Cloud Platform
  • In a GPU Node, it means the GPU Node SRN
Resource nameResource name
  • In the GPU Node service, it refers to the GPU Node name
Resource IDUnique resource ID in the service
ConstructorUser who created the service
Creation date and timeService creation date and time
ModifierUser who edited the service information
Modification date and timeDate and time the service information was modified
Server nameserver name
Node poolA collection of nodes that can be grouped into the same Cluster Fabric
Cluster Fabric nameUser-created Cluster Fabric name
Image/VersionServer OS image and version
Server typeCPU, memory, GPU, information display
Planned ComputeResource status with Planned Compute configured
LockIndicates whether Lock is enabled/disabled
  • When Lock is enabled, it prevents server termination/start/stop actions, avoiding accidental operations
  • If you need to change the Lock property value, click the Edit button to set it
NetworkGPU Node network information
  • VPC name, regular Subnet name, IP, Public NAT IP, and status
Block StorageBlock Storage information connected to the server
  • Volume name, disk type, capacity, status
Init ScriptView the Init Script content entered during server creation
Table. GPU Node detailed information tab items

note
If the VPC does not have an Internet Gateway attached, you cannot attach a Public NAT IP.

Tag

On the GPU Node List page’s Tag Tab, you can view the selected resource’s tag information, and add, modify, or delete it.

CategoryDetailed description
Tag listTag list
  • You can view the Key and Value information of the tag
  • Up to 50 tags can be added per resource
  • When entering a tag, you can search and select from the list of previously created Keys and Values
Table. GPU Node Tag Tab Items

Job History

GPU Node List page’s Job History Tab allows you to view the job history of the selected resource.

CategoryDetailed description
Task History ListResource Change History
  • Check operation details, operation time, resource type, resource name, event topic, operation result, operator information
  • Detailed Search button provides detailed search functionality
Table. GPU Node Job History Tab Detailed Information Items

Control GPU Node operation

If you need server control and management functions for the created GPU Node resources, you can perform tasks on the GPU Node List or GPU Node Details page. You can start and stop the running GPU Node resources.

Getting Started with GPU Node

You can start a stopped (Stopped) GPU Node. To start the GPU Node, follow these steps.

  1. All Services > Compute > Multi-node GPU Cluster Click the menu. 1. Go to the Service Home page of the Multi-node GPU Cluster.
  2. On the Service Home page, click the GPU Node menu. 2. Go to the GPU Node list page.
    • GPU Node List page, after selecting individual or multiple servers with the checkbox, you can Start via the More button at the top.
  3. GPU Node List page, click Resources. 3. Go to the GPU Node Details page.
    • On the GPU Node Details page, click the Start button at the top to start the server.
  4. Check the server status and complete the status change.

Stop GPU Node

You can stop a GPU node that is (Active). To stop the GPU Node, follow the steps below.

  1. All Services > Compute > Multi-node GPU Cluster Click the menu. 1. Go to the Service Home page of the Multi-node GPU Cluster.
  2. On the Service Home page, click the GPU Node menu. 2. Go to the GPU Node list page.
    • GPU Node List page, after selecting individual or multiple servers with the checkboxes, you can control them using the Stop button at the top.
  3. On the GPU Node List page, click Resources. 3. Go to the GPU Node Details page.
    • Click the 중지 button at the top of the GPU Node 상세 page to stop the server.
  4. Check the server status and complete the status change.

Terminate GPU Node

You can terminate unused GPU nodes to reduce operating costs. However, if you terminate the service, the running service may be discontinued immediately, so you should proceed with the termination only after fully considering the impact that may arise from the service interruption.

Caution
Please be aware that data cannot be recovered after terminating the service.

To cancel a GPU Node, follow these steps.

  1. All Services > Compute > Multi-node GPU Server Click the menu. 1. Go to the Service Home page of the Multi-node GPU Cluster.
  2. On the Service Home page, click the Cluster Fabric menu. 2. Go to the Cluster Fabric List page.
  3. On the Cluster Fabric List page, select the resources to cancel, and click the Cancel Service button.
    • Resources that use the same Cluster Fabric can be terminated simultaneously.
  4. When termination is complete, check on the GPU Node List page whether the resources have been terminated.
Notice

The cases where a GPU Node cannot be terminated are as follows.

  • When Block Storage(BM) is connected: Please disconnect the Block Storage(BM) connection first.
  • When File Storage is connected: Please disconnect the File Storage connection first.
  • If Lock is set: Please change the Lock setting to disabled and try again.
  • If the selection includes a server that cannot be terminated simultaneously: Please re-select only resources that can be terminated.
  • If the server you want to decommission has a different Cluster Fabric: Select only resources that use the same Cluster Fabric.
Reference
If all GPU Nodes in the Cluster Fabric are deleted, the Cluster Fabric is automatically deleted.
Monitoring Metrics
Cluster Fabric Management