The page has been translated by Gen AI.

Overview

Service Overview

GPU Server is a virtualized computing service that lets you freely allocate and use infrastructure resources such as CPU, GPU, and memory provided by the server, without having to purchase them individually, and allocate as much as needed at the required time. It is suitable for tasks that require fast computation speed, such as AI model experiments, predictions, and inference in a cloud environment, and you can flexibly select and use resources with optimized performance according to the type and scale of the work. The GPU Server provides the following features.

Provided features

  • GPU Server Management: Through a web-based console, users can directly Self Service create, delete, and modify GPU servers from provisioning to monitoring and billing.
  • Product offering by GPU quantity: Depending on the project’s purpose and scale, you can freely select the number of H100/A100 GPUs to configure a virtual server.
  • High‑Performance GPU Provision: We provide high‑performance GPU servers at physical‑server level using a pass‑through method.
  • Storage Connection: Provides additional attached storage besides the OS disk. * Block Storage, File Storage, Object Storage can be connected and used.
  • Strong Security Enforcement: Through the Security Group service, control inbound/outbound traffic exchanged with the external Internet or other VPC (Virtual Private Cloud) to securely protect the server.
  • Monitoring: You can view monitoring information such as the status of CPU, Memory, Disk, and GPU, which are computing resources, through the Cloud Monitoring service.
  • Network Configuration Management: The server’s subnet/IP can be conveniently changed from the values set at initial creation. * NAT IP provides a management feature that lets you enable or disable it as needed.
  • Key Pair method: To ensure a secure OS access method, we provide a Key Pair method instead of ID/PW login.
  • Image Management: You can create and manage Custom Images, and it provides sharing functionality between projects.
  • ServiceWatch Service Integration: You can monitor data through the ServiceWatch service.

Components

GPU Server provides GPUs, NVSwitch, and NVLink on top of virtualized computing resources.

caution
  • NVSwitch can be enabled and used only for instance types that allocate eight GPUs on a single GPU server.

Specifications by GPU Type

GPU (Graphic Processing Unit) performs the calculations required to generate the images that compose a computer screen, and because it is specialized for parallel processing, it can handle large amounts of data quickly, processing large‑scale parallel operations such as artificial intelligence (AI) and data analysis. The specifications of the GPU Types provided by the GPU Server service are as follows.

CategoryA100 TypeH100 TypeB300 Type
GPU ArchitectureNVIDIA AmpereNVIDIA HopperNVIDIA Blackwell Ultra
GPU Memory80 GiB80 GiB268 GiB
GPU Transistors54 billion 7N TSMC80 billion 4N TSMC208 billion 4NP TSMC
FP16 Tensor Core (Dense)312 TFLOPs989 TFLOPs2.25 PFLOPs
FP8 Tensor Core (Dense)Unsupported1,979 TFLOPs4.5 PFLOPs
FP4 Tensor Core (Dense)UnsupportedUnsupported13.5 PFLOPs
GPU Memory Bandwidth2,039 GB/s HBM2e3,352 GB/s HBM38 TB/s HBM3e
NVLink performanceNVLink 3NVLink 4NVLink 5
NVLink Signaling Rate25 GB/s (x12)25 GB/s (x18)50 GB/s (x18)
NVSwitch GPU-to-GPU bandwidth600 GB/s900 GB/s1.8 TB/s
Total NVSwitch aggregate bandwidth4.8 TB/s7.2 TB/s14.4 TB/s
Table. Specifications by GPU Type

NPU Type Specifications

The NPU (Network Processing Unit) is a processor specialized for AI inference operations, performing generative AI and various AI inference workloads based on high throughput and power efficiency. The specifications of the NPU Type provided by the GPU Server service are as follows.

CategoryFuriosa RNGD
ArchitectureTensor Contraction Processer
BF16256 TFLOPS
FP8512 TFLOPS
Memory BandwidthHBM3 1.5 TB/s
Memory CapacityHBM3 48 GB
Interconnect InterfacePCIe Gen5 x16
Table. Specifications by NPU Type

Server type

The server types offered by the GPU Server are as follows. For detailed information about the server types provided by GPU Server, see GPU Server 서버 타입.

CategoryServer typeCPU vCoreMemory(GB)GPU/NPU quantity
GPU-A100-1g1v16a1162341
GPU-A100-1g1v32a2324682
GPU-A100-1g1v64a4649364
GPU-A100-1g1v128a81281,8728
GPU-H100-2g2v12h1122341
GPU-H100-2g2v24h2244682
GPU-H100-2g2v48h4489364
GPU-H100-2g2v96h8961,8728
GPU-B300-3g3v16b1164801
GPU-B300-3g3v32b2329602
GPU-B300-3g3v64b4641,9204
GPU-B300-3g3v128b81283,8408
NPU-RNGD-1n1v8r181061
NPU-RNGD-1n1v16r2162122
NPU-RNGD-1n1v32r4324244
NPU-RNGD-1n1v64r8648488
Table. GPU Server server type

OS and driver version

The operating systems (OS) supported by the GPU Server are as follows. Please be aware that B300-type GPUs are supported only from a certain GPU version onward, so choose your image accordingly.

OSOS versionDriver versionServer type classification
Ubuntu24.04ND 580.126.20GPU-B300-3, GPU-H100-2, GPU-A100-1
Ubuntu24.04ND 570.195.03GPU-H100-2, GPU-A100-1
Ubuntu24.04FRD 2026.2.0NPU-RNGD-1
Ubuntu22.04ND 535.183.06GPU-H100-2, GPU-A100-1
RHEL9.6ND 580.126.20GPU-B300-3, GPU-H100-2, GPU-A100-1
RHEL8.1ND 580.126.20GPU-B300-3, GPU-H100-2, GPU-A100-1
RHEL8.1ND 535.183.06GPU-H100-2, GPU-A100-1
Table. GPU Server OS and driver version

Constraints

The GPU Server has the following constraints.

  • We are providing Ubuntu22.04 as the current OS version.
  • GPU is provided via Pass-through.
  • Physical GPU cards are created in units of 1, 2, 4, or 8 cards each. * –>

Preceding Service

This service must be installed in advance before creating it. Please prepare by referring to the user guide provided in advance.

Service CategoryserviceDetailed description
NetworkingVPCA service that provides an isolated virtual network in a cloud environment
NetworkingSecurity GroupVirtual firewall that controls server traffic
Table. GPU Server Preliminary Service
Release Note
Server type