1 - Cloud Hadoop

1.1 - Overview

Service Overview

Cloud Hadoop is a service for easily and quickly analyzing large-scale data, providing a Hadoop cluster (computing resources, management tools, and applications) used for big data processing and analysis in the SCP environment.

Features

Cloud Hadoop provides an automated cluster creation service through the Hadoop Manager and the Hadoop Ecosystem(ecosystem) composed of Spark, HDFS(Hadoop distributed file system), Hive, etc., enabling anyone to easily build, optimize, or flexibly scale infrastructure for big data analysis.

Service Diagram

Architecture diagram
Figure. Cloud Hadoop Architecture Diagram

Provided features

Cloud Hadoop provides the following features.

  • Provide Hadoop Cluster as a cloud service

    • Providing a Hadoop Cluster through automated cluster installation in the SDS Cloud environment
    • Perform essential operational activities for cluster management (cluster operation/monitoring)
    • Provides a Hadoop ecosystem with verified interoperability and allows users to access the server (VM)
  • Offer the Hadoop service stack as separate products (increase nodes per product)

    • Minimum node allocation per product for stable service operation
    • Providing diverse product selection opportunities to meet user needs and reduce costs
  • Providing user-friendly features for Hadoop services

    • Provides installation and management functions for each Hadoop ecosystem, optimal configuration values, and version management features.
    • Provide an integrated monitoring dashboard for system resources
    • Provides Service Failure Alert feature

Component

We package the major components of the Hadoop ecosystem to deliver an enterprise data cloud.

Service Configuration

Cloud Hadoop provides the following services.

  • Basic Installation Service
    • HDFS 3.3.6
    • YARN 3.3.6
    • Hbase 2.4.17
    • Hive 3.1.2
    • Tez 0.9.1
    • Hue 4.11.0
    • Solr 8.11.4
    • Spark 3.4.1
    • Zookeeper 3.8.5
  • Additional Option Service
    • Data Governance: Atlas 2.1.0, Ranger 2.1.0
    • Analytical Data Warehouse: Iceberg 1.8.0, Kyuubi 1.10.2
    • Data Ingestion: Sqoop 1.4.7, Kafka 3.9.1, Flume 1.11.0

Server type

The server types supported by Cloud Hadoop are as follows.

Category
exampleDetailed description
Server typeStandardProvided server types
  • Standard: Configured with the commonly used standard specifications (vCPU, Memory)
  • High Capacity: Large-capacity server specifications of 24 cores or more
Server sizes1v4m32Provided server specifications
  • vCPU 4
Table. Cloud Hadoop server type

The minimum specifications for using Cloud Hadoop are as follows.

Category
AlgebraInstance size (user-selected value)
Master2(fixed)CPU: 4 Core
Memory: 32 GB
Worker3(minimum)CPU: 4 Core
Memory: 32 GB
Data GovernanceNone
Analytical Data WarehouseNone
Ingestion3 (minimum)CPU: 4 Core
Memory: 32 GB
Table. Cloud Hadoop Minimum Specifications

Provision status by region

Cloud Hadoop is available in the following environments.

regionProvision status
Korea West (kr-west1)Provide
Korea East (kr-east1)Provide
South Korea South1(kr-south1)Not provided
South Korea South 2 (kr-south2)Not provided
South Korea 3 (kr-south3)Not provided
Table. Cloud Hadoop regional availability status

Preliminary Service

This is a list of services that need to be pre-configured before creating the service. Please refer to the guide provided for each service and prepare in advance.

Service CategoryserviceDetailed description
NetworkingVPCA service that provides an isolated virtual network in a cloud environment
NetworkingSecurity GroupVirtual firewall that controls server traffic
StorageObject StorageObject storage that simplifies data storage and retrieval
Table. Cloud Hadoop preliminary service

1.1.1 - ServiceWatch metric

You can view Virtual Server metrics in ServiceWatch for servers created in Cloud Hadoop. Like Virtual Server, the metrics provided by default monitoring are data collected at 5‑minute intervals. In the Virtual Server detailed view, enabling detailed monitoring allows you to view data collected at 1‑minute intervals. For more details, Virtual Server > Enable ServiceWatch Detailed Monitoring

information
  • The basic monitoring and detailed monitoring of Cloud Hadoop are provided with the same metrics as Virtual Server, and the namespace is also provided as Virtual Server.
Reference
Refer to the ServiceWatch guide for how to view metrics in ServiceWatch.
Reference
Refer to the ServiceWatch Agent guide for how to collect metrics using the ServiceWatch Agent.

Basic Metrics

The following are the basic metrics for the Virtual Server namespace.

The indicators whose names are displayed in bold below are the key indicators selected from the basic metrics provided by Virtual Server. Key metrics are used to build service dashboards that are automatically created for each service in ServiceWatch.

Each metric provides guidance in the user guide on which statistical value is meaningful when viewing that metric, and among the meaningful statistics, the values shown in bold are the primary statistics. In the service dashboard, you can view key metrics using primary statistical values.

Performance itemsDetailed descriptionunitmeaningful statistics
Instance StateInstance status display
  • 1 - Active
  • 0 - Off
None
  • Total
CPU UsageCPU usagePercent
  • Average
  • Maximum
  • Minimum
Disk Read BytesBytes read from block device (bytes)Bytes
  • Total
  • Average
  • Maximum
  • Minimum
Disk Read RequestsNumber of read requests on a block deviceCount
  • Total
  • Average
  • Maximum
  • Minimum
Disk Write BytesWrite capacity (bytes) on block deviceBytes
  • Total
  • Average
  • Maximum
  • Minimum
Disk Write RequestsNumber of write requests on block deviceCount
  • Total
  • Average
  • Maximum
  • Minimum
Network In BytesReceived bytes on the network interfaceBytes
  • Total
  • Average
  • Maximum
  • Minimum
Network In DroppedNumber of packet drops received on the network interfaceCount
  • Total
  • Average
  • Maximum
  • Minimum
Network In PacketsNumber of packets received on the network interfaceCount
  • Total
  • Average
  • Maximum
  • Minimum
Network Out BytesData transmitted from the network interface (bytes)Bytes
  • Total
  • Average
  • Maximum
  • Minimum
Network Out DroppedNumber of packet drops transmitted from the network interfaceCount
  • Total
  • Average
  • Maximum
  • Minimum
Network Out PacketsNumber of packets transmitted on the network interfaceCount
  • Total
  • Average
  • Maximum
  • Minimum
Table. Virtual Server Basic Metrics

1.2 - How

Users can create the service by entering the required Cloud Hadoop information and selecting detailed options through the Samsung Cloud Platform Console.

Create Cloud Hadoop

You can create and use the Cloud Hadoop service from the Samsung Cloud Platform Console.

To create Cloud Hadoop, follow the steps below.

  1. All Services > Data Analytics > Cloud Hadoop Click the menu. 1. Go to the Service Home page of Cloud Hadoop.

  2. On the Service Home page, click the Create Cloud Hadoop button. 2. Go to the Create Cloud Hadoop page.

  3. Cloud Hadoop Creation page: enter the information required to create the service and select detailed options.

    • In the Image and version selection area, select the required information.
      Category
      Required
      Detailed description
      imageRequiredSelect the type of image provided
      • Cloud Hadoop With Ubuntu 22.04
      Image versionRequiredSelect version of the selected image
      • Provide version list of the provided image
      Table. Cloud Hadoop image and version selection options
    • In the Service Information Input area, enter or select the required information.
      Category
      Required status
      Detailed description
      Server name PrefixRequiredThe server name on which Cloud Hadoop will be installed
      • must start with a lowercase English letter and be entered using lowercase letters, numbers, and the special character (-) with a length of 3 to 13 characters
      • A postfix such as 001, 002 is appended based on the server name to generate the actual server name
      Cluster nameRequiredCluster name formed by the servers
      • Enter using English letters, 3 ~ 20 characters
      • A cluster is a unit that groups multiple servers
      Planned ComputeSelectFor servers where Cloud Hadoop is installed, selecting a Planned Compute contract allows usage at a discounted price
      • All Services > Financial Management > Planned Compute menu can be requested
      Master Node > Master Node CountRequiredNumber of Master nodes
      • The number of Master nodes is fixed at two per Hadoop cluster
      • A Master node is the node where Hadoop Master is installed and provides the default HA (high availability) configuration
      • The Master node includes Hadoop Manager and various Hadoop ecosystem components installed together
      Master Node > Server TypeRequiredCPU and Memory types for distributed data processing
      • Standard-1: standard specifications commonly used
      • High Capacity-2: large-capacity server with 24 vCores or more
      • Recommended specifications: vCPU 8, Memory 64G
      Master Node > Block StorageRequiredBlock Storage type to be used for the Master node
      • Basic OS: Area where the engine is installed
      • DATA: Data file storage area
        • After selecting the storage type, enter the capacity (see Block Storage 생성하기 for details on each Block Storage type)
          • SSD: High‑performance standard volume
          • HDD: Standard volume
        • Capacity can be entered in multiples of 8 within the range 25 to 1,536
      • Delete on termination: When the server is terminated, the volume is terminated as well, but volumes with snapshots are not deleted even when Delete on termination is enabled.
      • Add Disk: Data storage area
        • After selecting Use, enter the capacity of the storage
        • Click the + button to add storage, or the x button to delete. Up to 9 can be added
        • Capacity can be entered in multiples of 8 within the range 25 to 1,536
      Worker Node > Worker Node countRequiredNumber of Worker nodes
      • Worker nodes can be selected from 3 to 90
      • Worker nodes are the nodes where Hadoop data nodes and the Resource Manager are installed, and they process and store distributed data
      Worker Node > Server TypeRequiredCPU and Memory types for distributed data processing
      • Standard-1: Standard specification commonly used
      • High Capacity-2: Large server with 24 vCores or more
      • Recommended specifications: vCPU 8, Memory 64G
      Worker Node > Block StorageRequiredBlock Storage type to be used on the Worker node
      • Basic OS: Area where the engine is installed
      • DATA: Data file storage area
        • After selecting the storage type, enter the capacity (refer to Block Storage creation for details on each Block Storage type)
          • SSD: High‑performance standard volume
          • HDD: Standard volume
        • Capacity can be entered in multiples of 8 within the range 25 to 1,536
      • Delete on termination: When the server is terminated, the volume is terminated as well; volumes with snapshots are not deleted even when Delete on termination is enabled.
      • Add Disk: Data storage area
        • Select Use, then enter the capacity of the storage
        • Click the + button to add storage, or the x button to delete. Up to 9 can be added
        • Capacity can be entered in multiples of 8 within the range 25 to 1,536
      Data GovernanceSelectAdditional Hadoop ecosystem installation for data governance
      • If you select Use, Atlas and Ranger are installed automatically
      • Cannot be modified or removed after creation
      Analytical Data WarehouseSelectionAdditional installation of Hadoop ecosystem for fast data analysis
      • If you select Use, Iceberg and Kyuubi will be installed automatically
      • Cannot be modified or removed after creation
      Data IngestionSelectAdditional installation of Hadoop ecosystem for data collection and loading
      • If you select Use, Kafka, Flume, and Sqoop will be installed automatically
      Data Ingestion > Ingestion Node countSelectNumber of Ingestion nodes
      • Ingestion nodes can be selected from 3 to 10
      • Worker nodes are the nodes where Hadoop data nodes and the Resource Manager are installed, and they process and store distributed data
      Data Ingestion > Server TypeSelectCPU and Memory types for distributed data processing
      • Standard-1: Standard specification commonly used
      • High Capacity-2: Large server with 24 vCores or more
      • Recommended specifications: vCPU 8, Memory 64G
      Data Ingestion > Block StorageSelectBlock Storage type to be used for the Ingestion node
      • Basic OS: Area where the engine is installed
      • DATA: Data file storage area
        • After selecting the storage type, enter the capacity (see Block Storage 생성하기 for details on each Block Storage type)
          • SSD: High‑performance general volume
          • HDD: General volume
        • Capacity can be entered in multiples of 8 within the range 25 to 1,536
      • Add Disk: Data storage area
        • After selecting Use, enter the storage capacity
        • Click the + button to add storage, or the x button to delete. Up to 9 can be added
        • Capacity can be entered in multiples of 8 within the range 25 to 1,536, and up to 9 can be created
      Object Storage bucketSelectionObject Storage to be used in the cluster
      • After selecting Bucket selection, select the Object Storage bucket
      • You can add up to 10, to delete click the x button
      • After adding a bucket, to set access permission for that bucket, select server resources from the All Services > Object Storage list > the relevant Object Storage Details > Access Control > Allow Server Resources menu
      Table. Cloud Hadoop Service Information Input Items
    • Required Information Input area: enter or select the required information.
      Category
      Required
      Detailed description
      Enter PrivateLink information

      | Enter the authentication key for PrivateLink connection

      • Create authentication key: IAM > My info > Authentication Key Management tab > Create Authentication Key button click
      • Copy authentication key: IAM > My info > Authentication Key Management tab > Click the generated authentication key > Authentication Key Details > Basic Information tab > Authentication Key > View button click > Copy the authentication key from the popup window
      • Access Key: Enter Access Key; can be entered only when first applying for the service
      • Secret Key: Enter Secret Key; can be entered only when first applying for the service
      | | Cloud Hadoop Manager account |

      | Enter the account and password to log in to Cloud Hadoop Manager

      • Account Name: Enter the account to use for login
      • Password: Enter the password to use for login
      • Confirm Password: Re-enter the password
      | | Network > Common Settings |

      | Network settings for servers created by the service

      • Select when you want to apply the same settings to all installed servers
      • Select the pre‑created VPC, Availability Zone, and Subnet
      • IP is generated automatically
      • Public NAT: Available when the VPC is connected to an Internet Gateway and the Subnet is of type Public. If Use is checked, NAT IP can be configured
      | | Network > Server-specific Settings | | Network settings for servers created by the service
      • Select when you want to apply different settings for each server being installed
      • Select a pre‑created VPC, Availability Zone, and Subnet
      • Displayed automatically based on the selected node
      • Enter the IP for each server
      • Public NAT: Available when the VPC is connected to an Internet Gateway and the Subnet is of Public type. Checking Use enables NAT IP configuration. For more information, see Public IP 생성하기
      | | Security Group | | Add Security Group
      • Click the Select button to choose from the list
      • Before the service creation is complete, you can delete the added Security Group by clicking the x button on the right
      | | Keypair | | Select the user authentication method used when connecting to a Virtual Server
      • Default login account per OS
        • Alma Linux: almalinux
        • Oracle Linux: cloud-user
        • RHEL:cloud-user
        • Rocky Linux: rocky
        • Ubuntu: ubuntu
        • Windows: sysadmin
      |

      Table. Cloud Hadoop required information entry fields
      Caution
      • For PrivateLink connections, you must enter an authentication key generated as a permanent key, and you must not delete that key. *
        If the authentication key expires or is deleted, making the key invalid, it may cause issues with resource changes and service termination in the Cloud Hadoop service.
      • When using a public subnet and assigning a public IP, you may be exposed to security attacks such as external hacking and malware infection.
      information
      • When creating a Cloud Hadoop service, only one Security Group can be selected, but after the service is created, up to four Security Groups can be selected, including the initially chosen Security Group. * However, the Security Group selected when creating the service for the first time cannot be modified or deleted.
      • If Cloud Hadoop is installed correctly, API communication between the installed Cloud Hadoop service and the Samsung Cloud Platform Console may occur continuously for the following reasons.
        • Changing resources of Cloud Hadoop service (adding nodes and resources)
        • Changing the state of the Cloud Hadoop service (start, stop, restart, and termination)
        • Check the status of the Cloud Hadoop service (Health Check)
    • In the Additional Information Input area, enter or select the required information.
      Category
      Required
      Detailed description
      time zoneSelect the time zone the Database will use
      tagSelectionAdd Tag
      • Add Tag Click the button to create and add a tag, or add an existing tag
      • Up to 50 can be added
      • The newly added tags are applied after the service creation is completed
      Table. Cloud Hadoop additional information input items
  4. Summary Check the detailed information and estimated charges generated in the panel, and click the Complete button.

    • Once creation is complete, check the created resource on the Resource List page.

Check Cloud Hadoop detailed information

The Cloud Hadoop service allows you to view and edit the full list of resources and detailed information. Cloud Hadoop Details page consists of Details, Tags, Job History tabs.

To view detailed information about the Cloud Hadoop service, follow these steps.

  1. All Services > Data Analytics > Cloud Hadoop Please click the menu. 1. Go to the Service Home page of Cloud Hadoop.
  2. On the Service Home page, click the Cloud Hadoop menu. 2. Go to the Cloud Hadoop List page.
  3. On the Cloud Hadoop List page, click the resource to view detailed information. 3. Cloud Hadoop Details Navigate to the page.
    • Cloud Hadoop Details At the top of the page, status information and additional feature information are displayed.
      CategoryDetailed description
      statusCloud Hadoop service status
      • Creating: Creating
      • Running: Creation complete, service available
      • Updating: Updating settings
      • Stopping: Stopping
      • Starting: Starting
      • Stopped: Stopped
      • Restarting: Restarting
      • Terminating: Terminating service
      • Error: Error during creation or service abnormal state
      • Undeployed: Error during deployment
      StartStart operating the discontinued service
      StopForce terminate service
      RestartRestart the service
      Add Worker NodeAdd a server with the same specifications as the previously created Worker node to the cluster.
      Service terminationTerminate the entire Cloud Hadoop service and server
      Table. Cloud Hadoop status information and additional features
Reference
  • The status indicator shows the status of the Cloud Hadoop service, and the server status can be checked in the server information.
  • Start, Stop, Restart buttons control only the Cloud Hadoop service, while server control can be managed from the Compute > Virtual Server list.

Detailed Information

On the Cloud Hadoop List page, you can view detailed information of the selected resource and edit the information if needed.

CategoryDetailed description
Server InformationServer information configured in this cluster
serviceService name
Resource TypeResource Type
SRNUnique resource ID in Samsung Cloud Platform
  • means the cluster SRN
Resource nameResource name
  • means cluster name
Resource IDUnique resource ID in the service
ConstructorUser who created the service
Creation Date/TimeService creation date and time
ModifierUser who edited the service information
Modification date and timeDate and time the service information was modified
Image versionOS and service image version
Cluster nameCluster name of the configured servers
Planned ComputeResource status with Planned Compute configured
Manager access URLCloud Hadoop Manager access URL
time zoneThe standard time zone for the service
PrivateLink informationAccess Key, Secret Key information
NetworkVPC, Availability Zone, Subnet information
Security GroupSecurity Group List
Keypair nameCreated/Selected Keypair Name
Basic ServiceCloud Hadoop Basic Service Stack List
Option ServiceCloud Hadoop option service stack list
  • Data Governance, Analytical Data Warehouse, Data Ingestion
MasterServer type, base OS, and Disk information for the Master node
  • If you need to modify the server type, click the Edit button next to the server type to configure it
    • Modifying the server type requires a server reboot
  • If you need to expand storage, click the Edit button next to the storage capacity to expand it
  • If you need to add storage, click the Add Disk button to add it
WorkerServer type, default OS, and disk information for the Worker node
IngestionServer type, default OS, and Disk information for the Ingestion node
Object Storage bucketObject Storage List
Table. Cloud Hadoop detailed information items

Tag

Cloud Hadoop List page lets you view the tag information of the selected resource, and add, modify, or delete it.

CategoryDetailed description
Tag ListTag list
  • You can view the Key, Value information of the tag
  • Up to 50 tags can be added per resource
  • When entering tags, you can search and select from the list of previously created Keys and Values
Table. Cloud Hadoop tag tab items

Job History

Cloud Hadoop List page allows you to view the operation history of the selected resource.

CategoryDetailed description
Task History ListResource Change History
  • Check operation details, operation date and time, resource type, resource name, operation result, and operator information
  • Click a resource in the list to display the Operation History Details popup
  • Provides detailed search functionality via the Detailed Search button
Table. Cloud Hadoop Job History Tab Detailed Information Items

Managing Cloud Hadoop Resources

If you need to modify the existing configuration options of a created Cloud Hadoop resource or require additional configuration, you can perform the work on the Cloud Hadoop Details page.

If expansion of the Cloud Hadoop cluster is needed due to increased workload or other reasons, you can add Worker nodes with the same specifications as the existing Worker nodes.

Notice
  • Each Cloud Hadoop cluster can use up to 10 Worker nodes.
  • When adding nodes, all settings except the number of nodes to add and the IP/NAT IP are fixed to the configuration entered during the service application.
  • If adding a node fails, contact the Samsung Cloud Platform service desk for troubleshooting.

Worker Node 추가 (네트워크 설정: 공통 설정) {#network}

You can add a Worker node to a Cloud Hadoop cluster whose network settings are created as common settings.

  1. Click the All Services > Data Analytics > Cloud Hadoop menu. 1. Navigate to the Service Home page of Cloud Hadoop.
  2. On the Service Home page, click the Cloud Hadoop menu. 2. Go to the Cloud Hadoop List page.
  3. Cloud Hadoop List On the page, click the resource where you want to add a node. 3. Cloud Hadoop Details Navigate to the page.
  4. Click the Add Worker Node button. 4. Go to the Add Worker Node page.
  5. After selecting Worker Node count, click the Complete button.
Reference
  • All settings, including the server name of each Worker node, are fixed to the configuration entered when applying for the service.

Add Worker Node (Network configuration: per-server settings)

You can add Worker nodes to a Cloud Hadoop cluster whose network configuration is set as per-server settings.

To add a Worker node, follow the steps below.

  1. All Services > Data Analytics > Cloud Hadoop menu, please click. 1. Navigate to the Service Home page of Cloud Hadoop.
  2. On the Service Home page, click the Cloud Hadoop menu. 2. Cloud Hadoop List navigate to the page.
  3. Cloud Hadoop List page, click the resource you want to add a node to. 3. Cloud Hadoop Details Go to the page.
  4. Add Worker Node Click the button. 4. Go to the Add Worker Node page.
  5. Please select the Worker Node count. 5. The server configuration area is automatically added based on the number of selected nodes.
  6. In the added server configuration area, enter IP and NAT IP, then click the Complete button.
Reference
  • All settings, including the server name of each Worker node, are fixed to the configuration entered when applying for the service.

Security Group Change

To change the Security Group of Cloud Hadoop, follow these steps.

  1. All Services > Data Analytics > Cloud Hadoop Click the menu. 1. Navigate to the Service Home page of Cloud Hadoop.
  2. On the Service Home page, click the Cloud Hadoop menu. 2. Go to the Cloud Hadoop List page.
  3. On the Cloud Hadoop List page, click the resource whose Security Group you want to change. 3. Cloud Hadoop Details Go to the page.
  4. Click the Edit button of Security Group on the detail information page. 4. Security Group selection The popup window opens.
  5. Search for the Security Group you want to add, then select the checkbox. 5. The selected Security Group is displayed in the list below.
  6. Click Confirm. 6. The selected Security Group will be applied.
Information
  • When creating a Cloud Hadoop service, you can select up to four Security Groups, including the Security Group you chose. * However, the Security Group selected when creating the service for the first time cannot be modified or deleted.

Add optional service

You can additionally install the Cloud Hadoop ecosystem (Data Governance, Analytical Data Warehouse, Data Ingestion).

Data Governance/Analytical Data Warehouse addition

To install Data Governance and Analytical Data Warehouse additionally, follow the steps below.

  1. All Services > Data Analytics > Cloud Hadoop Click the menu. 1. Navigate to the Service Home page of Cloud Hadoop.
  2. On the Service Home page, click the Cloud Hadoop menu. 2. Cloud Hadoop List navigate to the page.
  3. Cloud Hadoop List Click the resource on the page where you want to add an optional service. 3. Cloud Hadoop Details Go to the page.
  4. On the detail information page, click the Add button for the option service you want to add. 4. The notification popup opens.
  5. After reviewing the contents of the popup window, click the Confirm button. 5. The option service will be added automatically.
    • It may take some time depending on the scale.

Data Ingestion addition

To install Data Ingestion additionally, follow the steps below.

  1. All Services > Data Analytics > Cloud Hadoop menu, please click. 1. Go to the Service Home page of Cloud Hadoop.
  2. On the Service Home page, click the Cloud Hadoop menu. 2. Go to the Cloud Hadoop List page.
  3. Cloud Hadoop List On the page, click the resource where you want to add an optional service. 3. Cloud Hadoop Details navigate to the page.
  4. On the detail information page, click the Add button of Data Ingestion. 4. Add Data Ingestion Navigate to the page.
  5. After selecting the number of Ingestion Nodes, server type, and storage type and capacity, click the Complete button. 5. The option service will be added automatically.
    • It may take some time depending on the scale.

Change Server Type

You can change the server type of the Master node, Worker node, or Ingestion node in Cloud Hadoop.

To change the server type, follow the steps below.

Caution
  • If the server type is configured as Standard, it cannot be changed to High Capacity. * If you want to change to High Capacity, create a new service.
  • If you modify the server type, a server restart is required. * Please separately verify any software license modifications or software settings and their implementation due to specification changes.
  1. Click the All Services > Data Analytics > Cloud Hadoop menu. 1. Go to the Service Home page of Cloud Hadoop.
  2. On the Service Home page, click the Cloud Hadoop menu. 2. Go to the Cloud Hadoop List page.
  3. Cloud Hadoop List page, click the resource whose server type you want to change. 3. Cloud Hadoop Details Navigate to the page.
  4. On the detail information page, click the Edit button of the Server Type for the node you want to change. 4. Edit Server Type The popup window opens.
  5. After selecting the server type, click the Confirm button. 5. The notification popup opens.
    • Scale-Down of server type is not allowed.
  6. After reviewing the contents of the popup window, click the Confirm button.
    • The entire server on the node will be updated to the requested specifications, and the Cloud Hadoop cluster will restart.

Expanding Storage

Storage added to the data area can be expanded up to a maximum of 12 TB based on the initially allocated capacity. You can expand storage without stopping Cloud Hadoop operation, and if it is configured as a cluster, all nodes are expanded simultaneously.

notice
  • Storage capacity cannot be reduced and can only be expanded.
  • It can be expanded up to a maximum of 12 TB, and if more than 12 TB is required, it can be expanded through a service request.
  • It may take some time for the expansion to be completed after a request for expansion.

To increase storage capacity, follow the steps below.

  1. All Services > Data Analytics > Cloud Hadoop Click the menu. 1. Navigate to the Service Home page of Cloud Hadoop.
  2. On the Service Home page, click the Cloud Hadoop menu. 2. Cloud Hadoop List Navigate to the page.
  3. On the Cloud Hadoop List page, click the resource you want to expand capacity for. 3. Cloud Hadoop Details Go to the page.
  4. On the detail information page, click the Edit button of the node’s Disk you want to expand. 4. Disk Edit The popup window opens.
  5. After entering the number of units, click the Confirm button. 5. The notification popup opens.
    • You can set the capacity by entering the number of units provided in 8 GB increments.
  6. After reviewing the contents of the popup window, click the Confirm button.
    • It may take some time depending on the scale.

Add storage

If the storage space allocated to the data area exceeds 12 TB, additional storage can be added. When configured as a cluster, they are added simultaneously for each node type.

안내
  • The storage capacity can be set up to a maximum of 12 TB.
  • It may take some time for a storage addition request to be fully completed.

Follow the steps below to add storage.

  1. All Services > Data Analytics > Cloud Hadoop Click the menu. 1. Navigate to the Service Home page of Cloud Hadoop.
  2. On the Service Home page, click the Cloud Hadoop menu. 2. Go to the Cloud Hadoop List page.
  3. On the Cloud Hadoop List page, click the resource where you want to add storage. 3. Cloud Hadoop Details Navigate to the page.
  4. On the detail information page, click the Add Disk button of the node you want to add storage to. 4. Add Disk The popup window opens.
  5. Select the disk type, enter the capacity, and then click the Confirm button. 5. The notification popup opens.
    • If encryption is configured on the existing Block Storage, encryption will also be applied to the additional Disk.
    • If you configure it by selecting HDD, performance degradation may occur.
  6. After reviewing the contents of the popup window, click the Confirm button.
    • It may take some time depending on the scale.

Connecting to Cloud Hadoop

Follow these steps to access Cloud Hadoop.

  1. Check the IP of the Windows system (PC) that will connect to Cloud Hadoop.
    • Since external access is required, you need to check the system’s NAT IP.
  2. Add the following content to the hosts file on Windows.
    • VM host IP of the Cloud Hadoop cluster
    • VM host name of the Cloud Hadoop cluster
  3. Add the following rule to the Security Group you selected when applying for the Cloud Hadoop service.
    • Category: Inbound
    • Protocol: TCP
    • Target address: Windows system IP
    • Port: 7080
  4. On the Windows system you want to connect to, launch the Chrome browser and then access the Cloud Hadoop Manager URL.

Apache Hadoop Ecosystem Target IP/Port Information

Item | Protocol | Source | Target IP | Port | Remarks

ItemProtocalSourceTarget IPPortRemarks
ManagerTCPUser IPManager7080Cloud Hadoop Manager
HDFSCPUser IPMaster8042nodemanager web http
HDFSTCPUser IPMaster8044nodemanager web https
HDFSCPUser IPMaster8088resource manager web http
HDFSTCPUser IPMaster8090resource manager web https
HDFSTCPUser IPMaster8188timelneservice web http
HDFSTCPUser IPMaster8190timelneservice web https
HDFSTCPUser IPMaster9093alert manager
HDFSTCPUser IPMaster17000hbase master
HDFSTCPUser IPMaster17010hbase master web
HDFSTCPUser IPMaster17030hbase regionserver info
HDFSTCPUser IPMaster19090hbase thriftserver
HDFSTCPUser IPMaster19095hbase thriftserver info
HDFSTCPUser IPMaster19888Job History Server Web
HDFSTCPUser IPMaster50070name node web http
HDFSTCPUser IPMaster50075data node web http
AtlasTCPUser IPMaster21000atlas web http
AtlasTCPUser IPMaster21443atlas web https
HiveTCPUser IPMaster10000Hive sever2 thrift binary
HiveTCPUser IPMaster10001Hive sever2 thrift http
HiveTCPUser IPMaster10004Hive sever2 web binary
HiveTCPUser IPMaster10002Hive sever2 web http
HiveTCPUser IPMaster10005Hive sever2 HA web http
KerberosTCPUser IPMaster88key distribution server
KerberosTCPUser IPMaster749kadmin server
RangerTCPUser IPMaster9292ranger kms http
RangerTCPUser IPMaster6080ranger web http
SolrTCPUser IPMaster8983solr
SolrTCPUser IPMaster8988solr HA web http
SparkTCPUser IPMaster18080spark history server web http
SparkTCPUser IPMaster18082spark history server web https
TezTCPUser IPMaster8780tez ui
MonitoringTCPUser IPMaster7100prometheus web http
CmakTCPUser IPMaster19000cmak web http
HA ProxyTCPUser IPMaster38404HA Proxy web http
HueTCPUser IPMaster8000HUE web http
HueTCPUser IPMaster8005Hue HA web http
LLAPTCPUser IPMaster15002llap web http
Table. Hadoop ecosystem Target IP/Port information items

Terminate Cloud Hadoop

You can terminate unused Cloud Hadoop to reduce operating costs.

Caution
  • Data cannot be recovered after the service is terminated.
  • When the service is canceled, both the Cloud Hadoop service and the servers are terminated.
  • If you cancel the service, the active service will be stopped immediately. * Proceed with the termination after fully considering the impact that may arise from service interruption.

To cancel the service, follow these steps.

  1. Click the All Services > Data Analytics > Cloud Hadoop menu. 1. Navigate to the Service Home page of Cloud Hadoop.
  2. On the Service Home page, click the Cloud Hadoop menu. 2. Cloud Hadoop List Navigate to the page.
  3. On the Cloud Hadoop List page, select the resource to be terminated, then click the Terminate Service button. 3. The notification popup opens.
  4. Check the contents of the popup window, enter the name of the resource to be terminated, and then click the Confirm button.
  5. When the termination request is completed, check on the Cloud Hadoop list page whether the resource has been terminated.
    • It may take some time depending on the scale.

1.3 - API Reference

API Reference

1.4 - Release Note

Cloud Hadoop

2025.12.16
NEW Official release of Cloud Hadoop service
  • The Cloud Hadoop service for easy and fast analysis of large-scale data has been launched.
  • We provide an automated cluster creation service through the Hadoop Ecosystem and Hadoop Manager.

2 - Event Streams

2.1 - Overview

Service Overview

Event Streams provides fully managed creation and configuration of the open-source Apache Kafka for large-scale, high-volume message data processing. Samsung Cloud Platform automates the creation and configuration of Apache Kafka through a web-based console, allowing users to configure the main components of Apache Kafka—Broker, Zookeeper, and AKHQ—in either a single or clustered setup.

The Event Streams cluster consists of multiple Broker nodes; you can install between 1 and 10 Brokers, typically deploying three or more. Zookeeper can be installed separately to manage the distributed Brokers, but if not installed separately, it is installed on the Broker nodes. Additionally, we provide AKHQ (Apache Kafka HQ), a tool for managing Kafka, allowing users to perform cluster operation and management through it.

Provided features

Event Streams provides the following features.

  • Auto Provisioning: You can configure and set up an Apache Kafka cluster via the UI.
  • Operation Control Management: Provides functionality to control the status of running servers. In addition to starting and stopping the cluster, restarting is possible to apply configuration changes.
  • AKHQ provision: We provide AKHQ, a tool for managing Kafka, enabling users to manage and monitor clusters.
  • Add Broker node: If expansion is required to improve cluster performance and stability, you can add a node with the same specifications as the existing Broker nodes.
  • Parameter management: You can configure and modify parameters related to performance improvement and security.
  • Monitoring: CPU, memory, performance monitoring information can be accessed via Cloud Monitoring and Servicewatch.

Component

Event Streams provides pre‑validated engine versions and various server types in accordance with its open‑source support policy. Users can select and use them based on the scale of the service they wish to configure.

Engine version

The engine versions supported by Event Streams are as follows.

Technical support can be used until the supplier’s EoTS (End of Technical Service) date, and the EOS date when new creation is stopped is set to six months before the EoTS date.

The EOS and EoTS dates may change according to the supplier’s policy, so please refer to the supplier’s license management policy page for details.

Provided versionEoS DateEoTS Date
3.8.02026-07 (planned)2026-12-02
3.9.12026-09 (planned)2027-02-19
Table. Event Streams Supported Engine Versions

Server Type

The server types supported by Event Streams are as follows.

For detailed information about the server types provided by Event Streams, refer to Event Streams Server Types.

Standard ess1v2m4
CategoryexampleDetailed description
Server typeStandardProvided server types
  • Standard: Standard configuration (vCPU, Memory) commonly used
  • High Capacity: Large-capacity server specifications of 24 vCores or more
Server specificationsess1Provided server specifications
  • ess1, ess2: Standard specifications (vCPU, Memory) commonly used
  • esh2: Large-capacity server specifications
    • Providing servers with 24 vCores or more
Server specificationsv2Number of vCores
  • v2: 2 virtual cores
Server specificationsm4Memory capacity
  • m4: 4GB Memory
Table. Event Streams Server Type Components

Preliminary Service

This is a list of services that must be pre-configured before creating the service. Please refer to the guide provided for each service and prepare in advance.

Service CategoryserviceDetailed description
NetworkingVPCA service that provides an isolated virtual network in a cloud environment
Table. Event Streams Preliminary Services

2.1.1 - Server type

Event Streams server type

Event Streams provides server types composed of various combinations such as CPU, Memory, and Network Bandwidth. When creating Event Streams, Apache Kafka is installed according to the server type selected for the intended use.

Reference
The server types offered may vary depending on the region and AZ.

The server types supported by Event Streams are as follows.

Standard ess1v2m4
Category
ExampleDetailed description
Server typeStandardProvided server type categories
  • Standard: Configured with the commonly used standard specifications (vCPU, Memory)
  • High Capacity: Large‑capacity server specifications beyond Standard
Server specificationsess1Provided server type classification and generation
  • ess1: s means standard specification, and 1 indicates the generation
  • esh2: h means high-capacity server specification, and 2 indicates the generation
서버 사양v2Number of vCores
  • v2: 2 virtual cores
Server specificationsm4Memory capacity
  • m4: 4GB Memory
Table. Event Streams server type format
Reference

Check the node’s minimum specifications as shown below and select the server type.

CategoryvCPUMemory
Broker2 vCore4 GB
Zookeeper1 vCore2 GB

ess1 server type

The ess1 server type of Event Streams is offered with standard specifications (vCPU, Memory) and is suitable for various database workloads.

  • Intel 3rd‑generation (Ice Lake) Xeon Gold 6342 Processor up to 3.3 GHz
  • Supports up to 16 vCPUs and 64 GB of memory
  • Maximum networking speed of 12.5 Gbps
CategoryServer typevCPUMemoryNetwork Bandwidth
Standardess1v1m21 vCore2 GBUp to 10 Gbps
Standardess1v2m42 vCore4 GBUp to 10 Gbps
Standardess1v2m82 vCore8 GBUp to 10 Gbps
Standardess1v4m84 vCore8 GBUp to 10 Gbps
Standardess1v4m164 vCore16 GBUp to 10 Gbps
Standardess1v8m168 vCore16 GBMaximum 10 Gbps
Standardess1v8m328 vCore32 GBUp to 10 Gbps
Standardess1v16m3216 vCore32 GBUp to 12.5 Gbps
Standardess1v16m6416 vCore64 GBUp to 12.5 Gbps
Table. Event Streams server type specifications - ess1 server type

ess2 server type

The ess2 server type of Event Streams is offered with standard specifications (vCPU, Memory) and is suitable for various database workloads.

  • Intel 4th‑generation (Sapphire Rapids) Xeon Gold 6448H Processor up to 3.2 GHz
  • Supports up to 16 vCPUs and 64 GB of memory
  • Maximum networking speed of 12.5 Gbps
CategoryServer typeCPU vCoreMemoryNetwork Bandwidth(Gbps)
Standardess2v1m21 vCore2 GBUp to 10 Gbps
Standardess2v2m42 vCore4 GBUp to 10 Gbps
Standardess2v2m82 vCore8 GBUp to 10 Gbps
Standardess2v4m84 vCore8 GBUp to 10 Gbps
Standardess2v4m164 vCore16 GBUp to 10 Gbps
Standardess2v8m168 vCore16 GBUp to 10 Gbps
Standardess2v8m328 vCore32 GBUp to 10 Gbps
Standardess2v16m3216 vCore32 GBUp to 12.5 Gbps
Standardess2v16m6416 vCore64 GBUp to 12.5 Gbps
Table. Event Streams server type specifications - ess2 server type

esh2 server type

The esh2 server type of Event Streams is provided with high-capacity server specifications and is suitable for database workloads for large-scale data processing.

  • Intel 4th‑generation (Sapphire Rapids) Xeon Gold 6448H Processor up to 3.2 GHz
  • Supports up to 32 vCPUs and 128 GB of memory
  • Maximum networking speed of 25 Gbps
CategoryServer typevCPU: 2 (2 vCPU)MemoryNetwork Bandwidth
High Capacityesh2v32m6432 vCore64 GBMaximum 25 Gbps
High Capacityesh2v32m12832 vCore128 GBMaximum 25 Gbps
Table. Event Streams server type specifications - esh2 server type

ess3 server type

The ess3 server type of Event Streams is offered with standard specifications (vCPU, Memory) and is suitable for various database workloads.

  • Intel 6th‑generation (Granite Rapids) Xeon 6737P Processor up to 4.0 GHz
  • Supports up to 16 vCPUs and 256 GB of memory
  • Maximum networking speed of 12.5 Gbps
CategoryServer typeCPU vCoreMemoryNetwork Bandwidth(Gbps)
Standardess3v1m21 vCore2 GBMaximum 10 Gbps
Standardess3v2m42 vCore4 GBUp to 10 Gbps
Standardess3v2m82 vCore8 GBUp to 10 Gbps
Standardess3v4m84 vCore8 GBUp to 10 Gbps
Standardess3v4m164 vCore16 GBMaximum 10 Gbps
Standardess3v8m168 vCore16 GBUp to 10 Gbps
Standardess3v8m328 vCore32 GBUp to 10 Gbps
Standardess3v16m3216 vCore32 GBUp to 12.5 Gbps
Standardess3v16m6416 vCore64 GBUp to 12.5 Gbps
Table. Event Streams server type specifications - ess3 server type

esh3 server type

The esh3 server type of Event Streams is provided with high-capacity server specifications and is suitable for database workloads for large-scale data processing.

  • Intel 6th‑generation (Granite Rapids) Xeon 6738P Processor up to 4.1 GHz
  • Supports up to 128 vCPUs and 1,536 GB of memory
  • Maximum networking speed of 25 Gbps
CategoryServer typevCPUMemoryNetwork Bandwidth
High Capacityesh3v32m6432 vCore64 GBMaximum 25 Gbps
High Capacityesh3v32m12832 vCore128 GBMaximum 25 Gbps
Table. Event Streams server type specifications - esh3 server type