本頁面說明如何將超磁碟附加至 Managed Service for Apache Spark 叢集中的虛擬機器 (VM)。您可以為主要、主要 worker 和次要 worker 節點群組,分別設定磁碟。除了開機磁碟和附加至叢集節點的任何本機 SSD 外,這些磁碟也會附加至叢集節點。
事前準備
-
In the Cloud de Confiance console, on the project selector page, select or create a Cloud de Confiance project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that you have the permissions required to complete this guide.
-
Verify that billing is enabled for your Cloud de Confiance project.
Enable the Managed Service for Apache Spark API, if it is not already enabled.
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.
必要的角色
如要執行本頁的範例,您必須具備特定 IAM 角色。視機構政策而定,系統可能已授予這些角色。如要查看角色授予情況,請參閱「是否需要授予角色?」一文。
如要進一步瞭解如何授予角色,請參閱「管理專案、資料夾和機構的存取權」。
使用者角色
如要取得建立 Managed Service for Apache Spark 叢集所需的權限,請要求管理員授予您下列 IAM 角色:
- 專案的 Dataproc 編輯者 (
roles/dataproc.editor) - Compute Engine 預設服務帳戶的服務帳戶使用者 (
roles/iam.serviceAccountUser)
服務帳戶角色
為確保 Compute Engine 預設服務帳戶具備建立 Managed Service for Apache Spark 叢集的必要權限,請要求管理員在專案中,將 Dataproc 工作者 (roles/dataproc.worker) IAM 角色授予 Compute Engine 預設服務帳戶。
磁碟特性
附加的磁碟具有下列特性:
- 生命週期:附加磁碟的生命週期與附加的 VM 相同。Managed Service for Apache Spark 會在建立 VM 時建立磁碟,並在刪除 VM 時刪除磁碟。
- 不可變更性:叢集建立後,您就無法更新連接磁碟的屬性,例如大小、IOPS 或處理量。
- 掛接和使用:Managed Service for Apache Spark 會在
/mnt/N掛接磁碟,其中N是正整數 (例如/mnt/1、/mnt/2)。HDFS 和暫存資料 (例如 Shuffle 輸出) 會使用附加的磁碟,而不是開機永久磁碟。
磁碟設定
將磁碟附加至 Managed Service for Apache Spark 叢集節點時,可以指定下列磁碟設定參數:
磁碟類型 - 必填:要連接至 VM 執行個體的磁碟類型。 系統支援下列超磁碟:
hyperdisk-balancedhyperdisk-extremehyperdisk-mlhyperdisk-throughput
hyperdisk balanced high availability類型和永久磁碟無法附加至叢集節點。大小 (選用):磁碟大小。這個值必須是整數,後面加上
GB代表 GB,TB代表 TB。舉例來說,10GB會附加 10 GB 的磁碟。詳情請參閱「Hyperdisk 大小限制」。IOPS - 選用:指出要為連結磁碟佈建的 IOPS。這個參數會設定每秒磁碟 I/O 作業的上限。詳情請參閱「預設效能等級」。
總處理量 - 選填:指出要為連結磁碟佈建的總處理量。這個參數會設定每秒
MiB的處理量上限。詳情請參閱「預設效能等級」。
將磁碟連接至叢集
使用 gcloud CLI 或 Dataproc API 建立 Managed Service for Apache Spark 叢集時,可以附加磁碟並指定磁碟設定。
gcloud CLI
如要在建立叢集時附加磁碟,請搭配
gcloud dataproc clusters create指令使用--master-attached-disks、--worker-attached-disks或--secondary-worker-attached-disks標記。每個標記都接受以分號分隔的磁碟設定清單。 每個磁碟設定都是以半形逗號分隔的鍵/值組合清單,適用於
type、size、iops和throughput(請參閱「磁碟設定」)。
範例:下列指令會建立叢集,並將兩個 Hyperdisk 附加至每個主要 worker 節點。
gcloud dataproc clusters create CLUSTER_NAME \
--region=REGION \
--worker-attached-disks='type=hyperdisk-balanced,size=100GB,iops=5000,throughput=200;type=hyperdisk-throughput,size=9000GB'
API
如要附加磁碟,請在
masterConfig、workerConfig或secondaryWorkerConfig執行個體群組的diskConfig物件中加入attachedDiskConfigs陣列。在
clusters.createAPI 要求的內文中提供設定。
範例:下列 JSON 程式碼片段顯示 attachedDiskConfigs 陣列,其中附加了兩個超磁碟。
[
{
"type": "hyperdisk-balanced",
"diskSizeGb": 100,
"provisionedIops": 5000,
"provisionedThroughput": 200
},
{
"type": "hyperdisk-throughput",
"diskSizeGb": 9000
}
]