Subcloud Enrollment of a Factory-Cloned System

This feature enables mass deployment of subclouds by cloning a fully configured factory-installed subcloud and enrolling the cloned systems into a Distributed Cloud system. Instead of installing and configuring each subcloud individually, two approaches are supported:

  • Option A: Install and Restore from a Factory Clone — Install a minimal StarlingX system on the target, then restore from a backup taken on the factory-installed subcloud.

  • Option B: Disk Clone — Create a bit-for-bit disk image of the golden subcloud and write it to the target hardware.

Prerequisites

  • The subclouds must be AIO-SX subclouds.

  • All target hardware must be identical to the golden subcloud hardware (same firmware, NIC layout, disk topology, and CPU architecture).

  • Network connectivity (management and OAM) between the System Controller and the target subcloud locations must be established before enrollment.

  • The subcloud system install or restore from a clone must meet the hardware and network requirements. Because all subclouds are installed with the same image, their network configuration will be identical after installation. Their platform networks (OAM, management, admin, and so on) must be L2-isolated to avoid IP conflicts.

  • A well-known CA certificate must be installed on the subclouds before the factory install. It must also be installed on the System Controller before enrollment, which is the same requirement as for normal enrollment.

  • BMC credentials — Each target’s BMC credentials must be configured independently and must not be included in the clone.

  • Applications that reference hardware, OS, or machine UUIDs may need to be updated or reinstalled after enrollment. UUIDs are regenerated during enrollment to differentiate systems.

  • Prepare the factory-installed system (source system) before creating the clone image:

    ~(keystone_admin)]$ sudo pre-factory-install-clone.sh
    

    Note

    The system shuts down gracefully after preparation. Ensure that all services and applications can be recovered after cloning.

Option A — Install and Restore from a Factory Clone

Use this option to install a minimal StarlingX system on the target hardware and restore it from a backup taken on the factory-installed subcloud.

Procedure

  1. Update /opt/platform-backup/factory/<software_version>/backup_restore_values.yaml on the factory-installed system (source system) and add the following:

    ansible_become_pass: <password>
    ansible_ssh_pass: <password>
    backup_filename: factory_backup.tgz
    restore_registry_filesystem: true
    registry_backup_filename: localhost_image_registry_backup_*.tgz # put the actual file name
    

    Note

    The password must match the sysadmin password used during the factory install.

  2. Compress and transfer the factory backup.

    ~(keystone_admin)]$ cd /opt/platform-backup/factory/
    ~(keystone_admin)]$ tar czf <software_version>.tgz <software_version>/
    
  3. Transfer <software_version>.tgz to the target server after installation is complete, then extract it.

    ~(keystone_admin)]$ mkdir -p /opt/platform-backup/factory/
    ~(keystone_admin)]$ sudo tar xzf <software_version>.tgz -C /opt/platform-backup/factory/
    
  4. Restore the target system. For the full restore procedure, see Restore Without Reinstall (Factory-Installed).

    1. (Optional) Update the sysadmin password if it differs from ansible_become_pass:

      ~(keystone_admin)]$ echo "sysadmin:<password>" | chpasswd
      
    2. Run the restore playbook:

      ~(keystone_admin)]$ sudo -u sysadmin ansible-playbook \
          /usr/share/ansible/stx-ansible/playbooks/restore_platform.yml \
          -e "@/opt/platform-backup/factory/<software_version>/backup_restore_values.yaml"
      
  5. Unlock controller-0 and wait for it to come back online with status available and the task cleared.

  6. Complete the restore.

    ~(keystone_admin)]$ system restore-complete
    

    Following is the expected restoration:

    Backup tarball: around 10 GB, restoration playbook execution time: 6 to 8 minutes, total execution time: around 30 minutes calculating the install and unlock time.

Option B — Disk Clone

Use this option to create a bit-for-bit disk image of the golden subcloud and write it directly to the target hardware.

Procedure

  1. Create an image from the source server. The following steps use Clonezilla Live as an example; other tools may also be used.

    1. Download Clonezilla Live from https://clonezilla.org/clonezilla-live.php.

    2. Insert the Clonezilla Live image into the source server and power it on.

      Note

      Assign an IPv4 address to the server manually or via DHCP.

    3. Select device-image → Work with disks or partitions using images.

    4. Select the image repository (NFS server): nfs_server → nfs v4 → Use NFS server as image repository. Clonezilla will mount the remote directory via NFS.

    5. Select cloning mode: beginner and action: savedisk → Save local disk as image.

    6. Set an image name or accept the default timestamp.

    7. Select all disks on the server to save.

    8. Select the following clone options:

      Compression

      Default (use parallel gzip/zstd)

      FSCK

      Choose -fsck-y and it will check the saved image is restorable

      Encryption

      -senc: Do not encrypt the image

      Copy logs

      No, skip copying logs to USB

      Action when finished

      -p choose: Enter command prompt

    9. Monitor progress until completion. Several steps require entering yes to continue.

    10. Verify the image from the remote SSH image repository. Expected contents include:

      • Partition images (compressed)

      • Logical volume images (cgts-vg-*, docker-*, etcd-*)

      • Disk metadata files

      • blkdev.list, dev-fs.list, parts, and so on

    11. Power off the machine and eject the Clonezilla Live image.

  2. Restore the image to a target server.

    1. Download Clonezilla Live from https://clonezilla.org/clonezilla-live.php.

    2. Insert the Clonezilla Live image into the source server and power it on.

    3. Select device-image → nfs_server → Expert → restoredisk.

    4. Select the target disks to restore.

    5. Select the following restore options:

      FSCK

      Yes, check the image before restoring

      Action when finished

      -p choose: Enter command prompt

    6. Monitor the restore until completion.

    7. Verify that all partitions are restored to the disks, then power off the machine and eject the Clonezilla Live image. Following is the expected restoration:

      image size: around 50 GB, time cost: 10 - 12 minutes

Subcloud Enrollment of Cloned System

For subcloud enrollment, follow the normal subcloud enrollment process. Unique data for platform services is updated during enrollment, ensuring those services use correct data after enrollment.

Note

K8s etcd root CA certificate cannot be updated from the cloned source post enrollment. K8s root CA is not updated during the enrollment, however it can be updated with the existing orchestration. See Kubernetes Root CA Certificate Update for Distributed Cloud Orchestration.

For the enrollment procedure, see Enroll a Factory Installed Non Distributed Standalone System as a Subcloud.