VMWare VCF Automated deployment
MetalSoft provides an application extension for deploying VMware Cloud Foundation (VCF) automatically. Supported versions are VCF 9.1.0.0 (default) and VCF 9.0.1, selected per site through the vcf_version config variable, via the vmware-cloud-foundation9 extension. The extension deploys the management domain and, optionally, a VI workload domain on a second set of hosts (see Workload domains). The extension’s README documents every input and the full deployment flow — this page covers the MetalSoft-side setup.
The older extension for VCF 5.2.1 is still available at vmware-cloud-foundation.
What changed vs. VCF 5.2.x
Section titled “What changed vs. VCF 5.2.x”- Cloud Builder is gone — it is replaced by the VCF Installer appliance (the same OVA as SDDC Manager). On 9.0.1, per the Broadcom BOM, the VCF Installer 9.0.2.0 OVA is used to deploy the 9.0.1 components — this is intentional, not a version mismatch.
- The installer embeds no product binaries — everything (vCenter, NSX, SDDC Manager, VCF Operations, …) is pulled from an offline depot mirror over HTTPS with basic auth.
- No license keys at bring-up — the deployment runs in 90-day evaluation mode; licensing is applied post-deploy in VCF Operations.
- Larger management footprint — a minimum of 7 management VMs are deployed (vCenter, 3× NSX Manager + VIP, SDDC Manager, VCF Operations, Operations collector, plus Operations fleet management on 9.0.1); VCF Automation is optional.
What changed in VCF 9.1
Section titled “What changed in VCF 9.1”- The standalone VCF Operations fleet management appliance is gone — it is replaced by VCF Management Services, a License Server and an Identity Broker, all mandatory for a new fleet. They need 5 extra FQDNs (
sr1,ic1,fc1,lic1,idb1) and a 12-address IP pool (vcfms-ip-pool) onvcf-mgmt. - VCF Automation needs a 5-address pool (was 2) and a separate platform FQDN (
auto1p). - ESX hosts must run ESX 9.1.0.0-25370933 (OS template
esxi-9-25370933-cluster-node). - The depot mirror must be populated with 9.1 content by VCF Download Tool 9.1, which authenticates with a Software Depot ID + Activation Code from the Broadcom VCF Business Services console (download tokens no longer work). The installer’s own offline-depot configuration (basic auth against the mirror) is unchanged.
At a site that must stay on 9.0.1, set the config variables vcf_version=9.0.1.0 and leave ova_name empty (it is derived from the version). The 9.1-only DNS records and IP pools are still allocated by the platform but are not used.
Requirements
Section titled “Requirements”- Hosts: minimum 3 (4 recommended and default) vSAN-capable hosts for the management domain, each with at least two NICs at 10 Gbps or faster. Identical hardware (vendor, model, CPU, memory, disks, NICs) is strongly recommended. Check the VMware compatibility guides for supported hardware.
- ESXi template: the hosts must be pre-installed with the exact ESX build of the site’s
vcf_version— ESX 9.1.0.0-25370933 via OS template esxi-9.1-25370933-cluster-node for 9.1.0.0; a separate template, esxi-9.0-24957456-cluster-node (ESX 9.0.1.0-24957456), exists for 9.0.1 (see below). - Offline depot: an offline depot mirror reachable from the installer, serving
/PROD(the product version catalog and component binaries, including the ESX ISO of the selected version) over HTTPS with basic auth. =======
- Cloud Builder is gone — the VCF Installer appliance (the same OVA as SDDC Manager) replaces it. Per the Broadcom 9.0.1 BOM, the VCF Installer 9.0.2.0 OVA deploys the 9.0.1 components — this is intentional, not a version mismatch.
- The installer embeds no product binaries — it pulls everything (vCenter, NSX, SDDC Manager, VCF Operations, …) from an offline depot mirror over HTTPS with basic auth.
- No license keys at bring-up — the deployment runs in 90-day evaluation mode; apply licensing post-deploy in VCF Operations.
- Larger management footprint — the extension deploys a minimum of 7 management VMs (vCenter, 3× NSX Manager + VIP, SDDC Manager, VCF Operations, Operations fleet management, Operations collector); VCF Automation is optional.
Requirements
Section titled “Requirements”-
Hosts: minimum 3 (4 recommended and default) vSAN-capable hosts for the management domain, each with at least two NICs at 10 Gbps or faster. MetalSoft strongly recommends identical hardware (vendor, model, CPU, memory, disks, NICs). Check the VMware compatibility guides for supported hardware.
-
ESXi template: the hosts must be pre-installed with exactly ESX 9.0.1.0-24957456 — OS template esxi-9.0-24957456-cluster-node (see below).
-
Offline depot: an offline depot mirror reachable from the installer, serving
/PROD(the product version catalog and component binaries, including the ESX 9.0.1 ISO) over HTTPS with basic auth. -
DNS: a routable DNS zone for all management FQDNs (do not use a
.localdomain), and a DNS provisioning workflow extension (e.g. PowerDNS) installed and enabled at the site so that the records the platform allocates actually resolve. VCF validation requires the full forward and reverse (PTR) round-trip to succeed. -
Ansible Runner: enabled on the site controller — see Enabling the Ansible Runner Capability.
Registering the ESXi 9 template
Section titled “Registering the ESXi 9 template”There is one template per supported VCF version, each installing exactly the ESX build of its BOM:
| VCF version | Template (os-templates repo path) | ESX build |
|---|---|---|
| 9.1.0.0 (default) | ESXi/9.1/esxi-9.1-25370933-cluster-node | 9.1.0.0-25370933 |
| 9.0.1 | ESXi/9.0/esxi-9.0-24957456-cluster-node | 9.0.1.0-24957456 |
Register the one matching the site’s vcf_version (or both, if sites on different versions share the Global Controller). The commands below use the 9.1 template — for 9.0.1 substitute ESXi/9.0/esxi-9.0-24957456-cluster-node.
1a. Create the MetalSoft ESXi template (with internet access):
metalcloud-cli os-template list-repometalcloud-cli os-template create-from-repo ESXi/9.1/esxi-9.1-25370933-cluster-node1b. Create the MetalSoft ESXI template (airgapped) In an air-gapped environment, on a machine with internet access clone the github repository
git clone https://github.com/metalsoft-io/os-templatesCopy the os-templates directory to the machine where the CLI can be executed that has access to the Global Controller. The os-templates directory should contain the .git files.
metalcloud-cli os-template list-repo --repo-url /path/to/os-templatesmetalcloud-cli os-template create-from-repo ESXi/9.1/esxi-9.1-25370933-cluster-node --repo-url /path/to/os-templates- Update the template’s ISO URL asset to point at the matching ESX installer ISO hosted on your repository (9.1.0.0-25370933 for the 9.1 template, 9.0.1.0-24957456 for the 9.0 one) — the extension’s preparation role verifies the ESXi version and that the host certificate CN matches the host FQDN.
Registering the VCF 9 extension
Section titled “Registering the VCF 9 extension”- Clone the repository:
git clone https://github.com/metalsoft-io/metalsoft-extensionscd metalsoft-extensions/vmware-cloud-foundation9- Build and upload the Ansible bundle (playbooks at the archive root — see Structuring Ansible bundles):
(cd ansible && zip -r ../vmware-cloud-foundation9-v2.4.3.zip . -x '*.DS_Store' -x '*.zip' -x 'README.md')Upload the zip to a repository reachable by the Global Controller and set its URL as the vcf-ansible-bundle asset’s url in extension.json.
- Build and publish the execution environment image — the container image in which the Site Controller runs the extension’s Ansible tasks; every dependency the playbooks use must be baked into it. Build it from the bundled
execution-environment.yml(it pulls in the VMware SDKs, Galaxy collections and CLI tools the playbooks need):
ansible-builder build --file execution-environment.yml \ --tag <registry>/<namespace>/vcf9-ee:9.0.1 \ --container-runtime docker \ --extra-build-cli-args "--platform=linux/amd64"The image must be linux/amd64 (site controllers run x86_64 and the EE compiles native wheels). Push it to a registry reachable by the Site Controller, and update the ee-vcf9-9-0-1 asset’s registry fields (and, for air-gapped sites, host a docker save | gzip tarball at the asset’s url).
-
Adapt
extension.jsonto your environment before registering — at minimum:- change every
*_passwordinput default — the committed values are placeholders; new values must satisfy the inputs’validationRegEx(^[A-Za-z0-9!@#$%^&*+]{12,20}$— the charset all VCF components accept); - the
infrastructure.logicalNetworks[].profileLabelvalues must match logical network profiles that exist at the target site (see below); - both asset URLs from steps 2 and 3;
- optionally, adjust the
configVarsdefaults (depot location, appliance sizing, etc. — see step 6) so operators don’t have to override them at every site.
- change every
-
Register, publish and (optionally) make the extension public:
metalcloud-cli extension create "VMware Cloud Foundation 9 (VCF)" application "VMware Cloud Foundation (VCF)" --definition-source extension.jsonmetalcloud-cli extension publish <id-of-created-extension>metalcloud-cli extension make-public <id-of-created-extension>-
Activate the extension on the site level and configure its configVars. After publishing, the extension must also be activated per site, on the site level configuration page, and its site configuration saved before the first deploy. All the admin/site-level settings live in configVars and are set once per site (not per deployment); config vars without a default — currently
depot_password— are mandatory at that point. The most important ones:DNSResolvers— set it to your site’s internal resolvers (comma-separated). MetalSoft configures these resolvers on all hosts and appliances, and they must be able to resolve the management DNS records; handing VCF a public resolver breaks bring-up.depot_hostname/depot_port/depot_username/depot_password— point at your offline depot mirror.vcf_version—9.1.0.0(default) or9.0.1.0;ova_name/ova_urlare derived from it and can be left empty (on 9.0.1 the 9.0.2.0 installer OVA is intentional).nsx_manager_size/vcenter_vm_size/vcf_ops_appliance_size— appliance sizing (NSXsmallno longer exists in 9.x).deploy_vcf_automation/vcf_automation_internal_cluster_cidr— enable VCF Automation and adjust its internal CIDR if it would overlap existing networks.deploy_without_license_keys— must staytrueon VCF 9 (bring-up runs in evaluation mode).wld_nsx_manager_count— number of NSX Manager nodes deployed for the VI workload domain (1–3, default 3).
See the extension README for the full config var table.
At instance-create time the operator only fills in the deployment-specific inputs: server type, node count and OS template for the management domain, the appliance passwords and, optionally, vcf_instance_name (when unset the playbooks derive <extension_instance_id>-vcf). The form also asks for the workload domain inputs — wld_domain_instance_count (default 0 = no workload domain), the workload server type / OS template and the wld_* appliance passwords. The server type and OS template must be selected even when the count is 0, because these input types cannot carry defaults.
Preparing logical networks
Section titled “Preparing logical networks”The extension declares 4 logical networks; a logical network profile with the matching label must exist at every site where VCF will be deployed:
Go to Admin > Fabrics > fabric > Logical networks and create the following (the labels need to match the extension definition’s profileLabel values):
vcf-mgmt— tagged: true, Active-Active redundancy. Provides the default route and carries all the management appliance IP allocations and DNS records. Ensure this network has access to the Site Controller, in-band (set the correct VLAN id and/or IP ranges).vcf-vsan— tagged: true, Active-Active redundancy. Carries an IP pool handed to VCF for vSAN vmkernel addressing.vcf-vmotion— tagged: true, Active-Active redundancy. Carries an IP pool for vMotion vmkernel addressing.vcf-nsx— tagged: true, Active-Active redundancy. Carries an IP pool for NSX TEP addressing.
Set the IPv4 allocation strategy to Auto with sufficiently large subnets — the extension requests two sets of IP pools, one for the management domain (8 addresses for vSAN, 8 for vMotion, 20 for NSX) and one for the workload domain (16 for vSAN, 16 for vMotion, 20 for NSX), plus the management allocations (installer, SDDC Manager, vCenter, 3× NSX Manager + VIP, the VCF Operations appliances, on 9.1 the VCF Management Services / License Server / Identity Broker FQDNs and 12-address pool, and, optionally, VCF Automation) and the workload domain appliances (vCenter, 3× NSX Manager + VIP) on vcf-mgmt. All allocations, including the workload domain ones, are reserved at deploy time even when wld_domain_instance_count is 0.
Scaling the management domain
Section titled “Scaling the management domain”Scale the management domain by editing the mgmt_domain_instance_count input on the deployed instance (default 4, min 3, max 16). On save, the platform adjusts the servers and runs the onEdit lifecycle: it removes and decommissions outgoing hosts from the cluster in SDDC Manager at preDeploy (while still reachable), and prepares, commissions, and adds new hosts to the cluster at postDeploy. Scaling requires a successfully completed initial deployment, and the new hosts must meet the same hardware/ESX build requirements as the initial ones.
Workload domains
Section titled “Workload domains”Besides the management domain, the extension can create a VI workload domain — its own vCenter, its own NSX Manager cluster and a vSAN cluster on a separate set of hosts — through the SDDC Manager API. The workload domain is driven by a second instance array (workload) and controlled entirely by the wld_domain_instance_count input:
0(default) — management domain only. Nothing is created in SDDC Manager and thewld_*outputs are empty strings.3–16— a VI workload domain is created after the management bring-up (~2–3 hours extra). Values1and2are rejected at the input, since vSAN requires at least 3 hosts.
The workload domain uses an isolated SSO domain (required by VCF 9 for VI domains) — its vCenter login is administrator@wld-<domain>.sso.local with the wld_vcenter_sso_admin_password input as password. Naming mirrors the management domain with a w prefix: domain <extension_instance_id>-w, cluster <extension_instance_id>-w-cl1, appliances w-vcs1 / w-nsx1(a|b|c) in the site’s default DNS zone.
Requirements, in addition to the management domain ones:
- A completed management domain bring-up — the workload flow talks to the live SDDC Manager.
- Workload hosts on the same ESX build as the management hosts, vSAN-capable, meeting the same hardware requirements (they can be a different server type than the management hosts).
- A vLCM cluster image available on SDDC Manager — the management domain image is reused automatically.
- The
w-*DNS records resolving forward and reverse (same DNS provisioning extension as for the management records), and free addresses in the workload IP pools.
Lifecycle — all through editing wld_domain_instance_count on the deployed instance (the onEdit lifecycle, same as scaling the management domain):
| Edit | What happens |
|---|---|
0 → N (N ≥ 3) | Creates the workload domain: host preparation, network pool creation, host commissioning, domain spec validation and creation. Can also be done at initial deploy by setting the count at create time. |
N → N+k | New hosts are prepared, commissioned and added to the workload cluster. |
N → N-k (result ≥ 3) | Outgoing hosts are removed from the workload cluster and decommissioned, one at a time. |
N → 0 | Deletes the entire workload domain — its vSAN datastore and everything on it. |
Warning:
N → 0is irreversible. Migrate or remove user VMs and delete any NSX Edge clusters in the domain first (SDDC Manager refuses the deletion otherwise); the extension does not check for user workloads before deleting.
A failed workload lifecycle task can be retried from the platform without manual cleanup — the playbooks probe the SDDC Manager state first (existing domain, already commissioned hosts, existing network pool) and resume where they left off.
Deploy a test VCF cluster
Section titled “Deploy a test VCF cluster”From the Infrastructure Designer, add a new VCF cluster from the list, fill in the form, and deploy. The initial bring-up takes several hours (depot binary downloads plus ~3 hours for the management domain bring-up on physical hosts).
Outputs
Section titled “Outputs”After bring-up, the extension instance exposes the following outputs with the entry-point URLs and identifiers of the deployed stack. The URLs are derived from the management FQDNs allocated on vcf-mgmt and are refreshed on every successful run (including scaling). No credentials are exported — the appliance passwords remain inputs:
installer_url— the VCF Installer UI.sddc_manager_url— SDDC Manager.vcenter_url— the management vCenter.nsx_manager_url— the NSX Manager VIP.vcf_ops_url— VCF Operations.ops_fleet_mgmt_url— VCF Operations fleet management (9.0.1 only; empty on 9.1).ops_collector_url— VCF Operations collector.vcf_automation_url— VCF Automation (empty unlessdeploy_vcf_automationis enabled).sddc_id— the SDDC identifier used for naming (<extension_instance_id>-mby default).vcenter_sso_username— the management vCenter login (administrator@vsphere.localunless the SSO domain is overridden).datacenter_name/cluster_name— the vCenter datacenter and cluster names (default<sddc_id>-dc1and<sddc_id>-cl1).
Workload domain outputs (empty strings while wld_domain_instance_count is 0):
wld_vcenter_url— the workload domain vCenter.wld_nsx_manager_url— the workload domain NSX Manager VIP.wld_vcenter_sso_username— the workload vCenter login in its isolated SSO domain (administrator@wld-<domain>.sso.local).wld_domain_name— the workload domain name in SDDC Manager (<extension_instance_id>-w).