Download PDF

VMware Cloud Foundation 9 Upgrade and Application Modernization

Julius M. Nicolescu

September 2026

Technical Architecture and Implementation Roadmap

Client Strategic enterprise client — three datacenters (West US, East US, West Europe)
Prepared by Julius M. Nicolescu
Document date 2026-09-09
Scope Architecture and roadmap for upgrading three standalone vSphere 7/8 environments to VMware Cloud Foundation (VCF) 9.1: importing each existing environment into VCF as an upgraded workload domain, and building new, dedicated workload domains to bring VMware vSphere Kubernetes Service (VKS) 3.7 into production, with minimal workload disruption, simplified routing and load-balancer integration, and standardized automated operations.

Chapter Index


Introduction — The Problem We Are Trying to Solve

This document defines the VCF 9 upgrade architecture for a strategic enterprise client, led by the client’s VCF 9 Upgrade Architect. The engagement has two parts: upgrading the client’s three existing, standalone vSphere 7/8 environments by importing each into VMware Cloud Foundation (VCF) 9.1 as a managed workload domain, and building new, dedicated workload domains to bring VMware vSphere Kubernetes Service (VKS) online and scale and standardize modern container workloads across the client’s infrastructure footprint. It establishes the target-state design, the constraints the design must satisfy, and the phased upgrade path required to reach production readiness.

Current State and Assumptions

Objective

This document presents the VCF 9 upgrade path for the three standalone vSphere environments — each imported into VCF 9.1 as an upgraded workload domain — together with the new, VKS-dedicated workload domain design needed to bring VKS 3.7 into production. The design objective is the fastest, lowest-risk path to production readiness across all three physical locations.

Design Constraints

Constraint 1 — Minimal Workload Disruption. The upgrade methodology must maintain continuous infrastructure availability and policy alignment for existing production workloads running on the vSphere 8 environments. The design defines two tracks that run side by side — a brownfield upgrade of the existing environment and a greenfield build for VKS — that together eliminate risk to live traffic during the VCF 9.1 deployment.

Constraint 2 — Simplified Routing and Load Balancer Integration. The design must account for the pre-existing VMware NSX fabric and F5 BIG-IP load balancers as they integrate into the new VCF 9.1 and VKS 3.7 architecture, including network simplification options and alternative design choices (for example, vDS networking with Avi Load Balancer as a side-by-side or greenfield replacement). The design must provide resilient container-native ingress/egress routing, cloud-native microservices routing, and multi-site traffic management.

Constraint 3 — Standardized Automated Operations. The target architecture must adhere strictly to upstream Kubernetes standards and, at the later stages of the upgrade roadmap, fully leverage VCF SDDC Manager automated operations to manage the multi-cluster infrastructure lifecycle, ensuring repeatable deployment patterns for the rapid rollout of VKS 3.7.

Required Architectural Views

The design addresses three required architectural views, each mapped to a chapter in this document:

  1. Global Multi-Site SDDC Topology — a high-level physical-to-virtual architectural view showing how the three physical datacenters are consolidated and managed under a federated, multi-site model. Addressed in Chapter 4.
  2. Traffic Engineering and Packet Flow — the end-to-end routing path and policy enforcement points from the F5 BIG-IP edge, through the NSX software-defined overlay, into the VKS container-native ingress controllers, under a mandate for strict segmentation between namespaces. Addressed in Chapter 2, Chapter 6, Chapter 8, Chapter 9, and Chapter 10.
  3. The Upgrade Pipeline Sequence — a phased implementation timeline defining the fastest operational path from vSphere 8 to a fully operational VKS environment. Addressed in Chapter 3 and the brownfield/greenfield chapters that follow it.

A note on release versions: the detailed design chapters in this document are built against the VCF 9.1 / vSphere Supervisor design library, which is the current VCF 9.x release line that the VKS 3.7 Kubernetes distribution is consumed through. Wherever this document references “VCF 9.1,” that is the platform release this VKS 3.7 rollout runs on — there is no separate VCF version to plan for.

Problem Statement Summary

Stripped of surrounding detail, this is the problem the rest of the document solves:

Dimension Today Target
Platform Three standalone vSphere 7/8 environments, no shared management plane, ~45,000 cores total One VCF 9.1 Fleet, three VCF Instances (West US, East US, West Europe) under shared fleet-level operations
Support posture vSphere 7 exited General Support Oct 2, 2025; vSphere 8 runs out Oct 11, 2027 All hosts on a VCF 9.1 / ESXi 9.1-qualified, currently supported build
Kubernetes None — infrastructure is entirely VM-centric VKS 3.7 Kubernetes clusters running natively via vSphere Supervisor, in new workload domains built specifically for that purpose, alongside the existing VM workloads in their upgraded (imported) workload domains
Networking VMware NSX partially integrated across the US sites only; West Europe not yet on the fabric A consistent NSX (Segment Networking or VPC Networking) design applied uniformly across all three sites, with a documented Federation posture
Application delivery F5 BIG-IP DNS/GTM (global) and F5 BIG-IP LTM (local) exclusively, hardware-appliance-based, with no container-native ingress A resilient, container-native ingress/egress path, with a clear, staged option to consolidate onto VMware Avi Load Balancer fleet-wide, retiring F5 without a disruptive cutover
Operations Per-site, manual administration; no centralized visibility or lifecycle automation Centralized VCF Operations, License Server, and SDDC Manager-driven lifecycle automation, enabling repeatable, fleet-wide rollout patterns
Risk tolerance Zero tolerance for disruption to live production workloads during the transition Import (not migrate) each existing environment into an upgraded workload domain, leaving running VMs untouched, while a separate, new workload domain is built independently to bring VKS online

The chapters that follow document, in order: what the client operates today (Chapters 1–2), the overall roadmap and how its stages connect (Chapter 3), the multi-site fleet foundation every stage depends on (Chapter 4), the upgrade of the existing environment and the new, VKS-dedicated workload domain built alongside it — brownfield and greenfield — (Chapters 5–8), the Stage 3 load-balancer choice (Chapter 9), and the Stage 4 completion of the traffic-management modernization that retires F5 (Chapter 10).

Executive Summary — Recommended Path

This engagement has two workstreams, run side by side rather than as competing alternatives: upgrading the client’s existing vSphere 8 environment into VCF 9.1 by importing it as-is, and building a new, dedicated workload domain to bring VKS 3.7 into production. Because all three of the client’s sites already run vSphere with NSX, import — not migration or rebuild — is the correct mechanism for the upgrade workstream everywhere it applies.

  1. Stand up a new VCF 9.1 Fleet spanning all three sites, following the multi-site fleet design in Chapter 4, with US WEST hosting the fleet-wide services (VCF Operations, VCF Automation, License Server, SDDC Manager).
  2. Upgrade workstream — import, not migrate, each site’s existing vCenter and NSX environment into the fleet as a workload domain, per Chapter 5 — production VMs are never moved or rebuilt, only brought under VCF management. Where a site’s imported cluster has spare capacity, Supervisor can also be activated on it directly, using the Single Management Zone with Combined Workload Zones model, per Chapter 6.
  3. VKS workstream — build a new, dedicated workload domain using the Multi-Rack Layer 2 vSphere Cluster Model and the Centralized Connectivity Model, per Chapter 7, and activate the Three Management Zones Supervisor model on it, per Chapter 8 — giving VKS its own purpose-built capacity and zone-level resilience, independent of the imported estate.
  4. Choose a Stage 3 load-balancer for Supervisor on each workload domain — the built-in NSX Load Balancer, or Avi Load Balancer where Layer 7, WAF, GSLB, or analytics are required — per Chapter 9.
  5. Complete the private cloud operating model at Stage 4 by extending Avi fleet-wide and retiring F5 through the seven-phase, wave-based migration in Chapter 10.

The two workstreams do not block each other. The imported workload domain modernizes the existing vSphere 8 estate’s management plane without touching a single running VM, and can optionally take on some VKS capacity of its own if a site has room for it. The new, VKS-dedicated workload domain (Chapter 7 and Chapter 8) is built independently, on its own timeline, so VKS gets purpose-built capacity and zone-level resilience rather than being retrofitted onto infrastructure sized for something else.


Chapter 1: Existing State, Technical Documentation

Client: Strategic enterprise, 3 datacenters (West US, East US, West Europe)

Scope of this document: What the client has today, and what it takes for that hardware to qualify for VMware Cloud Foundation (VCF) 9.1. No future-state design is covered here.


1. Why vSphere support dates matter

The client’s three datacenters run a mix of vSphere 7 and vSphere 8. Both versions are still usable today, but they are on different support clocks — and vSphere 7 is close to running out.

Product family General Availability End of General Support End of Technical Guidance
vSphere 7.x / ESXi 7.x / vCenter 7.x Apr 2, 2020 Oct 2, 2025 Apr 2, 2027
vSphere 8.x / ESXi 8.x / vCenter 8.x Oct 11, 2022 Oct 11, 2027 Oct 11, 2029

General Support is the period where Broadcom actively patches, fixes bugs, and answers support tickets. Once this ends, the client is running unsupported software for production workloads. Technical Guidance is a reduced tier after General Support ends — existing documentation and best-effort guidance only, no new patches or bug fixes.

vSphere 7 already exited General Support on October 2, 2025. Any hosts still running vSphere 7 are outside the normal support window today. This is the most urgent driver for the upgrade. vSphere 8 has runway until October 2027. It is not an emergency, but it is also not permanent — it does not eliminate the need to move to VCF 9.1 / ESXi 9.1, it just means vSphere 8 hosts have more breathing room than vSphere 7 hosts.

The support clock, not just the desire to adopt Kubernetes, is what makes this an active modernization project rather than a “someday” initiative.


2. Sizing the existing environment: two conservative estimates

The client’s infrastructure footprint is stated as ~45,000 physical cores spread across three datacenters, on Dell hardware, with a 50/50 split between rack and blade/modular servers. Because the client did not supply exact server counts, two conservative, Intel-based estimates are used to translate “45,000 cores” into a concrete, physical picture: how many servers, what models, how many racks. Both estimates land on the same ~45,000 cores — they simply assume different-generation Dell hardware, which changes the server count and rack footprint.

2.1 Estimate A — 2021 Intel-generation hardware (higher server count)

This estimate assumes the client’s fleet was purchased around 2021, using 3rd Generation Intel Xeon Scalable CPUs — a conservative (lower cores-per-server) assumption, which produces a higher server and rack count.

Assumptions - 80 physical cores per server (2 × Intel Xeon 3rd Gen, up to 40 cores each) - 50/50 split between rack and blade/modular servers - Evenly distributed across the three datacenters

Hardware used in this estimate

Server type Dell model CPU generation Max cores/server
Rack server PowerEdge R750 2 × Intel Xeon 3rd Gen 80
Blade/modular sled PowerEdge MX750c 2 × Intel Xeon 3rd Gen 80

Totals

Type Count Cores/server Total cores
R750 rack servers 282 80 22,560
MX750c blade/modular servers 281 80 22,480
Total 563 servers 45,040 cores

Per-datacenter split

Datacenter Rack servers Blade servers Total servers Approx. cores
West US 94 94 188 15,040
East US 94 94 188 15,040
West Europe 94 93 187 14,960
Total 282 281 563 45,040

Rack layout per datacenter

2021 Estimate — Datacenter Rack Layout (11 Racks per Site)

Rack servers (shown for the largest site; other sites are the same pattern at 94 servers):

Rack R750 servers Used U
Rack 1 20 40U
Rack 2 20 40U
Rack 3 20 40U
Rack 4 20 40U
Rack 5 14 28U
Total 94 188U

Blade/modular servers (MX750c sleds, 8 sleds per MX7000 chassis, max 2 chassis per rack):

Rack MX7000 chassis Sled capacity Used U
Rack 6 2 16 sleds 14U
Rack 7 2 16 sleds 14U
Rack 8 2 16 sleds 14U
Rack 9 2 16 sleds 14U
Rack 10 2 16 sleds 14U
Rack 11 2 16 sleds 14U
Total 12 chassis 96 sled slots 84U

Rack count per datacenter (West Europe shown — 187 servers, one fewer blade sled than West US/East US):

Category Quantity Racks
R750 rack servers 94 servers 5 racks
MX750c blade servers 93 sleds / 12 MX7000 chassis 6 racks
Total West Europe 187 servers 11 racks

West US and East US follow the same layout with 94 rack servers and 94 blade sleds each, also landing at 11 racks per site.

If the client’s hardware is from the 2021 era, it takes 563 servers across roughly 33 racks total (11 per site) to deliver the ~45,000 cores they operate today.

2.2 Estimate B — 2023 Intel-generation hardware (lower server count)

This estimate assumes a refresh cycle to 4th Generation Intel Xeon Scalable CPUs — more cores per server, so the same ~45,000 cores are delivered by fewer, denser servers.

Assumptions - 112 physical cores per server (2 × Intel Xeon 4th Gen, up to 56 cores each) - 50/50 split between rack and blade/modular servers - Evenly distributed across the three datacenters (67 rack + 67 blade per site)

Hardware used in this estimate

Server type Dell model CPU generation Cores/server
Rack server PowerEdge R760 2 × Intel Xeon 4th Gen 112
Blade/modular server PowerEdge MX760c 2 × Intel Xeon 4th Gen 112

Totals

Type Count Cores/server Total cores
R760 rack servers 201 112 22,512
MX760c blade/modular servers 201 112 22,512
Total 402 servers 45,024 cores

Per-datacenter split

Datacenter Rack servers Blade servers Total servers Approx. cores
West US 67 67 134 15,008
East US 67 67 134 15,008
West Europe 67 67 134 15,008
Total 201 201 402 45,024

Rack layout per datacenter (identical for all three sites under this estimate — 67 rack + 67 blade servers each)

2023 Estimate — Datacenter Rack Layout (9 Racks per Site)

Rack servers:

Rack R760 servers Used U
Rack 1 20 40U
Rack 2 20 40U
Rack 3 20 40U
Rack 4 7 14U
Total 67 134U

Blade/modular servers (MX760c sleds, 8 sleds per MX7000 chassis, max 2 chassis per rack):

Rack MX7000 chassis MX760c sled capacity Used U
Rack 5 2 16 sleds 14U
Rack 6 2 16 sleds 14U
Rack 7 2 16 sleds 14U
Rack 8 2 16 sleds 14U
Rack 9 1 8 sleds 7U
Total 9 chassis 72 sled slots 63U

Rack count per datacenter:

Category Quantity Racks required
R760 rack servers 67 servers 4 racks
MX760c blade/modular servers 67 sleds / 9 MX7000 chassis 5 racks
Total per datacenter 134 servers 9 racks

This layout is the same at West US, East US, and West Europe.

If the client’s hardware is from the 2023 era, the same ~45,000 cores fit in 402 servers across 27 racks total (9 per site) — 161 fewer servers and 6 fewer racks total than Estimate A, purely because newer CPUs pack more cores per socket.

2.3 Why two estimates instead of one

The client stated a core count but not a hardware generation or exact server count. Rather than guess, both a lower bound and upper bound are provided using real, currently-supported Dell server families:

Estimate A (2021-era) Estimate B (2023-era)
Servers 563 402
Cores/server 80 112
Total cores 45,040 45,024
Racks per site 11 9
Total racks (3 sites) 33 27

Both estimates are deliberately conservative — they use mainstream, well-documented Dell rack/blade models rather than the newest or highest-density options, so the resulting server and rack counts represent a realistic worst case, not a best case. The true environment likely sits somewhere in this range (a mixed fleet), which is why both models are carried forward together rather than collapsed into a single number. Whichever end of the range the real environment sits closer to, the hardware in both estimates is confirmed compatible with VCF 9.1 (see Section 3 below), so the range does not block planning.


3. Hardware qualification for VCF 9.1

Moving to VCF 9.1 is not just a software upgrade — the underlying hardware (server, storage controller, network adapter, firmware) must be certified for ESXi 9.1, and the storage must be either a certified vSAN ReadyNode configuration or a supported external storage array.

The qualification chain:

VCF 9.1 support
      =  ESXi 9.1 support
      +  supported storage controller / NIC / firmware
      +  vSAN ReadyNode configuration
         OR supported external storage array

All checks are made against the Broadcom Compatibility Guide (the authoritative source for VCF/vSphere hardware compatibility): https://compatibilityguide.broadcom.com/search?program=server&persona=live&column=partnerName&order=asc

3.1 Estimate A hardware (2021-era: R750 / MX750c) — confirmed supported

Dell generation Models ESXi 9.1 support
yx5x / 2021 generation R750, R750xs, R650, R650xs, MX750c, C6520 Yes — ESXi 9.0 and 9.1

Dell’s own ESXi 9.x compatibility matrix confirms the R750 and MX750c, with Intel Xeon SP 83xx/63xx/53xx/43xx CPUs, support both ESXi 9.0 and ESXi 9.1. Source: https://www.dell.com/support/manuals/en-lk/vmware-esxi-9-x/vmware_9.x_compatibility_matrix_pub/dell-yx5x-poweredge-systems

3.2 Estimate B hardware (2023-era: R760 / MX760c) — confirmed supported

Dell generation Models ESXi 9.1 support
yx6x / 2023 generation R760, R660, MX760c, C6620 Yes — ESXi 9.0 and 9.1

Dell’s matrix confirms the R760, R660, MX760c, and C6620, with Intel Xeon SP 64xx/54xx/44xx/34xx CPUs (and Xeon SP 94xx on R760/R660/C6620), support both ESXi 9.0 and ESXi 9.1. Source: https://www.dell.com/support/manuals/en-us/vmware-esxi-9-x/vmware_9.x_compatibility_matrix_pub/intel-sapphire-rapids-processor-and-raptor-lake-processor

3.3 What this means for planning


Summary

Question Answer
Is vSphere 7 still supported? No — General Support ended Oct 2, 2025. This is the primary urgency driver.
Is vSphere 8 still supported? Yes, until Oct 11, 2027 — there is runway, but not indefinitely.
How big is the environment? ~45,000 cores across 3 sites, modeled two ways: 563 servers (2021-era, 80 cores/server) or 402 servers (2023-era, 112 cores/server).
Does the hardware qualify for VCF 9.1? Yes — both the 2021-era (R750/MX750c) and 2023-era (R760/MX760c) Dell families are confirmed on the ESXi 9.1 compatibility list. Storage/controller/NIC/firmware and vSAN ReadyNode vs. external array still need to be confirmed per site during discovery.

Chapter 2: Existing State, External and Global Traffic Management

Client: Strategic enterprise, 3 datacenters (West US, East US, West Europe)

Scope of this document: How user traffic reaches applications today. This describes the current state only. No future-state design is covered here.


1. The short version

Today, all application delivery — deciding which datacenter serves a user, and then load balancing inside that datacenter — is handled entirely by F5 BIG-IP. There is no VCF and no VKS in this picture yet. The systems receiving that traffic are plain vSphere 7/8 virtual machines, running in three separate, standalone datacenters: West US, East US, and West Europe.

There are two F5 components in play, and they do two different jobs:

Component Also known as Job Scope
F5 BIG-IP DNS Formerly “GTM” (Global Traffic Manager) Decides which datacenter should handle a user’s request Global — sits above all three sites
F5 BIG-IP LTM Local Traffic Manager Decides which server/VM inside that datacenter handles the request Local — one instance per site

2. How a request flows today

Existing State Request Flow

The path a request takes is the same three-hop pattern at every site: a global DNS decision, a local load-balancing decision, then the backend VM. Each datacenter box in the diagram also carries its rack footprint from the two conservative sizing estimates, so the traffic-management view and the physical-capacity view tie back to the same three sites:

Datacenter Racks — 2021 estimate Racks — 2023 estimate Rack size
West US 11 9 42U
East US 11 9 42U
West Europe 11 9 42U

Step by step, matching the numbered hops in the diagram:

  1. DNS Query. A user’s device resolves the application’s DNS name. That request lands on F5 BIG-IP DNS/GTM, not a plain DNS server — GTM is authoritative for the application’s public name.
  2. Best Site Selected. BIG-IP DNS/GTM continuously monitors the health and performance of all three datacenters — is the site up, how loaded is it, how far away is the user — and returns the IP address of the best datacenter for that user right now. This is the global decision: which site handles the request. The datacenters themselves take no part in this decision; it is made entirely at the GTM tier.
  3. VIP / Local LB. The user’s traffic then arrives at that site’s F5 BIG-IP LTM. LTM owns the local decision — it terminates the connection at a virtual IP (VIP) for the application, checks the health of individual backend VMs, applies SSL/TLS termination, enforces session persistence, and load-balances the request across the correct pool of VMs using Layer 4/Layer 7 policy.
  4. Delivery. The request finally reaches the target application, running as a virtual machine on vSphere 7 or vSphere 8 inside that datacenter’s rack footprint (11 racks under the 2021 estimate, 9 racks under the 2023 estimate — see the rack-elevation diagrams in the technical documentation).

The key distinction to remember: BIG-IP DNS/GTM makes the global “which site” decision once, up front. BIG-IP LTM makes the local “which server” decision, on every request, inside the chosen site. Together they form the entire application delivery control plane today — there is no other load balancing or ingress layer in front of these VM workloads, and the vSphere layer has no awareness of the other two sites.


3. What sits behind F5 today

Behind the F5 layer, the three datacenters are not yet unified:


4. Why this matters for the modernization project

This existing state is the baseline every future design has to work with and around:


Summary

Question Answer
Who decides which datacenter serves a user? F5 BIG-IP DNS/GTM (global tier)
Who load-balances inside a datacenter? F5 BIG-IP LTM (local tier), one instance per site
What receives the traffic today? vSphere 7/8 virtual machines — no VCF, no VKS
Is there any container-native ingress today? No — this is introduced in a later phase, out of scope for this document

Chapter 3: VKS Adoption Journey, The Four Stage Roadmap

Client: Strategic enterprise, 3 datacenters (West US, East US, West Europe)

Scope of this document: How the client moves from its current, standalone vSphere environment toward VMware Cloud Foundation (VCF) and VMware vSphere Kubernetes Service (VKS), one stage at a time. This is the map of the whole project: every stage below now links to the detailed solution document that was built for it, so this document is the index as much as it is the narrative.

VKS Adoption Journey

1. The journey

The client does not need to adopt the entire VMware Cloud Foundation and VKS stack in one move. The journey happens in four progressive stages, each one adding capability on top of the last, without requiring a rebuild of what came before. A workload domain can sit at any stage for as long as the organization needs — there is no forced timeline to reach Stage 4.

Stages 2 and 3 each have two tracks that run side by side — an upgrade track, brownfield, that imports and modernizes what’s already running, and a VKS track, greenfield, that builds a new, dedicated workload domain for Kubernetes — and this project has a detailed design document for both tracks at both stages, plus everything needed for Stage 4.

2. The four stages at a glance

Stage Name What’s added What it enables Documents
1 Traditional Environment vCenter + ESX only Baseline virtualization — this is where the client is today 01, 02
2 Foundation Introduced VCF Operations, License Server, a VCF Fleet across all three sites Centralized visibility and licensing — the operational foundation for everything after it 04 (fleet architecture), 05 (upgrade track — brownfield import), 07 (VKS track — greenfield)
3 VKS Adoption vSphere Supervisor and VKS, primarily on a new, dedicated workload domain Kubernetes workloads in production, alongside the existing VMs in their upgraded workload domain 06 (upgrade track — Supervisor on the imported workload domain), 08 (VKS track — Supervisor on the new, dedicated workload domain), 09 (load balancer choice — Avi, under either networking stack)
4 Full VKS Operations NSX, SDDC Manager, VCF Automation, VCF Mgmt. Services Full private cloud operating model — self-service, policy-driven automation, fleet-wide lifecycle, and application delivery for the client’s broader estate 10 (F5 → Avi migration)

The key architectural message: each stage is a superset of the one before it. Moving from Stage 1 to Stage 2 does not touch the running VMs — it adds a management layer on top of the existing vCenter and ESX hosts. This is why the client can start modernizing now without putting live production workloads at risk, and why each stage below has a document the client can hand to an implementation team on its own.


3. Stage 1 — Traditional Environment (where the client is today)

This is the client’s current state, fully described in Existing State — Technical Documentation and Existing State — External and Global Traffic Management.

What’s running: - vCenter, managing traditional virtualization operations - ESX hosts underneath it, running VM workloads - F5 BIG-IP DNS/GTM and LTM, handling all global and local traffic management in front of those VMs

What’s not present yet: - No VMware Cloud Foundation - No centralized, multi-site operations layer - No Kubernetes services — infrastructure operations are entirely VM-centric - No software-defined load balancing — F5 is hardware-appliance-based

In plain terms: this is standalone vSphere, administered site by site, the same way the client runs West US, East US, and West Europe today. It is a normal, valid operating model — it’s just the starting point, not the destination.


4. Stage 2 — Platform Foundation Modernization

This is the first real step of the journey, and it is deliberately low-risk: it adds a management and operations layer on top of the existing vCenter and ESX hosts, without replacing them.

What gets introduced: - Upgrade to vSphere 9.1 — brings the hosts onto currently-supported, VCF 9.1-qualified software - VCF Operations — centralized visibility across the environment (monitoring, capacity, health) - License Server — centralized license management, instead of managing licenses per site - Enhanced password and certificate management — tightening operational security practices as part of the same modernization pass - A VCF Fleet spanning all three sites — West US, East US, and West Europe become three VCF Instances under one fleet, detailed in VCF Fleet Architecture

What does not change at this stage: - The same vCenter continues to manage the same ESX hosts - No Kubernetes services are enabled yet — this stage is entirely about the operational foundation, not about running containers - Existing VM workloads are not disrupted; VCF Operations and License Server sit alongside the existing management plane, they do not replace it - F5 keeps doing its job unchanged — Stage 2 does not touch traffic management

Why this stage matters: Stage 2 is the foundation every later stage depends on. Centralized operations visibility and licensing have to exist before it makes sense to roll out Kubernetes (Stage 3) or a full private cloud operating model (Stage 4) across three sites — otherwise each site would be repeating the same manual setup independently.

Two tracks into Stage 2

Every site brings its own existing vSphere + NSX environment, and this project runs two tracks side by side rather than choosing one over the other — a full solution document exists for each:

Track Approach Document
Upgrade track — Brownfield import Deploy a new VCF 9 fleet and import each site’s existing vCenter + NSX environment into it as a workload domain, in place — this is how the existing vSphere 8 estate gets upgraded onto VCF 9.1 Brownfield: Import an Existing vCenter
VKS track — Greenfield Stand up a brand-new, dedicated workload domain using the Multi-Rack Layer 2 vSphere Cluster Model and the Centralized Connectivity Model, purpose-built to bring VKS into production Greenfield: Create New Workload Domains

Because every one of the client’s three sites already runs vSphere with NSX, import (not migration or rebuild) is how the existing environment gets upgraded — it’s the only supported way to bring an existing site into VCF 9.1 without abandoning what’s already there. The greenfield track runs alongside it: a new workload domain, built specifically to host VKS, rather than retrofitted onto the imported estate.

Multi-site reference design: for a client running three datacenters, Stage 2 is where the multi-site operating model gets established, using VMware’s reference blueprint VCF Fleet with Multiple Sites Across Multiple Regions — see VCF Fleet Architecture for the full design, including which VCF Instance hosts the fleet-wide services (VCF Operations, VCF Automation, License Server, SDDC Manager) versus which instances are lighter, instance-level peers.


5. Stage 3 — VKS Adoption

With the Stage 2 foundation in place — the imported workload domain and the new, VKS-dedicated workload domain built alongside it — Stage 3 turns on vSphere Supervisor, which is what actually makes VKS Kubernetes clusters possible. Kubernetes workloads run in the new, purpose-built workload domain, alongside the existing VMs that continue running, untouched, in their upgraded workload domain.

What gets introduced: - vSphere Supervisor, activated on top of the workload domain’s vSphere Zone(s) - VKS Clusters and vSphere Pods, consumed through vSphere Namespaces - A networking stack (NSX Segment Networking or NSX VPC Networking, depending on path) that gives every namespace its own Kubernetes API endpoint, ingress, and load balancer - A load balancer choice for that networking stack — see below

What does not change at this stage: - Existing VM workloads keep running unmodified, in namespaces of their own if needed - The workload domain’s vCenter, NSX, and storage from Stage 2 are reused, not rebuilt - Global traffic management (F5, still) is untouched — Stage 3 is about what happens inside a workload domain, not how users reach it from outside

Two Supervisor models, one per workload domain

Stage 3’s Supervisor deployment model depends on which workload domain it’s being activated on — the imported one, or the new VKS-dedicated one — since each has a different number of vSphere Zones to work with:

Workload domain Approach Zone model Document
Imported workload domain (upgrade track) The simplest supported Supervisor model: one vSphere Zone, combining management and workload, activated directly on the imported cluster where extra Kubernetes capacity is wanted alongside the existing VM estate Single Management Zone with Combined Workload Zones, NSX Segment Networking vSphere Supervisor — Brownfield Single Zone
New, VKS-dedicated workload domain (VKS track) Three vSphere Zones (one per cluster), each combining management and workload, activated on the new, purpose-built three-cluster workload domain Three Management Zones with Combined Workload Zones, NSX VPC Networking vSphere Supervisor — Greenfield Design

The practical difference: the imported workload domain’s single-zone model gives Supervisor host- and cluster-level resilience (via vSphere HA/DRS on the one zone it has); the new workload domain’s three-zone model gives Supervisor zone-level resilience — its three control plane VMs are spread one per zone, so losing an entire zone still leaves Supervisor with quorum. Both are valid designs for their respective workload domain — the choice follows directly from which workload domain Supervisor is being activated on, not a separate decision.

Networking detail matters here too: the imported workload domain uses NSX Segment Networking (a shared Tier-0 Gateway with an auto-created Tier-1 Gateway and load balancer per namespace); the new, VKS-dedicated workload domain uses NSX VPC Networking (one VPC per namespace, inside an NSX Project, aligned with VCF Automation’s project model). Both documents include the concrete IP/CIDR sizing requirements needed to actually build the networking, not just the conceptual model.

The load balancer choice: NSX Load Balancer or Avi Load Balancer

Document: vSphere Supervisor Load Balancer Option: Avi Load Balancer

Whichever networking stack a site ends up on, Supervisor needs a load balancer underneath it, and both networking stacks support the same two choices:

Networking Stack Load Balancer Options
NSX Virtual Private Clouds (VPC) NSX Load Balancer · Avi Load Balancer
NSX Segment Networking NSX Load Balancer · Avi Load Balancer

This is a load-balancer decision, not a zone-model or connectivity decision — it doesn’t change which workload domain (the imported one, or the new VKS-dedicated one) Supervisor is being activated on, and it doesn’t have to be decided at the same time as that choice. The NSX Load Balancer is the simpler built-in default. Avi Load Balancer is the richer choice — the same Controller/Service-Engine architecture used everywhere else in this project — and is worth choosing where a site needs Layer 7 load balancing, WAF, GSLB, or deep traffic analytics beyond what the built-in NSX Load Balancer provides. Document 09 covers the Avi option in full: the 1:1 Avi-Controller-to-NSX-Manager deployment model, the Centralized Transit Gateway data path for Supervisor traffic, and how Avi Controllers and Service Engines place onto the three-region VCF Fleet from document 04.


6. Stage 4 — Full Private Cloud Operating Model

This is the final stage: the client’s environment stops being “a workload domain that happens to run Kubernetes” and becomes a genuine private cloud operating model — self-service provisioning, policy-driven automation, fleet-wide lifecycle management, and, critically, a modern application-delivery layer that finishes replacing what F5 does today.

What gets introduced: - NSX, SDDC Manager, VCF Automation, and VCF Management Services — the full VCF control plane, fleet-wide (already touched on in the Stage 2 fleet architecture, fully active by this stage) - Avi Load Balancer extended beyond Supervisor to the client’s entire application estate, retiring F5 completely

Avi as the F5 replacement, fleet-wide

Document: F5 Replacement: Migrating to VMware Avi Load Balancer

This closes the loop all the way back to Stage 1: the F5 BIG-IP DNS/GTM and LTM estate described in the existing-state traffic management document is retired and replaced by the same Avi platform already available as a Stage 3 load-balancer option (document 09) — Avi GSLB takes over the global “which site” decision from F5 BIG-IP DNS/GTM, and Avi Virtual Services on Avi Service Engines take over the local “which server” decision from F5 BIG-IP LTM. Every F5 concept (pools, VIPs, health monitors, SSL profiles, iRules, partitions) has a direct Avi equivalent, mapped out in full in document 10.

The migration itself is deliberately not a cutover: F5 and Avi run in parallel through a seven-phase migration (discover F5 → build the Avi foundation → recreate local LTM services as Avi Virtual Services → recreate global GTM services as Avi GSLB → pilot one low-risk application → cut over production application by application → decommission F5 in waves). Nothing about this migration is blocked by, or blocks, any other stage — it can run on its own timeline once the Stage 4 fleet-wide control plane exists.

Why this is genuinely the last stage: Stage 4 isn’t complete with just NSX/SDDC Manager/VCF Automation turned on — a private cloud operating model still needs an answer for the application traffic that predates this whole project. Any site that chose Avi as its Stage 3 Supervisor load balancer (document 09) is already running Avi; this stage extends that same platform to carry everything F5 carries today, so the client ends up on one load-balancing platform — Avi — across VM, VCF, and VKS workloads, instead of running F5 and Avi side by side indefinitely. A site that chose the NSX Load Balancer for Supervisor in Stage 3 can still adopt Avi fleet-wide here — the two decisions are independent.

What does not change at this stage: - Workload domains, Supervisor, and VKS clusters built in Stages 2 and 3 are not rebuilt — Stage 4 adds automation, fleet-wide operations, and load balancing on top of them - The migration to Avi is per-application and reversible mid-flight (a rollback path back to F5 is kept open until an application is proven stable on Avi) — nothing is a forced, one-time cutover


7. How the stages connect, end to end

Reading the whole journey as one line, per site:

Stage 1 (standalone vCenter/ESX/F5) → Stage 2 (VCF Fleet established; the existing environment is upgraded via import, and a new, dedicated workload domain is built for VKS via greenfield) → Stage 3 (vSphere Supervisor activated on the new workload domain — and, where wanted, on the imported one too; VKS clusters start running behind the NSX Load Balancer or Avi Load Balancer) → Stage 4 (NSX/SDDC Manager/VCF Automation complete the private cloud operating model; Avi Load Balancer is extended fleet-wide to retire the F5 estate).

Because Stage 2 and Stage 3 each run an upgrade track and a VKS track side by side, and because the three sites don’t have to move through the journey in lockstep, the client can run West US, East US, and West Europe at different stages simultaneously — for example, one site already on Stage 4 while another is still validating Stage 3 — without any stage depending on the others being at the same point.


Summary

Question Answer
Where is the client today? Stage 1 — standalone vCenter + ESX + F5, no VCF, no Kubernetes
What is the first move? Stage 2 — establish the VCF Fleet, then run two tracks side by side: import each site’s existing environment (upgrade track) and build a new, dedicated workload domain for VKS (VKS track), adding VCF Operations and a License Server with no workload disruption
Does Stage 2 touch running VMs or F5? No — it adds an operations/licensing/fleet layer above the existing management plane; F5 is untouched
What are Stage 2’s two tracks? Upgrade track — brownfield import of the existing environment (doc 05); VKS track — a new, dedicated workload domain via greenfield (doc 07)
When is Kubernetes introduced? Stage 3 — vSphere Supervisor is activated on the new, VKS-dedicated workload domain (and optionally on the imported one too), enabling VKS clusters alongside the existing VMs
What are Stage 3’s two Supervisor models? Imported workload domain — single zone, NSX Segment Networking (doc 06); new VKS-dedicated workload domain — three zones, NSX VPC Networking (doc 08)
What load balancer options does Stage 3 offer, and where is that documented? NSX Load Balancer (built-in) or Avi Load Balancer, under either workload domain’s networking stack — see doc 09
What completes the journey at Stage 4? NSX, SDDC Manager, VCF Automation, and VCF Management Services fleet-wide, plus Avi Load Balancer extended to replace F5 across the client’s broader application estate (doc 10)
Does the client have to reach Stage 4 to get value? No — each stage is a valid, superset operating model on its own; there is no forced timeline
Do all three sites have to be at the same stage? No — each site can progress independently, since the upgrade and VKS tracks (and the Stage 3 load-balancer choice) run per site and per workload domain, and Stage 4’s Avi migration is per application

Chapter 4: VCF Fleet Architecture, Multiple Sites Across Multiple Regions

Client: Strategic enterprise, 3 datacenters (West US, East US, West Europe)

Scope of this document: The actual solution design for running VMware Cloud Foundation (VCF) across multiple, geographically separated sites as a single managed fleet. This is the Stage 2 (Platform Foundation Modernization) architecture referenced in the VKS Adoption Journey document.

VCF Fleet with Multiple Sites Across Multiple Regions — Architecture Overview

Source blueprint reviewed for this design: VCF Fleet with Multiple Sites Across Multiple Regions.


1. The solution

Instead of building three separate, disconnected VCF deployments — one per datacenter — the client builds one VCF Fleet: three VCF Instances, one per site, that operate independently day-to-day but are tied together under shared fleet-level management, shared identity, and shared operations visibility. The diagram maps this directly onto the client’s three datacenters: Region A — US WEST, Region B — US EAST, and Region C — EUROPE WEST.

2. The core idea: one “first” instance, everyone else is a peer

This is the single most important design decision in the whole blueprint, and it is easy to miss on first read:

In the diagram, this is why VCF Instance 1 (Region A — US WEST) has a full 4-column grid of components in its Management Domain Cluster, while VCF Instance 2 (Region B — US EAST) and VCF Instance 3 (Region C — EUROPE WEST) each have a lighter 2-column grid — neither duplicates fleet management, licensing, or automation. Both consume those services from Instance 1 over the network.

Site assignment for this design: US WEST is designated the “first” VCF Instance and hosts the fleet-wide services; US EAST and EUROPE WEST run as instance-level peers. US WEST was selected here as the fleet anchor as a working assumption (typically the site with the strongest network position and operational team); this can be revisited during detailed site planning if a different site is better positioned to hold the fleet role.

3. Walking through the diagram

Top bars (span all three sites): - Self-Service with VCF Automation — the self-service consumption layer. Any tenant, in any region, requests infrastructure through the same VCF Automation front end. - VCF Operations — the single-pane-of-glass monitoring and visibility layer across every site in the fleet.

Inside each VCF Instance — Management Domain Cluster: - Management vCenter / Workload vCenter — one vCenter manages the physical hosts, a separate vCenter manages workload placement. This separation is standard VCF practice, repeated identically at every site. - Management NSX Manager Cluster / Workload NSX Manager Cluster — the same split applied to networking: one NSX Manager for infrastructure networking, one for workload networking. - VCF management services — the baseline services every instance needs to operate (present at all three sites). - Fleet-only components (Instance 1 / US WEST in the diagram): VCF Operations Cluster, VCF Operations for Networks Cluster, VCF Automation Cluster, License Server, SDDC Manager, Virtual Network Appliance. - Instance-level component present at all three sites: Cloud Proxy — a lightweight local collector that feeds telemetry back to the fleet-level VCF Operations Cluster in US WEST.

Inside each VCF Instance — Workload Domain Cluster: - NSX Edge (×2) — provides north-south routing and load-balancing capacity for that site’s workloads. - Supervisor VM Cluster — the vSphere Supervisor control plane. This is the component that turns the workload domain into a Kubernetes platform (the mechanism behind VKS in Stage 3).

Below the instances — storage and network fabric: - Storage Site A1 / Storage Site B2 / Storage Site C3 — each site’s own primary storage. Optional asynchronous replication is shown between US WEST ↔︎ US EAST and US EAST ↔︎ EUROPE WEST. This is not synchronous, stretched storage — each site keeps its own storage, and data is optionally copied between adjacent sites for resilience, not shared in real time. Replication does not have to be limited to this chain — any site pair can replicate directly if the client’s disaster-recovery plan calls for it. - Virtual Networking — the NSX-backed overlay running on top of the physical fabric, local to each site. - Physical Network — the underlying physical fabric, connected between sites over standard Layer 3 routing (US WEST ↔︎ US EAST ↔︎ EUROPE WEST). There is no requirement for a stretched Layer 2 network between regions — ordinary routed IP connectivity between sites is sufficient, including across the US-to-Europe hop.

4. Key design decisions (from the reference blueprint)

The blueprint’s design profile makes a specific set of choices for this topology. Restated in plain English:

Area Decision What it means
Consumption Self-service via VCF Automation, plus direct vCenter access Tenants get a self-service portal; platform teams can still go straight to vCenter when needed
Site layout Multiple sites, multiple regions Built for geographic distribution from the start, not retrofitted
Availability Tolerates a single host failure, a single network path failure, a single rack failure, a whole cluster failure, or a single management component failure The design assumes and survives one failure at a time at any layer
Isolation Hypervisor-based, network-based, cluster-based, and tenant-based Multiple independent isolation mechanisms are layered, not relied on individually
Recoverability Supports component backup/restore, full instance backup/restore, and fleet-wide disaster recovery Recovery is designed at three scopes: one component, one whole site, or the whole fleet
VCF Automation Three-node, highly available deployment Survives a node failure without losing the self-service layer
Tenancy model Shared workload domain, multiple tenants isolated by vSphere Namespaces and network segmentation Tenants share physical infrastructure but are logically walled off — no dedicated cluster per tenant required
VCF Operations Three-node analytics cluster + Cloud Proxy per site (a second Cloud Proxy recommended for HA) Fleet-wide monitoring stays available even if one node or one site’s collector fails
Licensing Single, centralized License Server for the whole fleet One place to manage licenses instead of per-site license servers
Log management Three-node deployment behind a load balancer Centralized logging survives a single node failure
Identity / SSO One Identity Broker, fleet-wide single sign-on One login experience across every site in the fleet, not one login per site
Supervisor control plane Three control plane VMs per workload domain, protected by vSphere HA The Kubernetes control plane for VKS survives a single VM or host failure
Load balancing Virtual Network Appliance (VNA), VPC-style networking Load balancing for Supervisor and VKS workloads is delivered as part of the platform, not bolted on separately
Multi-tenant networking Dedicated gateways for tenants needing strict separation, shared gateways for common services Balances strict isolation where it’s required against efficiency where it isn’t
Rack design Single-rack vSphere cluster, single-rack NSX Edge cluster The simplest fault-domain model — each cluster fits and fails within one rack
Storage Single-tier vSAN ESA (all-flash, NVMe only) High-performance storage with a simpler architecture than older tiered vSAN designs

5. What this means for the client’s three sites

Mapped onto West US, East US, and West Europe:

Client site Diagram role Fleet role
West US Region A — US WEST VCF Instance 1 — hosts fleet-wide services (VCF Operations, VCF Automation, License Server, SDDC Manager) plus its own instance components
East US Region B — US EAST VCF Instance 2 — instance-level components only, consumes fleet services from US WEST
West Europe Region C — EUROPE WEST VCF Instance 3 — instance-level components only, consumes fleet services from US WEST

6. Operations monitoring: Cloud Proxy and VCF Operations across three regions

The Cloud Proxy component appears in every VCF Instance in the diagram above (Section 3), but it’s worth explaining on its own because it’s easy to confuse with a networking or traffic component. It is neither.

What Cloud Proxy actually does

Cloud Proxy is the local operations collector and integration bridge for a VCF Instance. It lets the central VCF Operations platform pull metrics, logs, events, tasks, alarms, and health data from the local components at that site — vCenter, ESXi, NSX, vSAN, SDDC Manager, and VCF Management Services — without those components having to talk to the central platform directly over the WAN.

A few things Cloud Proxy is not: - It does not carry application traffic. - It does not replace NSX, SDDC Manager, or F5 BIG-IP. - It is not a second copy of VCF Operations — it is a lightweight local collector that feeds the one central platform.

In a multi-site design, a Cloud Proxy is deployed close to the VCF Instance it monitors — one per region, at minimum — so that operational data collection stays local while visibility stays centralized.

The monitoring model

VCF Operations Monitoring Model — Three-Region Fleet

VCF Operations is deployed once, centrally, in the fleet’s primary management region — US WEST in this design, matching the same site that hosts the other fleet-wide services described in Section 2. It is the single console for dashboards, alerts, capacity planning, log analysis, compliance, lifecycle visibility, inventory, health, and issue tracking across the whole fleet. Broadcom documents VCF Operations as mandatory in VCF 9.x — every fleet has exactly one.

Every region runs its own Cloud Proxy, including US WEST itself. Each Cloud Proxy talks only to its own local VCF Instance and forwards what it collects up to the central VCF Operations platform. West US, East US, and West Europe each get one; a second Cloud Proxy per region is a reasonable addition for local high availability, but a second full VCF Operations cluster is not — East US and West Europe do not need their own VCF Operations deployment unless one of them requires a genuinely independent operations control plane (a separate fleet, data-sovereignty requirement, or mandated regional isolation).

Layer Deployment model
VCF Operations Deployed centrally, once, in the primary management region (US WEST)
Cloud Proxy Deployed in every region — at least one per VCF Instance
VCF Operations for Logs / log management Deployed centrally; every region forwards its logs to it
VCF Operations for Networks Optional but recommended — deploy collectors close to each region if flow analytics and NSX path analysis are required
Guest OS / application agents Optional — install product-managed agents only on the VMs or servers that need OS/application-level monitoring

What each region actually hosts

Region Component Required? Purpose
West US (primary) VCF Operations cluster Yes — central Main dashboards, alerts, capacity, health, compliance for the whole fleet
VCF Management Services Yes Fleet lifecycle, SDDC lifecycle, software depot, licensing services (mandatory per Broadcom)
VCF License Server Yes Centralized licensing for the fleet
Cloud Proxy Yes Local collection for the West US VCF Instance
VCF Log Management Recommended Centralized log collection and analysis
VCF Operations for Networks collector Optional Network visibility, NSX path analysis, flow data
East US Cloud Proxy Yes Collects metrics/events/health from the East US VCF Instance
Local log forwarding Yes, if using logs Sends vCenter, ESX, NSX, SDDC Manager, and workload logs to the central log platform
VCF Operations for Networks collector Optional Local flow and network collection
West Europe Cloud Proxy Yes Collects metrics/events/health from the West Europe VCF Instance
Local log forwarding Yes, if using logs Sends logs to the central log-management platform
VCF Operations for Networks collector Optional Local network flow/path collection

Neither East US nor West Europe gets its own VCF Operations cluster or its own License Server in this design — both are fleet-level services that live once, in US WEST, and are consumed by every region over the network, the same pattern already established for VCF Automation and SDDC Manager in Section 2.

7. Disaster recovery: asynchronous replication across three regions

Section 3 mentioned that storage replication between sites is optional and asynchronous. This section explains what that actually means and how disaster recovery (DR) is implemented on top of it.

What “asynchronous” actually means

Asynchronous replication copies selected VM workloads from one region to another after the write has already been committed at the source site. It is not zero data loss — that’s what synchronous replication provides, and synchronous replication requires very low, metro-distance latency that doesn’t exist between West US, East US, and West Europe. Instead, asynchronous replication gives the client a defined RPO (Recovery Point Objective) — a chosen, known amount of possible data loss, such as 5 minutes, 15 minutes, 1 hour, or 4 hours, depending on how the replication schedule is configured.

Asynchronous Replication — Write Flow

The write itself is never held up waiting for the remote site — production keeps running at full local speed, and the replica catches up shortly after.

Why the client wants this

Asynchronous replication in this design exists for disaster recovery and workload mobility, not for everyday local high availability (local HA is already handled inside each site by vSphere HA and the cluster design covered in Section 4).

Reason Explanation
Regional disaster recovery If West US is lost, critical workloads can be recovered in East US or West Europe
No metro-latency requirement Works across long distances where synchronous replication would be too latency-sensitive
Lower cost and complexity than stretched clusters Avoids stretching storage or cluster dependencies across continents
Controlled RPO/RTO The client defines how much data loss is acceptable and how quickly workloads must recover
Non-disruptive DR testing Recovery plans can be tested without stopping production
Migration support The same replication mechanism can help move workloads between sites, not just recover them
Ransomware / corruption recovery With multiple recovery points or snapshot retention, the client can recover from an earlier point in time, not just the latest (possibly already-corrupted) replica

Broadcom describes the VCF DR model as a primary site and recovery site architecture: production workloads are replicated to a secondary site so they can resume operations if the primary site becomes unavailable. Recovery sites do not need to be physically identical to the primary — they only need enough resources to run the protected workloads, not a mirror-image build.

The clean VCF 9.1 implementation combines:

Protection topology

VCF Disaster Recovery — Protection Topology

Each VCF Instance runs its own VLR / Site Recovery locally — DR is not a separate fourth platform, it’s a capability layered onto the same three VCF Instances already described in this document, with status rolling up into the same central VCF Operations console used for day-to-day monitoring (Section 6).

Two ways to structure who protects whom:

Recommendation: pair sites by application criticality and business ownership, not by blindly replicating everything everywhere — not every workload needs cross-region DR, and treating all of them identically adds cost without adding protection where it matters.

Example protection pairing:

Source site Recovery site Use case
West US East US Primary US DR
East US West US Reverse DR
West Europe East US (per Option A) Cross-region DR — confirm against data-residency requirements before finalizing; a dedicated EU-based recovery target may be required instead if data cannot leave the region

Summary

Question Answer
Is this three separate VCF deployments or one? One fleet — three VCF Instances (US WEST, US EAST, EUROPE WEST) under shared fleet-level management
Does every site run the same components? No — US WEST runs fleet-wide services in addition to its own; US EAST and EUROPE WEST run instance-level components only
Is storage shared or stretched between sites? No — each site keeps its own storage; cross-site replication is optional and asynchronous
What connects the sites at the network layer? Standard routed Layer 3 — no stretched Layer 2 required, even to Europe
Does this design already account for VKS? Yes — the Supervisor VM Cluster in each Workload Domain Cluster is the same component VKS activates in Stage 3
How is the fleet monitored? One central VCF Operations cluster in US WEST; a local Cloud Proxy in every region forwards data to it — no per-region VCF Operations cluster needed
Does replication between sites guarantee zero data loss? No — it’s asynchronous, giving a defined RPO (e.g. 15 minutes), not synchronous zero-data-loss replication
Who protects whom for DR? Recommended default: West US ↔︎ East US pair, plus West Europe → East US — confirm West Europe’s target against data-residency requirements

Chapter 5: Stage 2 Option 1, Brownfield Import of an Existing vCenter

Client: Strategic enterprise, 3 datacenters (West US, East US, West Europe)

Scope of this document: One clear message — because every client site already runs vSphere with NSX, the only supported way to upgrade West US, East US, and West Europe onto VMware Cloud Foundation (VCF) 9.1 is to deploy a new VCF 9 fleet and import each existing vSphere + NSX environment into it as a workload domain. This document explains that upgrade track in full: what it means, what it requires, how it works, and exactly what happens to the existing NSX network. This is Stage 2 of the modernization journey. The VKS track — building a new, dedicated workload domain to bring VKS into production — is covered separately and is out of scope here.

Reference: Import an Existing vCenter to Create a Workload Domain (Broadcom techdocs).


The mechanism this upgrade relies on

Import an Existing vCenter to Create a VCF 9.1 Workload Domain. VCF 9.1 can bring an already-running vCenter under VCF management directly. The existing vCenter, its clusters, hosts, storage, and NSX Manager all come in as-is — nothing is copied, moved, or rebuilt. This is the mechanism the entire upgrade track relies on.

Deploy New VCF 9 Environment / Import Workload Domain. The general pattern: stand up a brand-new VCF 9 fleet on unused hardware, then import an existing environment into it afterward.

Transition Path 1 — Deploy New VCF 9 Environment / Import Workload Domain

How this plays out for the client’s environment

Brownfield: Import an Existing vCenter to Create a VCF 9.1 Workload Domain

This is the same three-phase pattern shown above, expanded with every NSX-specific check, precheck, and validation step that matters for an environment that already has NSX deployed — the level of detail an implementation team actually needs.

What “import” actually means

This is not a migration tool — nothing is copied, moved, or rebuilt. Broadcom’s own framing is simple: you import the existing vCenter, clusters, hosts, storage, and NSX into VCF management. The VMs stay exactly where they are, and the existing NSX Manager — with its Tier-0/Tier-1 gateways, segments, and firewall policy — comes in as-is.

Before: - Standalone vSphere 8 environment (if a site is still on vSphere 7, it must be upgraded to a supported vSphere 8 build first — vSphere 7 cannot be imported directly) - vCenter 8.x, ESXi 8.x clusters, VM workloads - vSAN / FC / NFS / iSCSI storage - Existing NSX Manager (4.2.x), registered to vCenter, with Tier-0/Tier-1 gateways and segments already configured

After import: - A VCF 9.1 Instance now exists, with its Management Domain (SDDC Manager, VCF Operations, VCF Management Services) - The existing environment becomes an Imported VI Workload Domain inside that instance — same vCenter, same ESXi clusters, same VM workloads, same storage, same NSX Manager with the same T0/T1 gateways and segments, now visible and managed through VCF

One important constraint: the import is all-or-nothing at the vCenter level. Every cluster inside that vCenter comes in together — there is no cluster-by-cluster import. If some clusters should end up in different workload domains, or if the vCenter contains unsupported or legacy clusters, that has to be resolved (split, reorganized, or remediated) before the import runs, not after.

Requirements

Prerequisites (must be true before you start)

Minimum component versions

Component Minimum version
VCF Instance 9.0 or later
VMware vCenter 8.0 Update 3a or later
VMware ESX 8.0 Update 3 or later
NSX Manager 4.2 or later — required, existing registration expected

The vCenter and ESX minimums both start at 8.0 — there is no path that imports a vSphere 7 environment directly. A site still on vSphere 7 has to be upgraded to a supported vSphere 8 build first, as a standalone step, before the import prerequisites above can be satisfied.

One version trap worth flagging explicitly: NSX 9.1 does not support vCenter 8.0 Update 3a. If the plan is to bring the existing NSX Manager onto a shared NSX 9.1 instance (or upgrade it to a 9.1 NSX instance as part of this workstream), the vCenter must first be upgraded to 9.1 — importing at 8.0 U3a and expecting NSX 9.1 compatibility does not work. Since the client’s environment already runs NSX, this version pairing has to be checked and resolved before import, not discovered during it.

Additional NSX-specific checks and requirements

Because the client’s environment already has NSX deployed, these checks matter more than the general “no NSX yet” path documented by Broadcom — validate all of them before scheduling an import:

Supported configurations

These are the configuration boundaries the existing environment has to fit inside before import. Anything outside this list needs to be remediated first.

Storage — supported: - Enough free space for a full VCF deployment - A datastore shared across, accessible from, and writable by every host in the cluster - Any supported vSphere storage type - vSAN Stretched Clusters (minimum 3 ESX hosts per availability zone, plus a witness host) - Two-node vSAN clusters for ROBO deployments (2 hosts + 1 witness) - vSAN OSA clusters, provided deduplication/compression settings match across hosts

Storage — not supported: vSAN OSA clusters using compression-only configurations.

Network — supported: - vSphere Distributed Switch (VDS) 8.0 or later - Statically assigned VMkernel IPs (NSX Host TEPs are the one exception — DHCP is fine there) - A dedicated network for vSphere vMotion - VDS with LACP enabled - Clusters sharing a VDS - DNS with both forward and reverse records - A shared NSX Manager across multiple vCenter instances (as long as Enhanced Linked Mode isn’t in use) - NSX Bare Metal and VM Edge nodes

Network — not supported: Cisco virtual switches, a vCenter without a VDS, custom distributed port groups, non-default vCenter ports, dynamically allocated VMkernel IPs, or multiple NSX Managers managing a single vCenter.

Compute — supported: - The vCenter VM hosted on the default cluster in the management domain - Clusters using vSphere Configuration Profiles - Clusters using vSphere Lifecycle Manager images - Clusters using fully automated vSphere DRS - Standalone ESX hosts or single-host clusters, as long as at least one other compliant cluster also exists

Compute — not supported: vCenter with Enhanced Linked Mode, an NSX Manager already shared across different VCF instances, baseline-based lifecycle management, manual or partial DRS, partial cluster imports, vCenter HA (VCHA) clusters, or a vCenter that’s already connected to SDDC Manager.

What this means for the client: before importing any of the three datacenters, each site’s actual storage/network/compute configuration needs to be checked against this list. A site running Cisco virtual switching, VCHA, or partial DRS, for example, would need remediation before it qualifies — this is a concrete discovery task, not a formality.

The import workflow

Walking through the diagram’s middle phase:

  1. In VCF Operations, go to Operate → Inventory → Detailed View, select the target VCF Instance, and choose “Import a vCenter”.
  2. Enter a name for the new workload domain.
  3. Select the vCenter to import, plus its existing, already-registered NSX Manager.
  4. Confirm the certificate thumbprints.
  5. The vCenter registers as a Compute Manager in NSX. Clusters already NSX-prepared stay prepared; clusters that weren’t prepared before import remain unprepared.
  6. If Edge clusters exist on that NSX Manager (NSX 9.0+), the Edge node VMs are auto-discovered and pulled into VCF inventory — their credential passwords are reset as part of this step.
  7. Run the built-in prechecks — this is the gate that catches unsupported configurations before anything changes.
  8. Review the settings and click Finish.

A few behaviors worth knowing about going in: existing IPv4/IPv6 dual-stack networking carries over as-is (subject to the IPv4-only constraint noted in “Additional NSX-specific checks and requirements” above if the NSX instance is shared into an IPv4-only domain). And the import workflow automatically activates NSX on the Distributed Virtual Port Groups it finds, which in turn automatically turns on the Distributed Firewall (DFW) — with default-allow rules — across all DVPGs in the imported cluster. That’s a security-relevant side effect to plan for, not just a technical detail; it’s addressed directly in “Post-import lifecycle” below.

Post-import lifecycle: aligning to the VCF 9.1 BOM

Being imported does not mean being upgraded. Right after import, the workload domain is under VCF management but is still running whatever vCenter/ESXi/NSX versions it had going in. Getting it onto the VCF 9.1 Bill of Materials (BOM) is a separate, subsequent lifecycle step, shown as the third phase in the diagram:

  1. Imported vSphere 8 workload domain — VMs untouched, now visible and managed through VCF
  2. Review DFW activation on all DVPGs in the imported cluster — the import turns DFW on with default-allow rules for Layer 2 and Layer 3 traffic. This is not a hardened state; treat it as a checklist item on day one, not an assumption of “secure by default.” Two options: tighten the DFW rules to the client’s actual policy immediately, or — if DFW activation isn’t wanted yet — deactivate NSX on the affected DVPGs via the NSX UI or the Transport Node Collection API, and revisit later.
  3. Validate existing T0 / T1 gateways and segments — confirm the existing NSX logical topology came through unchanged (see “What happens to the existing NSX T0, T1 gateways, and segments” below for what to expect and how to validate this).
  4. Apply configuration updates — align local configuration to VCF-managed expectations
  5. Download / map VCF 9.1 binaries — pull the correct binaries for the target BOM
  6. Run VCF lifecycle prechecks — the gate before any upgrade is applied
  7. Upgrade NSX / vCenter / ESXi as required — brought up to the versions the VCF 9.1 BOM specifies
  8. Workload domain aligned to VCF 9.1 BOM — the end state, fully lifecycle-managed by SDDC Manager going forward

This two-step separation (import first, upgrade second) is what keeps the brownfield path low-risk: the import itself is a management-plane change, not a software upgrade, so it doesn’t carry upgrade risk. The upgrade to the VCF 9.1 BOM happens afterward, on its own schedule, with its own prechecks gate.

What happens to the existing NSX T0, T1 gateways, and segments

This is the question every network team asks first, and it deserves a direct answer: the import does not touch NSX logical network topology.

Why: the vCenter import operates at the compute manager registration level — it links (or confirms the link between) vCenter and its existing NSX Manager, and it determines which clusters are NSX-prepared. It does not read, rebuild, or recreate Tier-0 gateways, Tier-1 gateways, segments, groups, or DFW policies. Those are configuration objects that live inside the existing NSX Manager, and that NSX Manager is brought into VCF’s managed inventory as a whole, unchanged — not disassembled and reconstructed. Broadcom’s own description of the “existing NSX registration” path backs this up directly: clusters that are not already NSX-prepared are explicitly not reconfigured during import — the import only acts on what’s eligible, and leaves everything else alone.

What that means concretely:

Recommended validation step: Broadcom’s documentation does not explicitly state “T0/T1 and segments are preserved unchanged” in those exact words — that behavior is inferred from how the import is scoped (compute-manager registration, not network reconfiguration). Because of that, don’t just assume it — export or snapshot the existing NSX Manager’s Tier-0/Tier-1 gateway configuration, segment list, and DFW rules before the import, and diff that against the same configuration after the import completes, before signing off on the import as done. This turns an inference into a verified fact for this specific environment.

Where this fits for the client

Client scenario Supported approach
One vCenter, all clusters belong in one workload domain Good fit — straightforward import
One vCenter, but clusters should end up in different workload domains Split or reorganize before import
One vCenter with unsupported or legacy clusters Remediate or separate those clusters before import
A site still running vSphere 7 Upgrade to a supported vSphere 8 build first — vSphere 7 cannot be imported directly
NSX Federation in use at a site Only the Local Manager is imported — validate the Global Manager relationship and any stretched T0 gateways separately, per site
Multiple NSX Managers registered to one vCenter Not supported — consolidate to a single NSX Manager per vCenter before import

Post-import, two operational limitations to plan around: hosts cannot be added to an imported cluster without going through the vSphere Client (not through VCF Operations directly), and password management via the VCF Operations Console isn’t available for imported domains out of the box (Broadcom documents a workaround under KB 388859).


Summary

Question Answer
Which transition path applies to the client? Deploy a new VCF 9 fleet, then import each existing vSphere + NSX environment as a workload domain — there is no supported in-place upgrade of a vSphere + NSX environment straight to VCF; import is the only path
Does this move or rebuild the VMs? No — the VMs stay exactly where they are; only the management relationship changes
Can I import part of a vCenter? No — every cluster in that vCenter is imported together, no cluster-by-cluster selection
What versions does the source environment need? vCenter 8.0 U3a+, ESX 8.0 U3+, NSX Manager 4.2+ (required — existing registration expected), into a VCF 9.0+ Instance
Is the workload domain upgraded to VCF 9.1 automatically during import? No — import and BOM alignment are two separate steps; upgrade happens afterward with its own prechecks
Can a vSphere 7 environment be imported directly? No — it must be upgraded to a supported vSphere 8 build first
What’s the biggest hidden gotcha? NSX 9.1 does not support vCenter 8.0 U3a — upgrade vCenter to 9.1 first if a 9.1 NSX instance is the target
What happens to existing T0/T1 gateways and segments? They are retained unchanged — the import registers vCenter as a Compute Manager in NSX, it does not rebuild NSX logical topology. Validate with a before/after config diff (see “What happens to the existing NSX T0, T1 gateways, and segments”)
Does DFW get turned on automatically? Yes — on every DVPG in the imported cluster, with default-allow rules. Review and tighten immediately after import

Chapter 6: Stage 3 Option 1, Brownfield vSphere Supervisor Single Zone

Client: Strategic enterprise, 3 datacenters (West US, East US, West Europe)

Scope of this document: The Supervisor design for the imported, upgrade-track workload domain — Stage 3 of the modernization journey, for sites that want extra Kubernetes capacity on the imported estate itself, in addition to the new, VKS-dedicated workload domain covered separately. This covers the simplest supported Supervisor deployment model: one vSphere Zone, its storage, and its network connectivity. The Supervisor design for the new, VKS-dedicated workload domain (multiple vSphere Zones across multiple clusters) is a separate document and out of scope here.

References: Single Management Zone with Combined Workload Zones Model and NSX Segment Connectivity Model Overview (Broadcom techdocs), plus “vSphere Supervisor with NSX Segment Networking: Architecture” (VCF Solution Architecture and Design training material, page 2-35).


The solution in one paragraph

Once the imported workload domain exists (via the import path covered in the previous document), adding Kubernetes capacity to it is a matter of turning on vSphere Supervisor on top of it. The simplest way to do that is the Single Management Zone with Combined Workload Zones Model: one vSphere Zone, mapped to one vSphere Cluster, does double duty as both the Supervisor’s management zone and its workload zone. No second cluster, no extra zone, no extra complexity — just the existing (imported) cluster, activated for Supervisor.

Design decisions at a glance

vSphere Supervisor Design for VKS Consumption

Four decisions define this model, and all four are the simplest available option:

Decision Choice
How many management zones? One (the default)
How many workload zones? One — combined with the management zone, not separate
What kind of datastore? Zonal datastore (local to the zone)
When does it get switched on? Either at workload domain creation, or later from vCenter — the client’s choice

The zone model itself

Single Management Zone with Combined Workload Zones Model

Reading the diagram top to bottom, this is what “turning on Supervisor” actually builds:

Why this is the right starting point for the client: every one of the three imported workload domains (West US, East US, West Europe) already has exactly one vSphere Cluster per site in the simplest case. This model activates Supervisor directly on that cluster — no re-architecture, no second cluster to build first.

Key design considerations that shaped this model (per Broadcom’s design guidance): vSphere Supervisor cluster topology, vSphere Supervisor zone design, and control plane VM placement and sizing (Small / Medium / Large — sized to the expected number of namespaces and API load, not fixed by this model).

A hard requirement, not optional: the vSphere Zone selected for Supervisor must have vSphere HA enabled and vSphere DRS running in Fully Automated or Partially Automated mode, and it cannot already be assigned to another Supervisor instance. Both the vSphere Zone name and the Supervisor name must also follow RFC 1123 naming rules (lowercase, alphanumeric, hyphens only, starting with a letter or number) — worth checking early, since a bad name is a rejected activation, not a warning.

Storage: the zonal datastore

Storage Topology — Zonal Datastore

A zonal datastore is simply a datastore that’s scoped to one vSphere Zone:

In plain terms: storage doesn’t need to be redesigned for this model. Whatever datastore already backs the imported cluster becomes the zonal datastore for Supervisor, as long as it meets the standard requirements (shared across, accessible from, and writable by every host — the same storage rule already covered in the import requirements).

Networking: vSphere Supervisor with NSX Segment Networking

vSphere Supervisor with NSX Segment Networking — Architecture

This is the actual wiring diagram for how Supervisor’s networking works once NSX Segment Networking is in use — not just the abstract Tier-0/Tier-1 model, but where the load balancers, control plane VMs, pods, and ESX hosts each sit, and how they all reach the management plane.

The two halves of the diagram:

What the load balancer actually exposes, and why it’s different on each side: - On the vSphere Namespaces side, the load balancer publishes three things: the apiserver virtual IPs for the VKS Kubernetes Cluster (so users can reach that cluster’s Kubernetes API), the Kubernetes Load Balancer Service (for type: LoadBalancer services created inside that cluster), and Ingress Virtual IPs (for ingress traffic into that cluster). - On the Supervisor System Namespaces side, the load balancer only needs two: the apiserver virtual IPs for the Supervisor cluster itself, and Ingress virtual IPs for Supervisor-level ingress. There’s no separate Kubernetes Load Balancer Service entry here, because this side isn’t running tenant workloads — it’s running Supervisor itself.

The functional behaviors that matter operationally (from the architecture reference):

  1. NCP creates one shared Tier-1 Gateway for system namespaces, and a separate Tier-1 Gateway plus load balancer for each namespace, by default. Every Tier-1 Gateway connects up to the same Tier-0 Gateway and down to a default segment — this is automatic, not something to hand-configure per namespace.
  2. Workloads inside the same namespace share one SNAT IP for north-south traffic. From the outside, all outbound traffic from a namespace’s workloads looks like it’s coming from one address.
  3. East-west traffic between namespaces happens without SNAT. Namespace-to-namespace connectivity is preserved at the real source IP — SNAT is a north-south (outbound-to-the-world) behavior only, not something that gets in the way of workloads talking to each other.
  4. Each VKS cluster gets its own separate network segment inside its Tier-1 Gateway, specifically for isolation — one VKS cluster’s pod network is not the same broadcast domain as another’s.
  5. The Tier-0 Gateway is associated with the NSX Edge cluster, which is what actually provides the routing path out to the external network via the uplink VLAN.

Tying it back to the management plane: both halves ultimately depend on the same NSX Manager and vCenter, which sit on the Management Network — the same management plane that governs the rest of the imported workload domain. Nothing about enabling Supervisor networking introduces a second, separate management path.

Network Requirements for Supervisor with NSX Segment Networking

The diagram above shows the shape of NSX Segment Networking; this section gives the concrete numbers — how many IPs, which CIDR sizes, which components — needed to actually build it.

1. Management Network Requirements

Component Minimum Quantity Required Configuration
IPs for Supervisor control plane VMs Block of 5 A block of 5 consecutive DHCP or static IP addresses, assigned from the Management Network to the Supervisor control plane VMs.
Management Network Subnet 1 5 IP addresses for the Supervisor control plane: 1 for each of the 3 nodes, 1 for the virtual IP, 1 for rolling cluster upgrade.
Management traffic network 1 A Management Network that is routable to the ESXi hosts, vCenter, Supervisor, and the load balancer.
Management Network VLAN 1 The VLAN ID of the Management Network subnet.

2. Workload Network Requirements

Component Minimum Quantity Required Configuration
vSphere Pod CIDR range /23 private IP addresses A private CIDR range that provides IP addresses for vSphere Pods. These same addresses are also used for VKS cluster nodes.
Kubernetes services CIDR range /16 private IP addresses A private CIDR range to assign IP addresses to Kubernetes services.
Egress CIDR range /27 static IP addresses (minimum) A private CIDR used to determine the egress IP for Kubernetes services. Only one egress IP address is assigned per namespace on the Supervisor.
Ingress CIDR /27 static IP addresses (minimum) A private CIDR range used for the IP addresses of ingresses.

3. NSX Network Requirements

Component Minimum Quantity Required Configuration
Load balancer IPs 3 A set of 3 routable IPs for external connectivity: Kube API server, Management Proxy, and vSphere CSI.
VLANs 3 IP addresses for the tunnel endpoints (TEPs). Both the ESXi host TEPs and the Edge TEPs must be routable.
Tier-0 Uplink IP /24 private IP addresses The IP subnet used for the Tier-0 uplink.

Notes on TEPs and VLANs: - ESXi hosts and NSX Edge nodes both act as tunnel endpoints — a TEP IP is assigned to each host and each Edge node. - The ESXi host VTEP and the Edge VTEP must both have an MTU size greater than 1600. - Because ESXi host TEP IPs form an overlay tunnel with the Edge nodes’ TEP IPs, the VLAN IPs on both sides must be routable to each other. - A separate, additional VLAN is required to provide North-South connectivity to the Tier-0 Gateway. - IP pools can be shared across clusters, but the host overlay IP pool/VLAN must not be shared with the Edge overlay IP pool/VLAN — unless host TEP and Edge TEP traffic use different physical NICs, in which case they can share the same VLAN.

NSX Federation — a note for a 3-site client

Because the client operates three datacenters, NSX Federation is directly relevant, not a theoretical feature:

What this means for the client’s three sites


Summary

Question Answer
How many vSphere Zones does this model use? One — it serves both Supervisor management and application workloads
Does this require a second cluster? No — the same imported cluster is reused
What storage does Supervisor use? A zonal datastore, local to that one vSphere Zone
How many Control Plane VMs run? Three, for resilience — not one
Can a namespace run plain VMs instead of Kubernetes? Yes — a namespace can host either a VKS Cluster or plain Virtual Machines
What connects namespaces to the outside world? NSX Segment Networking — a shared Tier-0 Gateway, with NCP auto-creating a dedicated Tier-1 Gateway and load balancer per namespace
Does traffic between namespaces get SNAT’d? No — only north-south (outbound) traffic is SNAT’d; east-west traffic between namespaces keeps its real source IP
Does this model support NSX Federation? Yes — full support, which is the main reason to choose it over a VPC-based model
When can Supervisor be activated? Either at workload domain creation or later from vCenter, per site, independently

Chapter 7: Stage 2 Option 2, Greenfield Workload Domains

Client: Strategic enterprise, 3 datacenters (West US, East US, West Europe)

Scope of this document: The VKS track — standing up a brand-new, dedicated VCF workload domain to bring VKS into production — Stage 2 of the modernization journey, run alongside the brownfield upgrade track covered separately. This covers the two design building blocks that define this workload domain: the Multi-Rack Layer 2 vSphere Cluster Model (compute/fault-domain design) and the Centralized Connectivity Model with Shared Tier-0 Gateway Per Tenant (networking design). Stage 3 (enabling VKS on this workload domain) is a separate, later step and is out of scope here.

References: Multi-Rack Layer 2 vSphere Cluster Model and Centralized Connectivity Model (Broadcom techdocs).


The solution in one paragraph

Greenfield is the opposite of the brownfield import path covered elsewhere in this project: instead of bringing an existing vSphere + NSX environment under VCF management as-is, a brand-new workload domain is built from scratch on new or repurposed hardware. That means every design decision — compute layout, fault tolerance, network connectivity — gets made fresh, rather than inherited from whatever was already running. The two decisions that matter most for a new workload domain are covered here: how the compute cluster is laid out across racks for resilience, and how tenants get network connectivity to the outside world.

Compute: Multi-Rack Layer 2 vSphere Cluster Model

Multi-Rack Layer 2 vSphere Cluster Model

The problem this solves: a single-rack cluster protects against a host failure, but not against losing an entire rack (a failed top-of-rack switch, a power incident, a maintenance mistake). This model protects against exactly that, by spreading vSphere cluster hosts across multiple racks.

How it works, reading the diagram bottom to top: - Hosts sit physically in four separate racks (Rack 1 through Rack 4), two hosts per rack in this example. - All of this sits inside one Workload Domain — this is not three separate workload domains, it’s a single workload domain built with three internal clusters. - The diagram shows three Layer 2 vSphere Clusters, and each cluster is a vSphere Zone — Cluster 1 is Zone 1, Cluster 2 is Zone 2, Cluster 3 is Zone 3. Each is its own pool of capacity, independently managed by its own vSphere HA and DRS, and every one of the three is stretched across the same four physical racks. - Every infrastructure network a given cluster/zone needs — VM Management, ESX Management, vMotion, IP Storage (NFS/vSAN), NSX Host TEP — is stretched across all four racks as a single Layer 2 broadcast domain for that zone. A host in Rack 3 is on the exact same VM Management network as a host in Rack 1 within its own zone, not a routed, separate one. - Above that, each rack’s switches carry all three zones’ stretched Layer 2 traffic between racks, all within one set of Layer 2 Adjacencies.

Scenario: dedicated racks per zone (six racks)

Multi-Rack vSphere Cluster and Zone Model — Dedicated Racks Per Zone

The diagram above stretches all three zones across the same four racks. That’s one valid physical layout — it isn’t the only one. Here’s a second, equally valid scenario, not an alternative option to choose instead of the first: six racks are available, and each zone is confined to its own dedicated pair of racks rather than spreading across all of them.

Multi-Rack vSphere Cluster and Zone Model

One VCF Workload Domain contains three vSphere clusters. Each vSphere cluster is mapped to one vSphere Zone. The ESXi hosts for each cluster are distributed across multiple racks to provide rack-level resiliency within the zone.

The vSphere Zone is a logical availability construct associated with a vSphere cluster — it is not a rack-level object. Physical racks provide the underlying host placement and failure-domain design; the zone is the logical boundary layered on top of that physical placement, and the two don’t have to line up the same way every time.

Workload Domain
  |
  +-- vSphere Zone 1 / Cluster 1
  |     +-- Hosts distributed across Rack 1, Rack 2
  |
  +-- vSphere Zone 2 / Cluster 2
  |     +-- Hosts distributed across Rack 3, Rack 4
  |
  +-- vSphere Zone 3 / Cluster 3
        +-- Hosts distributed across Rack 5, Rack 6

How this differs from the first diagram: in the four-rack layout, all three zones share the same four racks — every zone’s hosts are physically interleaved with the other two zones’ hosts. In this six-rack scenario, each zone gets its own pair of racks, with no physical overlap between zones at all. Both give each zone the same logical protection (loss of one rack doesn’t take out the whole zone, because a zone’s hosts sit in two racks) — the difference is purely how the physical racks are assigned, not how the zones behave logically.

When dedicated racks per zone make more sense than shared racks: when the client wants a zone’s failure domain to be physically, not just logically, independent of the others — for example, if Zone 1’s racks are on a different power feed, a different network uplink pair, or even planned for a different physical location within the datacenter than Zones 2 and 3. Sharing racks (the first diagram) is more space-efficient; dedicating racks (this scenario) is more strictly isolated.

Why three zones: preparing for the Three Management Zones Supervisor model

This three-cluster layout isn’t compute design for its own sake — it’s building the exact physical foundation Stage 3’s greenfield Supervisor design needs: the Three Management Zones with Combined Workload Zones Model.

Requirements per zone (per cluster):

Each of the three zones (Cluster 1 / 2 / 3)
Minimum ESX hosts 3 minimum to tolerate one host failure per zone
vSphere HA Required
vSphere DRS Fully automated (recommended)
Lifecycle management Its own vSphere Lifecycle Manager image, independent of the other two zones
Stretched networks Its own VM Management, ESX Management, vMotion, IP Storage, NSX Host TEP — a separate broadcast domain per zone, not shared with the other two
NSX transport nodes Every host in the zone
Role Combined — runs both Supervisor management components and application workloads

Total requirements, across all three zones:

Why this is simpler than it sounds: because every network is one broadcast domain instead of one-per-rack, this model needs fewer VLANs and subnets than a Layer 3 multi-rack design would. The tradeoff is that Layer 2 has to actually be stretchable across those racks in the client’s physical network — this model assumes that capability exists, it doesn’t create it.

Resilience, concretely: protection operates at two levels — vSphere HA protects against a single host failing, and the multi-rack spread protects against an entire rack failing. A cluster sized to tolerate one host failure keeps running even if a whole rack goes dark, because the surviving hosts are physically in other racks.

Sizing, at minimum:

These are the general-purpose VCF cluster minimums from the design guide, for reference — they apply when a cluster is dedicated to one role only (a traditional, separate management domain, for example):

Cluster type Storage Minimum ESX hosts
Management Domain (Simple) vSAN 3
Management Domain (HA) vSAN 4
Additional Management Clusters vSAN 3
Management Domain (Simple) VMFS on FC/NFS 2
Management Domain (HA) VMFS on FC/NFS 4

In the three-zone Combined Workload Zones model used here, that traditional split doesn’t apply — each zone already carries both roles, which is why the per-zone minimum above is 3 hosts (the same floor as a simple management cluster), not a separate, larger number for a dedicated management cluster plus more for workload clusters on top of it.

Requirements worth calling out explicitly: - One set of clusters per VCF workload domain — this workload domain’s three zones are not shared with any other workload domain. - vSphere Lifecycle Manager images (not baselines) — this keeps firmware and vendor add-ons manageable as one unit. - vSphere HA enabled for every cluster, full stop. - Every ESX host configured as an NSX transport node, under a single overlay transport zone.

Practical recommendations from the design guide: use separate vSphere Distributed Switches to keep storage, NSX, and vMotion traffic apart (minimum 2 NICs per switch); prefer vSAN ESA as the primary storage engine; assign NSX Host TEP IPs from static pools rather than DHCP; and enable vSphere DRS in fully automated mode so load balances across racks without excessive vMotion churn.

Networking: Centralized Connectivity Model with Shared Tier-0 Gateway Per Tenant

Centralized Connectivity Model with Shared Tier-0 Gateway Per Tenant

The problem this solves: every tenant needs a path to the outside world, but building a dedicated Tier-0 Gateway (and its NSX Edge nodes) per tenant is expensive and often unnecessary. This model lets multiple tenants share one Tier-0 Gateway and its Edge node resources, while still keeping each tenant’s networking logically separate.

What is a “tenant,” concretely?

“Tenant” is a deliberately generic word in NSX’s design language — an NSX Project — and it can map to whatever organizational boundary actually needs network isolation. It doesn’t have to mean a different company or a different customer. For an enterprise client running its own private cloud, a tenant is most often an internal organizational boundary:

The one property every tenant shares, no matter what it represents organizationally: it gets its own Centralized Transit Gateway, its own VPC(s), and its own IP address space, and traffic never crosses from one tenant to another except through whatever explicit routing/firewall policy is deliberately configured — never by default.

How it works, reading the diagram top to bottom: - External Connectivity and BGP bring routes in from the physical network to a single, shared Tier-0 Gateway, running in Active/Active mode. - Each Tenant (an NSX Project) gets its own Centralized Transit Gateway (C-TGW), connected to that shared Tier-0. The C-TGW runs in Active/Standby mode and is where tenant-level NAT and VPN services attach. - Inside each tenant, one or more VPCs connect to that tenant’s C-TGW, each with its own VPC Gateway, its own NAT, and its own Subnet carrying the actual workload VMs. - This is the same building block VCF Automation’s tenancy models are built on, and it also supports vSphere Supervisor deployments directly — it’s not exclusive to one consumption model.

Why share Tier-0 instead of dedicating one per tenant: it’s simply more resource-efficient — fewer NSX Edge nodes to deploy and operate, while tenants still get logical isolation through their own C-TGW and VPCs. The tradeoff is that all inter-VPC and north-south traffic for every tenant ultimately passes through the same shared Edge nodes, so there’s a real (if manageable) risk of resource contention between tenants under heavy load — worth monitoring, not a reason to avoid the model outright.

When to choose the shared model vs. a dedicated Tier-0 per tenant:

Choose… When…
Shared Tier-0 (this model) Tenants have broadly similar traffic patterns, infrastructure is cost/resource constrained, and some shared-capacity risk is acceptable
Dedicated Tier-0 per tenant A tenant needs guaranteed performance, its own VPN or advanced routing, or strict traffic isolation is mandated

Requirements worth calling out explicitly: - A minimum of two NSX Edge nodes in the Edge cluster, for high availability. - An enterprise-routable external IP range (a /24 or larger is preferred) for NAT, public subnets, and load balancer VIPs. - A non-routable IP range (/23 or larger preferred) assigned to the private-TGW block, for inter-VPC connectivity within a tenant. - For vSphere Supervisor / VCF Automation specifically: N-S Services and Default Outbound NAT (Auto-SNAT) both need to be enabled on the Default VPC Connectivity Profile.

Practical recommendations from the design guide: run Tier-0 in Active/Active to scale out to eight NSX Edge nodes; use a unique private BGP ASN per Tier-0 to avoid loop-detection issues; peer each Edge node with two separate physical devices for redundancy; and for Supervisor’s load-balancing needs specifically, use Large/X-Large Edge appliances or bare-metal NSX Edge rather than the smaller form factors.

Where this fits for the client


Summary

Question Answer
What does “greenfield” mean here? A brand-new workload domain, built from scratch — not an import of an existing environment
What does the Multi-Rack Layer 2 model protect against? Loss of an entire rack, not just a single host — by spreading one cluster’s hosts across multiple racks
Does this model need more VLANs than a single-rack design? No — fewer than a Layer 3 multi-rack design, since each network is one stretched broadcast domain across racks
What’s the minimum cluster size? 3 hosts per zone, to tolerate one host failure — the same floor as a simple management cluster
How many clusters does the diagram show, and why? Three — one Workload Domain built from three vSphere Zones (Cluster 1 = Zone 1, Cluster 2 = Zone 2, Cluster 3 = Zone 3), each combining management and workload roles, all sharing the same four racks
What’s the minimum total host count across all three zones? 9 (3 zones × 3 hosts), before workload-driven sizing adds more
Why three zones instead of one? To prepare for the Three Management Zones with Combined Workload Zones Supervisor model — Supervisor survives the loss of an entire zone because its management components are spread across all three
Do the three zones have to share the same racks? No — two valid scenarios are shown: all three zones sharing the same four racks, or each zone dedicated to its own pair of racks (six racks total). Both give each zone the same logical protection
Is a vSphere Zone a physical, rack-level object? No — it’s a logical availability construct tied to a vSphere cluster. Physical racks determine host placement and failure domains underneath it, but the zone boundary itself is logical
What does “tenant” mean in this model? Any organizational boundary needing network isolation — e.g. HR vs. Finance, production vs. staging, or a business unit — not necessarily a different company
What does the Centralized Connectivity Model share across tenants? One Tier-0 Gateway and its NSX Edge node resources
What stays separate per tenant? Each tenant’s own Centralized Transit Gateway (C-TGW) and VPCs
Minimum NSX Edge nodes required? Two, for high availability
When should a tenant get a dedicated Tier-0 instead? When it needs guaranteed performance, its own VPN, or mandated strict traffic isolation

Chapter 8: Stage 3 Option 2, Greenfield vSphere Supervisor Three Zones

Client: Strategic enterprise, 3 datacenters (West US, East US, West Europe)

Scope of this document: The greenfield vSphere Supervisor design for enabling VKS — Stage 3, Option 2 of the modernization journey. This is the greenfield counterpart to the brownfield single-zone design covered earlier: instead of one combined zone, three vSphere Zones are activated across the three-cluster Workload Domain already built in the previous document. This covers the zone model itself, its cross-zone storage, its networking, and how it all wires together end to end for VKS consumption.

References: Three Management Zones with Combined Workload Zones Model and Centralized Connectivity Model (Broadcom techdocs).


The solution in one paragraph

The previous document built a Workload Domain out of three vSphere Clusters, one per rack pair or spread across shared racks. This document activates vSphere Supervisor across all three of those clusters at once — each cluster becomes a vSphere Zone, and every zone does double duty as both a management zone and a workload zone. The payoff: Supervisor’s own control plane survives the loss of an entire zone, not just a single host, because its three control plane VMs are spread one-per-zone instead of stacked in one place.

Design decisions at a glance

Decision Choice
vSphere Supervisor Management Zones Three
vSphere Supervisor Workload Zones Three — combined with management (not separate)
Zone datastore Cross-zone datastore (a multi-zone datastore)
Activation From the vCenter API, after workload domain creation — a Day 2 operation, not a Day 0 default

That last row matters operationally: the three-zone model is not what a workload domain gets by default. A workload domain defaults to a single zone; three-zone Supervisor has to be explicitly opted into and activated afterward.

The zone model itself

Three Management Zones with Combined Workload Zones Model

Reading the diagram top to bottom:

Storage: the cross-zone datastore

Storage Topology — Cross-Zone Datastore

This is the storage counterpart to the three-zone compute model, and it’s a deliberate departure from the single-zone design:

Why this matters in practice: a zonal datastore ties a namespace’s storage to wherever that one zone happens to be. A cross-zone datastore means a namespace whose workloads span multiple zones (like Namespace A’s VKS Clusters, or Namespace C’s VMs above) isn’t left with storage that only half-covers where its workloads actually run.

Networking: Centralized Connectivity Model

Centralized Connectivity Model with Shared Tier-0 Gateway Per Tenant

This is the same Centralized Connectivity Model covered in full detail in the previous document (shared Tier-0 Gateway, per-tenant Centralized Transit Gateway, VPCs) — it’s referenced here rather than re-explained, because it’s exactly what connects this three-zone Workload Domain to the outside world. The connectivity model doesn’t change when the compute layer goes from one zone to three; the workload domain still presents as one tenant-facing network boundary regardless of which zone a given workload actually lands in.

The next diagram shows precisely how that connectivity model plugs into Supervisor’s own components.

Putting it together: NSX VPC networking model

vSphere Supervisor Networking — NSX VPC Model

NSX VPC Networking is the default Supervisor networking model when Supervisor is activated during workload domain creation. It provides the richest alignment with VCF Automation and the most comprehensive tenant-style networking workflow:

Reading the diagram left to right:

Why this is the default, not just an option: because it’s built on NSX Projects and VPCs from the ground up, this model requires no additional connectivity design work to get tenant-style isolation — creating a new namespace and a new VPC inside its own NSX Project is the isolation boundary. That’s why Broadcom ships it as the out-of-the-box choice when Supervisor is activated at workload domain creation, ahead of the Centralized Connectivity Model covered above, which is the model to reach for when VPC-style consumption isn’t the goal.

Network Requirements for Supervisor with NSX VPC Networking

The diagram above shows the shape of NSX VPC networking; this section gives the concrete numbers — how many IPs, which CIDR sizes, which components — needed to actually build it.

1. Management Network Requirements

Component Minimum Quantity Required Configuration
IPs for Supervisor control plane VMs Block of 5 A block of 5 consecutive DHCP or static IP addresses, assigned from the Management Network to the Supervisor control plane VMs.
Management Network Subnet 1 5 IP addresses for the Supervisor control plane: 1 for each of the 3 nodes, 1 for the virtual IP, 1 for rolling cluster upgrade.
Management traffic network 1 A Management Network that is routable to the ESXi hosts, vCenter, the Supervisor, and the load balancer.
Management Network VLAN 1 The VLAN ID of the Management Network subnet.

2. Workload Network Requirements

Component Minimum Quantity Required Configuration
Kubernetes services CIDR range /16 private IP addresses A private CIDR range to assign IP addresses to Kubernetes services.
Private IP Block — Transit Gateway /16 private IP addresses A private CIDR range for inter-VPC connectivity. This range is not advertised by the Tier-0 Gateway.
Private VPC CIDR range /16 private IP addresses A list of IPv4 CIDRs for IP allocation of private segments. A separate IP pool is created for each VPC, matching its private CIDR range; pools are never shared across Namespaces. A /26 subnet from the range is reserved for the Avi Service Engine.
Load balancer IPs 5 A set of 5 routable IPs for external connectivity: Kube API server, Docker Registry, Kube-DNS, Management Proxy, and vSphere CSI.

3. VPC Network Requirements

Component Minimum Quantity Required Configuration
Tier-0 Uplink IP /24 private IP addresses The IP subnet used for the Tier-0 uplink (IP count depends on Edge configuration — see below).
Centralized Gateway 1 Required by Supervisor, because it provides the Edge nodes needed to configure load balancers.
NSX Project 1 NSX VPCs are created inside projects — the project specified during Supervisor enablement is where the VPC is created when a vSphere Namespace is created.
VPC Connectivity Profile 1 Private Transit Gateway IP Blocks must be configured.

Tier-0 Uplink IP count depends on the Edge configuration: - 1 IP — no Edge redundancy. - 4 IPs — BGP with Edge redundancy (2 IP addresses per Edge). - 3 IPs — static routes with Edge redundancy.

The Edge Management IP/subnet/gateway and the Uplink IP/subnet/gateway must all be unique.

VPC Connectivity Profile, additionally: if the gateway connection type is set to Centralized Gateway, External IP Blocks must be configured and Auto SNAT must be enabled.

Requirements

Universal Supervisor requirements (apply regardless of zone count): - vSphere HA enabled on every cluster. - vSphere DRS enabled, Fully Automated or Partially Automated. - Zone names and Supervisor names must be RFC 1123 compliant (lowercase, alphanumeric, hyphens, starting with a letter or number). - A compatible vSphere Zone must be available at activation time. - Storage policies defined for control plane VMs, ephemeral disks, and the image cache.

Requirements specific to the three-zone model: - Three available vSphere Zones, none of them already assigned to another Supervisor instance — each zone is dedicated, not shared with a different Supervisor deployment. - The High Availability Control Plane model is mandatory here, not optional — three zones need quorum distributed across all three, which only the HA control plane model provides. - Three-zone Supervisor is an opt-out-then-opt-in flow: a workload domain defaults to single-zone, so three-zone activation has to be explicitly chosen and performed as a Day 2 operation from the vCenter API after the workload domain already exists.

Practical recommendations: assign control plane VM IPs statically rather than relying on external DHCP; and if Supervisor is deployed on a vSAN stretched cluster, place control plane VMs on a management network that’s extended across the availability zones involved, so they stay reachable no matter which zone they land in.

Where this fits for the client


Summary

Question Answer
How many zones does this model use, and how does that compare to the brownfield design? Three, each combining management and workload — versus one combined zone in the brownfield single-zone model
Does each zone run its own Control Plane VM? Yes — one per zone, three total, giving Supervisor quorum that survives losing one whole zone
What kind of datastore does this model use? A cross-zone datastore — one logical datastore spanning all three zones, versus a zonal (single-zone) datastore in the brownfield model
Can a namespace’s workloads span more than one zone? Yes — namespaces are not locked to a single zone, unlike the single-zone model
What connectivity model does this use? The Centralized Connectivity Model — the same shared Tier-0 Gateway design covered in the previous document
Is three-zone Supervisor the default? No — a workload domain defaults to single-zone; three-zone activation is a deliberate Day 2 step from the vCenter API
Is the HA control plane model optional here? No — it’s mandatory for the three-zone model, unlike the single-zone model where it’s a choice
What is the default Supervisor networking model? NSX VPC Networking — activated automatically when Supervisor is enabled during workload domain creation, ahead of the Centralized Connectivity Model
How does NSX VPC Networking map tenants to network isolation? One VPC per vSphere Namespace, each inside an NSX Project — the VPC boundary is the tenant/consumption boundary
What separates management traffic from workload traffic end to end? Distinct networks throughout (Management Network vs. each VPC’s Private Subnet / Private Transit Subnet), with Supervisor’s own Default VPC kept separate from tenant VPCs

Chapter 9: Stage 3 Load Balancer Option, Avi Load Balancer

Client: Strategic enterprise, 3 datacenters (West US, East US, West Europe)

Scope of this document: This is a Stage 3 option, not a Stage 4 capability. vSphere Supervisor supports two load balancers, under either supported networking stack, and this document is the solution design for the Avi Load Balancer choice — covering how it deploys, how its traffic flows, and how it fits alongside the brownfield (NSX Segment Networking) and greenfield (NSX VPC Networking) Supervisor designs already built in this project.

Reference: Avi Load Balancer with NSX One to One Workload Domain Deployment in VMware Cloud Foundation (Broadcom techdocs).


Where this decision sits: vSphere Supervisor networking

vSphere Supervisor supports two networking stacks, and both of them support the same two load-balancer choices:

Networking Stack Load Balancer Options
NSX Virtual Private Clouds (VPC) — the greenfield design in this project NSX Load Balancer · Avi Load Balancer
NSX Segment Networking — the brownfield design in this project NSX Load Balancer · Avi Load Balancer

This is why Avi belongs in Stage 3, alongside the zone-model documents, rather than in Stage 4: it’s a load-balancer choice made when Supervisor is activated, independent of whether the workload domain underneath it took the brownfield or greenfield path. The NSX Load Balancer is the simpler, built-in default with no separate control plane to deploy; Avi Load Balancer is the richer option — Layer 7, WAF, GSLB, and deep analytics — and is the subject of the rest of this document.

The solution in one paragraph

Avi is VMware Cloud Foundation’s load balancer for VKS and vSphere Supervisor workloads — it handles Layer 4 and Layer 7 load balancing, web application security, and automatic scaling. It has two parts: Avi Controllers, the control plane, deployed and lifecycle-managed by VCF Operations (which also handles their service accounts, password rotation, and certificate management automatically); and Avi Service Engines (SEs), the data plane, which actually carry traffic and can be deployed into any workload domain cluster. The one design decision that shapes everything else is how many Avi Controller clusters to deploy relative to how many NSX Manager instances exist — one shared cluster, or one dedicated cluster per workload domain.

The core design choice: shared vs. dedicated Avi Controllers

This maps directly to how NSX itself is deployed:

NSX deployment Avi Controller deployment
NSX 1:Many — one NSX Manager instance shared across multiple workload domains One shared Avi Controller cluster serves all of them
NSX 1:1 — a dedicated NSX Manager instance per workload domain A dedicated Avi Controller cluster per NSX instance

Neither is universally “better” — it’s a direct consequence of how NSX was already deployed for each workload domain. Avi follows NSX’s tenancy pattern, it doesn’t set its own.

The 1:1 deployment model

Avi Load Balancer with NSX One-to-One Workload Domain Deployment

This is the fully isolated version of the model, and it’s the one worth designing for by default in a multi-tenant enterprise environment:

Requirements that apply regardless of shared or dedicated: - Avi must be deployed with a 1:1 relationship to its NSX Manager — a single Avi Controller cluster is never split across two different NSX Manager instances. - Avi Load Balancer must be integrated with NSX before Supervisor activation, not after — this is a sequencing requirement, not just a configuration step. - A 3-node Avi Controller cluster is the standard for availability, sized Small, Large, or X-Large depending on scale. - Controllers are deployed in the Management Domain via VCF Operations, with DRS anti-affinity rules automatically configured so the three controller nodes don’t end up on the same host.

Load balancing for vSphere Supervisor: Centralized Transit Gateway

Avi Load Balancer with Centralized Transit Gateway for vSphere Supervisor

This diagram answers a specific, practical question: once Avi is deployed, how does its traffic actually get from a client to a VKS workload, and how do the Avi Service Engines themselves get managed? The answer uses two separate paths that only meet at the Tier-0 Gateway.

The data path (left side, blue): a Client request hits the Physical Network, routes to the Tier-0 Gateway, and from there down to the Centralized Transit Gateway — the same Centralized Connectivity Model construct covered earlier in this project, now carrying Supervisor’s VPC traffic specifically. From the Transit Gateway, traffic reaches the VPC Gateway, which fans out over VPC Subnets to the actual VKS Workloads (VKS cluster nodes and plain VMs). The same VPC Gateway also connects to a VPC Services Subnet, which is how traffic actually reaches the Avi Service Engines doing the load balancing.

The management path (right side, green): the Tier-1 Gateway — a separate branch off the same Tier-0 — connects down to an AVI SE Mgmt Overlay Segment, an NSX overlay segment dedicated purely to management connectivity between the Avi Controllers and the Avi SEs. Broadcom’s design explicitly allows an alternative here too: an NSX VLAN segment can be used for this management connectivity instead of an overlay segment, if that fits the client’s existing network design better.

Why two paths instead of one: this keeps Service Engine data traffic (client requests reaching real workloads) and Service Engine management traffic (the Avi Controllers configuring and monitoring those same Service Engines) on genuinely separate networks. A data-path congestion event doesn’t interfere with the control plane’s ability to manage the Service Engines, and vice versa.

Where Avi fits across the VCF Fleet

VCF Fleet with Avi Load Balancer — AVI Controller and Service Engines Placement

This is the same three-region VCF Fleet from the Fleet Architecture design, with Avi’s two components dropped into place at every site:

Why this repeats identically at all three sites: Avi Controller lifecycle and Service Engine placement aren’t fleet-level services — they don’t get centralized in US WEST the way VCF Operations, VCF Automation, and the License Server are. Each site’s Avi Controller cluster manages that site’s own Service Engines only, which is consistent with the “blast radius” isolation already designed into every other management-plane component in this project.

Requirements and recommendations

Critical requirements: - Enable NSX Cloud when Avi is used as the load balancer for VPC-based Supervisor networking (the NSX VPC model covered in the Greenfield vSphere Supervisor Design). - Avi’s 1:1 relationship with NSX Manager (stated above) is not optional — it’s a hard requirement, not just the recommended pattern. - Avi must be integrated with NSX before Supervisor is activated.

Recommendations: - Deploy the 3-node Avi Controller cluster in the Management Domain, managed through VCF Operations, so certificate/password/service-account lifecycle is automated rather than manual. - Enable VPC mode in the NSX Cloud Account for project discovery — for VPC-based Supervisor deployments, VPC mode is actually on by default when Avi is deployed through VCF Operations. - Enable DHCP addressing for Service Engines so SE IPs are provisioned automatically rather than requiring manual allocation per SE. - In VPC deployments specifically, Service Engines communicate with the VPC over the VPC Services Subnet shown in the diagram — a separate NSX VLAN or overlay segment isn’t needed for that data path, only for SE management.

Where this fits for the client


Summary

Question Answer
What are the two parts of Avi? Avi Controllers (control plane, in the Management Domain) and Avi Service Engines (data plane, in workload domain clusters)
Can one Avi Controller cluster serve two different NSX Manager instances? No — the relationship is always 1:1 with NSX Manager
Who manages Avi Controller certificates and passwords? VCF Operations, automatically, as part of its lifecycle management of the controllers
Can Avi be integrated with NSX after Supervisor is already activated? No — integration must happen before Supervisor activation
What carries client-to-workload data traffic? The Centralized Transit Gateway → VPC Gateway → VPC Subnets / VPC Services Subnet path
What carries Avi Controller-to-Service-Engine management traffic? A separate path: Tier-1 Gateway → AVI SE Mgmt Overlay Segment (or an NSX VLAN segment as an alternative)
Is Avi required from day one of a VKS rollout? No — it’s one of two supported Stage 3 load-balancer options for Supervisor (the other being the built-in NSX Load Balancer); it can be chosen at Supervisor activation or adopted later
Which model fits this client’s three-datacenter footprint? The NSX 1:1 model — one dedicated Avi Controller cluster per workload domain, matching the existing per-domain NSX Manager pattern
Which Supervisor networking stacks support Avi? Both — NSX VPC Networking (greenfield) and NSX Segment Networking (brownfield) support Avi Load Balancer as well as the NSX Load Balancer

Chapter 10: Stage 4, F5 to Avi Load Balancer Migration

Client: Strategic enterprise, 3 datacenters (West US, East US, West Europe)

Scope of this document: The final document in this series, and a Stage 4 design. It answers one question: how does the client retire its existing F5 BIG-IP estate — described in Existing State — External and Global Traffic Management — and replace it with VMware Avi Load Balancer (formerly NSX Advanced Load Balancer), the same Avi platform already designed as a Stage 3 load-balancer option for VKS and vSphere Supervisor in vSphere Supervisor Load Balancer Option: Avi Load Balancer. This document extends Avi beyond Supervisor to the client’s entire application estate — the piece that actually makes Stage 4 a complete private cloud operating model.


1. The solution in one paragraph

F5 does two jobs today: BIG-IP DNS/GTM decides which of the three datacenters serves a user, and BIG-IP LTM load-balances inside the selected datacenter, in front of plain vSphere 7/8 VM workloads. Avi replaces both jobs with its own two-tier architecture: Avi GSLB takes over the global “which site” decision, and Avi Virtual Services running on Avi Service Engines take over the local “which server” decision. Nothing about the shape of the traffic-management problem changes — it is still a global tier sitting above a local tier — only the vendor and the underlying architecture change, from hardware-appliance-based F5 to the software-defined, Controller/Service-Engine model Avi already uses everywhere else in this project.

2. Mapping F5 to Avi

Every F5 concept the client operates today has a direct Avi equivalent. This table is the migration Rosetta Stone — refer back to it throughout the rest of this document:

F5 concept Avi equivalent
F5 BIG-IP DNS / GTM Avi GSLB
F5 BIG-IP LTM Avi Virtual Services on Avi Service Engines
F5 pool Avi Pool
F5 pool members Avi Pool Servers
F5 VIP Avi Virtual Service VIP
F5 health monitor Avi Health Monitor
F5 SSL profile Avi SSL Profile / Certificate
F5 iRules Avi HTTP policies / DataScripts / WAF policies (use-case dependent)
F5 partitions / tenants Avi Tenants / VCF Automation projects / NSX Projects
F5 analytics / logs Avi Analytics + VCF Operations integration

Do not treat this as a one-to-one automatic conversion. A few F5 features need real design decisions rather than a straight lookup — see Section 7, Migration cautions.

3. Target-state traffic flow

Target State — Request Flow: Avi GSLB to Avi Service Engines

This is the direct replacement for the flow in the existing-state document, hop for hop:

  1. DNS Query — a user’s device resolves the application name against Avi GSLB, which takes over DNS delegation from F5 BIG-IP DNS/GTM.
  2. Best Site Selected — Avi GSLB evaluates site health and applies a GSLB policy (active/active, active/passive, geo-based, priority-based, or ratio-based) and returns the best region’s Virtual Service — the same job GTM did, just on Avi.
  3. VIP / Service Engines — the regional Avi Virtual Service, running on distributed Avi Service Engines, terminates the connection, applies SSL, checks pool member health, and load-balances to the application. The application itself doesn’t have to change on day one: it can still be a plain vSphere VM, exactly as it is today, and later become a VCF or VKS workload without another redesign of this flow.

Where the Avi Controller fits: it is deliberately not drawn as a traffic hop. The Controller cluster is the control plane — it configures Virtual Services, runs health monitors, and pushes configuration to Service Engines — but it never sits in the path of a live request. Only the Service Engines carry traffic, which is exactly the control-plane/data-plane split VMware’s Avi architecture is built around.

4. Component design per region

Avi Component Design — Three-Region Deployment

Each of the three regions is built identically — there is no “primary” Avi site the way F5 GTM sometimes centralizes configuration. Per region:

Component Purpose
Avi Controller cluster (3-node: 1 leader, 2 followers) The region’s control plane — Virtual Services, pools, policies, certificates, health monitors, and analytics
Service Engine Group — Production Resource pool for Service Engines carrying production traffic
Service Engine Group — Non-Production A separate resource pool for dev/test/staging, isolated from production capacity
Service Engine Group — VKS / Kubernetes A dedicated resource pool for Service Engines fronting VKS ingress traffic via the Avi Kubernetes Operator (AKO)
VIP Network Where application Virtual Service addresses live — the client-facing side
Backend Network Where the app VMs, VKS nodes, or services actually live — the same north-south split F5 LTM already used
Integration: vCenter · NSX · VCF Operations The Controller cluster’s connections into the rest of the platform

Above all three regions sits Avi GSLB, configured with a single GSLB Service (for example app.company.com) whose pool members are the three regions’ Virtual Services. GSLB only ever talks to each region’s Controller cluster for health and configuration — it never touches application data traffic directly.

Separating Service Engines into Production, Non-Production, and VKS/Kubernetes groups mirrors the same isolation principle used everywhere else in this project (the “blast radius” thinking applied to NSX Manager, vCenter, and Avi Controllers in the earlier documents): a capacity problem or a bad deployment in one group doesn’t starve the others.

5. Migration approach: seven phases

F5 and Avi run in parallel for the entire migration — there is no big-bang cutover. F5 stays in production until every application has been individually validated on Avi and moved over.

Phase What happens
1. Discover and document F5 Build a full inventory of the current F5 configuration — GTM Wide IPs, GTM pools, LTM VIPs, pools, health monitors, SSL certificates, iRules, SNAT/NAT behavior, persistence settings, and any WAF/ASM policies. This inventory drives every later phase.
2. Build the Avi foundation Deploy Avi Controllers, cloud integration (vCenter/VCF, NSX if applicable), Service Engine Groups, VIP networks, IPAM/DNS, certificates, and VCF Operations monitoring — all built alongside F5, with zero production traffic on Avi yet.
3. Recreate local LTM services in Avi For every F5 LTM VIP, build the equivalent Avi Virtual Service — same pool members, same health monitor type, same SSL profile — and validate it under a temporary DNS name before it’s live.
4. Recreate global GTM services in Avi GSLB For every F5 GTM Wide IP, build the equivalent Avi GSLB Service with the same regional endpoints, then choose the right traffic policy per application (active/active, active/passive, geo, priority, or ratio-based).
5. Pilot migration Take one low-risk application through the full flow — build, test backend health, test SSL, test persistence, test application behavior, add it to GSLB, lower DNS TTL, shift a small percentage of traffic, monitor, then cut over fully. This pilot becomes the template for every application after it.
6. Production cutover, per application Repeat the pilot pattern for each remaining application: reduce DNS TTL, confirm Avi health checks are green, verify certificates, source-IP behavior, persistence, firewall rules, and logging — and keep a rollback path back to F5 available until the application is proven stable on Avi.
7. Decommission F5, by wave Retire F5 in waves, not all at once: non-critical internal apps first, then regional apps, then customer-facing apps, then critical apps with DR/GSLB dependencies, then whatever legacy/special cases are left. Don’t decommission F5 for an application until every dependency — DNS delegation, certificates, monitoring, firewall rules, NAT, iRules, WAF/ASM policies — has actually moved.

6. Why replace F5 with Avi

Reason Benefit
Software-defined architecture Removes dependency on hardware ADC appliances
VCF integration Native alignment with VCF 9.1, VCF Operations, VCF Automation, NSX, and VKS
Centralized analytics Better visibility into app health, latency, errors, and traffic behavior
Elastic scale-out Service Engines deploy and scale with demand instead of fixed hardware capacity
Kubernetes integration Native VKS ingress via the Avi Kubernetes Operator (AKO) and Gateway API support
Multi-tenancy Maps cleanly onto VCF Automation projects, NSX Projects, and Avi Tenants
Automation Simpler day-0 to day-2 lifecycle and application onboarding
GSLB replacement Directly replaces F5 BIG-IP DNS/GTM for global traffic management
Local ADC replacement Directly replaces F5 BIG-IP LTM for VIPs, pools, SSL, persistence, and health checks

VCF 9.1 recommends Avi specifically for environments that need advanced ADC capability — GSLB, deep traffic analytics, complex traffic management, and application security (WAF) — which describes this client’s current F5 use exactly. For VKS workloads, Avi is the recommended ingress path via AKO whenever an application needs Layer 7 balancing, WAF, or DNS integration beyond what a basic in-cluster ingress controller provides.

7. Migration cautions

A handful of F5 features don’t translate automatically and need explicit validation during Phase 1 and Phase 3:

F5 feature What to watch for
iRules Translate to Avi HTTP policies or DataScripts — validate behavior doesn’t silently change
ASM / Advanced WAF Map to Avi WAF policies; don’t assume rule-for-rule equivalence
Complex persistence Validate the Avi persistence profile actually reproduces the F5 behavior under test
SNAT behavior Confirm source-IP requirements are preserved for apps that depend on real client IPs
SSL profiles Validate cipher suites, certificate chains, and client-certificate auth match
TCP profiles Validate timeout and protocol behavior, especially for long-lived connections
GTM topology rules Rebuild as Avi GSLB policies — geo/topology logic needs to be re-verified, not assumed
Monitoring integrations Repoint any tooling that watches F5 metrics/logs to Avi Analytics and VCF Operations
DNS delegation Move DNS authority from F5 GTM to Avi GSLB carefully, application by application, not all at once

8. Where this fits for the client


Summary

Question Answer
What replaces F5 BIG-IP DNS/GTM? Avi GSLB
What replaces F5 BIG-IP LTM? Avi Virtual Services running on Avi Service Engines
Does the client cut over all at once? No — F5 and Avi run in parallel through a seven-phase, wave-based migration
Do all F5 features map automatically? No — iRules, WAF/ASM, persistence, SNAT, SSL, TCP profiles, and GTM topology rules need explicit validation
Is there one Avi Controller cluster for the whole fleet? No — each region gets its own 3-node Avi Controller cluster and Service Engine Groups, matching the existing three-datacenter boundaries
How are Service Engines separated? Into Production, Non-Production, and VKS/Kubernetes groups per region
Does this depend on VCF or VKS being live first? No — Avi can front plain vSphere VMs today and VCF/VKS workloads later without changing this design
What’s the end state? One load-balancing platform (Avi) across VM, VCF, and VKS workloads, replacing F5 across all three sites

Document Map — How the Chapters Fit Together

Stage Name Chapters
1 Traditional Environment (today) Chapter 1, Chapter 2
— Roadmap overview Chapter 3
2 Platform Foundation Modernization Chapter 4 (fleet), Chapter 5 (brownfield), Chapter 7 (greenfield)
3 VKS Adoption Chapter 6 (brownfield), Chapter 8 (greenfield), Chapter 9 (load balancer choice)
4 Full Private Cloud Operating Model Chapter 10

Each stage is a superset of the one before it: moving forward never requires rebuilding what came before, and the client’s three sites do not need to progress through the stages in lockstep.


Reference Table

Every external URL cited across this document’s source material, consolidated in one place. All 23 links were checked on 2026-07-08 and are live. Two Dell links return an automated-client block (HTTP 403) to a plain, headerless request — this is Dell’s bot protection, not a broken link; both were independently confirmed reachable and correct via a full browser-style fetch.

Design library and how-to documentation (Broadcom TechDocs)

# Reference Description Cited in URL
1 Import an Existing vCenter to Create a Workload Domain The mechanism behind the brownfield path — bringing an existing vCenter, its clusters, hosts, storage, and NSX Manager into VCF 9.1 management as-is Chapter 5 https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/building-your-private-cloud-infrastructure/working-with-workload-domains/import-an-existing-vcenter-to-create-a-workload-domain.html
2 VCF Fleet with Multiple Sites Across Multiple Regions (blueprint) The reference blueprint for running one VCF Fleet — three VCF Instances, one per region — under shared fleet-level management Chapter 4 https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/design/design-blueprints-for/infrastructure-modernization/vcf-fleet-with-multiple-sites-across-multiple-regions-blueprint.html
3 Multi-Rack Layer 2 vSphere Cluster Model Compute/fault-domain design for a vSphere cluster stretched across multiple racks within one availability zone Chapter 7 https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/design/design-library/cluster-models/multi-rack-cluster-detailed-design/layer-2-multi-rack-cluster.html
4 Centralized Connectivity Model with Shared Tier-0 Gateway Per Tenant Networking design where multiple tenants share one Tier-0 Gateway and NSX Edge cluster while keeping logical isolation via per-tenant Centralized Transit Gateways and VPCs Chapter 7, Chapter 8 https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/design/design-library/workload-connectivity-designs/centralized-connectivity-model.html
5 Single Management Zone with Combined Workload Zones Model The brownfield Supervisor zone model — one vSphere Zone serving both Supervisor management and application workloads Chapter 6 https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/design/design-library/self-service-iaas-deployment-models/vsphere-supervisor-zone-models/single-zone.html
6 Three Management Zones with Combined Workload Zones Model The greenfield Supervisor zone model — three vSphere Zones, each combining management and workload roles, giving Supervisor zone-level fault tolerance Chapter 8 https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/design/design-library/self-service-iaas-deployment-models/vsphere-supervisor-zone-models/three-management-zones-with-combined-workload-zones-model.html
7 NSX Segment Connectivity Model Overview Networking model used by the brownfield Supervisor design — a shared Tier-0 Gateway with an auto-created Tier-1 Gateway and load balancer per namespace Chapter 6 https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/design/design-library/workload-connectivity-designs/nsx-segment-network-connectivity-model.html
8 Avi Load Balancer Detailed Design (NSX 1:1 Workload Domain Deployment) Design guidance for deploying Avi Controllers and Service Engines, including the 1:1 Avi-to-NSX-Manager pattern used across this client’s three sites Chapter 9 https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/design/design-library/vcf-load-balancing-detailed-design/avi-load-balancer-detailed-design.html

Hardware compatibility references

# Reference Description Cited in URL
9 Broadcom Compatibility Guide (VCF / vSphere hardware search) The authoritative source for VCF 9.1 / ESXi 9.1 server, storage controller, NIC, and firmware compatibility Chapter 1 https://compatibilityguide.broadcom.com/search?program=server&persona=live&column=partnerName&order=asc
10 Dell ESXi 9.x Compatibility Matrix — yx5x / 2021 generation (R750, MX750c) Confirms the 2021-era Dell hardware estimate (R750 rack / MX750c blade) supports ESXi 9.0 and 9.1 Chapter 1 https://www.dell.com/support/manuals/en-lk/vmware-esxi-9-x/vmware_9.x_compatibility_matrix_pub/dell-yx5x-poweredge-systems?guid=guid-19353ff5-4005-40b6-96a6-93816d58a614&lang=en-us
11 Dell ESXi 9.x Compatibility Matrix — yx6x / 2023 generation (R760, MX760c) Confirms the 2023-era Dell hardware estimate (R760 rack / MX760c blade) supports ESXi 9.0 and 9.1 Chapter 1 https://www.dell.com/support/manuals/en-us/vmware-esxi-9-x/vmware_9.x_compatibility_matrix_pub/intel-sapphire-rapids-processor-and-raptor-lake-processor?guid=guid-bcf73f99-5539-4476-91ab-c56f57e49561&lang=en-us

Source reference diagrams (Broadcom original assets)

These are the original Broadcom PNG/SVG master images each corresponding draw.io diagram in this document was recreated from — kept here for provenance and future re-verification against the source.

# Reference Description Cited in URL
12 VKS Adoption Journey (source image) Master reference for diagrams/vks-adoption-journey.drawio.svg Chapter 3 https://techdocs.broadcom.com/content/broadcom/techdocs/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/_jcr_content/assetversioncopies/933c3a9a-0068-4c58-be91-2dc78d892264.original.png
13 VCF Fleet with Multiple Sites — Architecture Overview (source image) Master reference for diagrams/vcf-fleet-architecture-overview.drawio.svg Chapter 4 https://techdocs.broadcom.com/content/broadcom/techdocs/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/_jcr_content/assetversioncopies/eb3bca7c-61cc-4620-85f5-590333ce6438.original.svg
14 Single Management Zone with Combined Workload Zones Model (source image) Master reference for diagrams/vsphere-supervisor-single-zone-model.drawio.svg Chapter 6 https://techdocs.broadcom.com/content/broadcom/techdocs/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/_jcr_content/assetversioncopies/d1b97310-bfaa-4015-b129-2dce88c45714.original.png
15 Storage Topology — Zonal Datastore (source image) Master reference for diagrams/vsphere-zonal-datastore.drawio.svg Chapter 6 https://techdocs.broadcom.com/content/broadcom/techdocs/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/_jcr_content/assetversioncopies/18793ba3-1165-4fa6-a43f-bd46afb5df91.original.png
16 NSX Segment Connectivity Model Overview (source image) Master reference for diagrams/vsphere-supervisor-nsx-segment-networking.drawio.svg Chapter 6 https://techdocs.broadcom.com/content/broadcom/techdocs/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/_jcr_content/assetversioncopies/db2c9f8c-dbe8-44b3-a84e-01bf26d93309.original.png
17 Multi-Rack Layer 2 vSphere Cluster Model (source image) Master reference for diagrams/multi-rack-layer2-cluster-model.drawio.svg Chapter 7 https://techdocs.broadcom.com/content/broadcom/techdocs/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/_jcr_content/assetversioncopies/256720f8-287e-4983-805a-71e117dce746.original.svg
18 Centralized Connectivity Model with Shared Tier-0 Gateway Per Tenant (source image) Master reference for diagrams/centralized-connectivity-shared-tier0.drawio.svg Chapter 7, Chapter 8 https://techdocs.broadcom.com/content/broadcom/techdocs/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/_jcr_content/assetversioncopies/69b74bba-f545-4acd-bc8b-201a3c1447c4.original.svg
19 Three Management Zones with Combined Workload Zones Model (source image) Master reference for diagrams/three-management-zones-model.drawio.svg Chapter 8 https://techdocs.broadcom.com/content/broadcom/techdocs/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/_jcr_content/assetversioncopies/af407163-1605-414c-82f3-40fe56ee206e.original.png
20 Storage Topology — Cross-Zone Datastore (source image) Master reference for diagrams/cross-zone-datastore.drawio.svg Chapter 8 https://techdocs.broadcom.com/content/broadcom/techdocs/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/_jcr_content/assetversioncopies/d661fd40-8b91-4e11-8159-9bbd5e26e72c.original.png
21 vSphere Supervisor Deployment Model Using VPC-Based Networking (source image) Master reference for diagrams/vsphere-supervisor-nsx-vpc-model.drawio.svg Chapter 8 https://techdocs.broadcom.com/content/broadcom/techdocs/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/_jcr_content/assetversioncopies/778aa927-b1fa-480c-94cf-0a7757235171.original.png
22 Avi Load Balancer with NSX One-to-One Workload Domain Deployment (source image) Master reference for diagrams/avi-nsx-one-to-one-deployment.drawio.svg Chapter 9 https://techdocs.broadcom.com/content/broadcom/techdocs/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/_jcr_content/assetversioncopies/b3c19dfb-2a5a-4779-9426-e2c3135e59dd.original.svg
23 Avi Load Balancer with Centralized Transit Gateway for vSphere Supervisor (source image) Master reference for diagrams/avi-centralized-transit-gateway-supervisor.drawio.svg Chapter 9 https://techdocs.broadcom.com/content/broadcom/techdocs/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1/_jcr_content/assetversioncopies/ffd0dd34-ebbc-4a70-9f24-f1fe03cc25f4.original.svg