CloudForge
Published 15 min read

Argo CD vs Flux: An Enterprise Guide to Multi-Cluster GitOps

Compare Argo CD and Flux for enterprise multi-cluster GitOps: topology patterns, reference architecture, governance controls, operating model and adoption roadmap.

Argo CD and Flux are both mature, CNCF graduated GitOps controllers. Either can reconcile a Kubernetes estate from Git or an OCI registry. For an enterprise running clusters across environments, regions and compliance boundaries, the harder decision is not which tool has the longer feature list. It is where reconciliation runs, which side opens the network connection, and where cluster credentials are stored.

This guide is written for platform leads, architects, CTOs and security reviewers. It compares Argo CD 3.5 and Flux 2.9 as of October 2026, then sets out a reference architecture, governance controls, an operating model and a phased adoption roadmap that apply to either tool.

Executive summary

  • Both tools are production-grade. The functional gap between them narrows with each release. Choose on operating model and topology before features.
  • Topology matters more than the tool. Decide per trust zone whether a central control plane pushes to clusters or each cluster pulls its own state.
  • Choose Argo CD when many product teams need a shared web UI, SSO-mapped permissions and fleet templating from a central control plane.
  • Choose Flux when clusters must pull their own state with no inbound access and no central credential store, or when GitOps should be installed with the cluster as plain Kubernetes objects.
  • Never share a control plane across trust zones. Separate production from non-production, and give regulated environments their own.
If your estate looks like thisRecommended patternWhy
A few network zones, many product teams, dozens of clustersArgo CD: one highly available instance per environment tier, with ApplicationSetsShared UI and SSO permissions; push is acceptable inside one trust zone
Regulated workloads, private or edge clusters, no inbound access allowedFlux in each cluster, or the Argo CD agent where a central UI is mandatoryNo central credential store; failures stay local
Clusters created by Cluster API or Terraform at scaleFlux installed during provisioning; hub and spoke for the platform layerGitOps arrives with the cluster as Kubernetes objects
A platform team serving many product teamsFlux for platform add-ons, Argo CD for applications, with strict ownershipEach tool covers the layer it suits best

1. How GitOps reconciliation works

GitOps replaces imperative deploy scripts with a controller that continuously compares two things: the desired state in Git or an OCI registry, and the live state in the cluster. When they differ, the controller applies the difference. When someone changes the cluster by hand, the controller detects the drift and restores the approved state.

For an enterprise this matters for two reasons. The Git history becomes the change record, with an author, reviewer, timestamp and diff for every production change. And unapproved changes do not survive the next reconcile.

Figure 01 Reconcile loopGit is the source of truth. The controller keeps the cluster honest.
Continuous reconciliation
1 Merge to Git2 Controller fetches3 Diff desired vs live4 Apply and pruneevery interval or on webhookIn sync
deployment/checkout · prod-us-east
Desired (Git)
replicas: 3
Live (cluster)
replicas: 3
Image
checkout@sha256:9f2c…

Synced and healthy

Illustrative example. A manual scale-up creates drift; self-heal (Argo CD selfHeal, or Flux reconciliation) restores the approved state from Git.

Self-heal is a policy decision. Argo CD reverts drift when selfHeal is enabled on an Application; Flux reverts it on its regular reconciliation interval. Both need a documented break-glass path for incidents, covered in the governance section below.

2. Argo CD and Flux in 2026

Argo CD 3.5: an application delivery platform

Argo CD runs an API server, web UI, repository server and application controller. The Application is the unit of delivery, and ApplicationSet generates Applications from cluster inventories, Git directories or pull requests.

Version 3.5, released as a candidate in June 2026, added mutual TLS between internal components, Git commit signature verification through Source Integrity, ApplicationSet management in the UI, Helm 4 support and ApplicationSets in any namespace. It also promoted impersonation and the Source Hydrator to beta. Version 3.4 added a way to register a cluster without the central controller reconciling it, which supports estates that mix a central hub with agents.

Flux 2.9: a toolkit of composable controllers

Flux is a set of controllers for sources, Kustomize, Helm, notifications and image automation, configured entirely through Kubernetes custom resources. There is no required central server.

OCI artifact support reached general availability in Flux 2.6 and image update automation in Flux 2.7. Flux 2.9, released on June 30, 2026, added a CLI plugin system and field-level ignore rules for server-side apply drift detection. The Flux Operator, maintained by ControlPlane, adds a web UI, sharding and multi-tenancy lockdown, and is available through Red Hat OperatorHub.

3. The four multi-cluster patterns

Every multi-cluster design answers three questions: which side opens the network connection, where cluster credentials are stored, and what happens to the fleet when the central cluster fails.

Figure 02 Multi-cluster patternsWho opens the connection, and who holds the keys
Git / OCIdesired stateHub clusterArgo CD · UI, API, controllerkubeconfig × 3one per spokeprod-us-eastWorkloads onlyprod-eu-westWorkloads onlyedge-store-114Workloads only
Who connects
Argo CD on the hub opens a connection to every spoke's Kubernetes API.
Credentials
The hub stores a credential for every cluster.
If the hub fails
No cluster receives new syncs. Running workloads keep running.
Best for
One trust zone with many product teams that want a single UI.
Watch for
Cross-network API traffic and egress cost; the hub becomes a high-value target.
Argo CD connectionFlux connectionFetch from Git or OCIkubeconfig = stored cluster credential

Arrows point in the direction the connection is opened. The table below summarizes all four patterns.

PatternConnectionCredentialsIf the central cluster failsBest fit
Argo CD hub pushHub opens a connection to each spoke's Kubernetes APIHub stores one per clusterNo syncs anywhere; workloads keep runningOne trust zone, many product teams, one UI
Argo CD agent pullAgents on spokes connect out to the hub over gRPC with client certificatesNone stored on the hubSpokes keep current state; UI and new configuration pauseFirewalled or edge clusters that still need a central UI
Flux per-cluster pullEach cluster pulls from Git or OCI; nothing connects inboundEach cluster holds only its own read credentialNo hub to fail; impact stays in one clusterRegulated, edge and large uniform fleets
Flux hub and spokeA management cluster applies through kubeconfig secretsHub stores one per clusterNo reconciliation until the hub returnsCluster API estates with a management cluster

Argo CD hub push and Flux hub and spoke share the same risk profile: one cluster holds credentials for the fleet. The difference is operational. Argo CD provides a single UI and permission model; Flux provides plain Kubernetes objects that provisioning tools can create with the cluster, and a hub that can be sharded across several Flux instances.

The agent pattern is Argo CD's answer to the pull model. The argocd-agent project inverts the connection so that spokes connect out to a principal on the hub, authenticate with client certificates and reconcile locally. Its documentation lists high availability for the principal and full multi-tenancy as work in progress, so treat it as maturing for production use.

4. Enterprise reference architecture

The architecture below applies to either tool. It separates control planes by trust zone, keeps a single source of truth with a signed supply chain, and treats regulated environments as pull-only.

Figure 03 Reference architectureOne source of truth. A control plane per trust zone.
SOURCE OF TRUTHNON-PRODUCTION ZONEPRODUCTION ZONEREGULATED ZONEApp repositoriesproduct teamsPlatform config repoclusters/ platform/ apps/Policy repoKyverno / GatekeeperCI pipelinebuild · test · sign · SBOMOCI registrysigned images + configsArgo CD (HA)ApplicationSets · SSO rolesdev1 clusterstaging2 regionsControl planes are not sharedacross zones. Each is rebuiltfrom Terraform and Git.Argo CD (HA)AppProjects · sync windowsprod-aus-eastprod-beu-westPRNo standing human write access.Break-glass is logged andreconciled back into Git.payments-pciisolated account or VPCargocd-agentor a local Argo CDAgent connects out to the hubOwn audit scopeSHARED PLATFORM SERVICESSSO / OIDCSecrets: Vault or cloud KMSAdmission policyMetrics, logs, SIEM
Fetch desired stateReconcile to clustersPromotion by pull requestDashed boxes mark trust boundaries

The same architecture applies to either tool. Switch the tool to see which components fill each role.

Design principles

  1. Separate control planes per trust zone. Production, non-production and regulated environments never share a GitOps control plane.
  2. Pull across untrusted networks. Use hub push only inside one trust zone. Across zones, clusters pull.
  3. Make Git the change record. No routine kubectl in production. Exceptions follow a documented break-glass procedure.
  4. Promote immutable artifacts. Production references image digests and versioned configuration, never a moving branch or a latest tag.
  5. Split platform and application ownership. Platform add-ons and product workloads live in separate paths with separate owners.
  6. Enforce policy at admission. CI checks are advisory. Kyverno or OPA Gatekeeper in the cluster is the control.
  7. Keep secrets out of Git. Reference external stores through External Secrets Operator, or encrypt with SOPS.
  8. Manage the GitOps platform as code. Bootstrap it with Terraform, upgrade it through Git, and rehearse rebuilding it.

The supply chain controls behind principle 4 are covered in the DevSecOps pipeline security guide, and the platform boundaries behind principle 5 in the internal developer platform strategy guide.

5. Repository structure and promotion

Use one branch with a folder per environment, not a branch per environment. Environment branches drift apart through cherry-picks and make it hard to see what is running where. A folder layout makes every environment's state reviewable on main, and CODEOWNERS assigns approval rights per path.

platform-config/
├── clusters/                  # what runs where, one folder per cluster
│   ├── nonprod/dev-use1/
│   ├── nonprod/stg-use1/
│   └── prod/prd-use1/  prd-euw1/
├── platform/                  # add-ons owned by the platform team
│   ├── base/                  # ingress, cert-manager, external-secrets, kyverno
│   └── overlays/nonprod/  prod/
└── apps/                      # product teams own their folders via CODEOWNERS
    └── payments/
        ├── base/
        └── overlays/dev/  staging/  prod/   # prod pins image digests

Promotion is a pull request that copies a tested digest from the staging overlay to the production overlay. CI can open that pull request automatically after staging passes; a person with production approval rights merges it. Rollout across production regions then proceeds in waves, and an unhealthy wave stops the release.

Figure 04 Progressive promotionRelease in waves. Stop at the first unhealthy one.

Dev

ApplicationSet step 1
  • dev-1v2.14.1
Next step waits for Healthy

Staging

RollingSync step 2
  • stg-usv2.14.1
  • stg-euv2.14.1
Next step waits for Healthy

Production

RollingSync step 3 · maxUpdate 25%
  • prod-us-eastv2.14.1
  • prod-us-westv2.14.1
  • prod-eu-westv2.14.1
  • prod-ap-southv2.14.1

Every cluster is running v2.14.1. Start the release to roll out v2.15.0.

Illustrative rollout of one service across seven clusters. Argo CD uses ApplicationSet progressive syncs; Flux chains Kustomizations with dependsOn and wait: true.

Argo CD expresses waves with ApplicationSet progressive syncs. Each RollingSync step waits for the previous step's Applications to report Healthy, and maxUpdate limits how many production clusters change at once. Progressive syncs must be enabled on the ApplicationSet controller.

apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
  name: payments
  namespace: argocd
spec:
  goTemplate: true
  generators:
    - clusters:
        selector:
          matchLabels:
            fleet: payments
  strategy:
    type: RollingSync
    rollingSync:
      steps:
        - matchExpressions:
            - {key: env, operator: In, values: [staging]}
        - matchExpressions:
            - {key: env, operator: In, values: [prod]}
          maxUpdate: 25%        # one production cluster in four at a time
  template:
    metadata:
      name: 'payments-{{.name}}'
      labels:
        env: '{{index .metadata.labels "env"}}'
    spec:
      project: payments         # AppProject limits repositories and destinations
      source:
        repoURL: https://github.com/acme/platform-config
        targetRevision: main
        path: 'apps/payments/overlays/{{index .metadata.labels "env"}}'
      destination:
        server: '{{.server}}'
        namespace: payments

Flux expresses the same ordering with dependsOn between Kustomizations. Setting wait: true means a Kustomization reports Ready only when its workloads are healthy, so the next wave starts only after the previous one succeeds. In hub mode, kubeConfig points at the target cluster; in per-cluster mode it is omitted.

apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
  name: payments-prod-us-east
  namespace: flux-system
spec:
  interval: 10m
  retryInterval: 2m
  serviceAccountName: payments-deployer   # tenant RBAC
  sourceRef:
    kind: GitRepository
    name: platform-config
  path: ./apps/payments/overlays/prod
  prune: true
  wait: true
  dependsOn:
    - name: payments-staging
  kubeConfig:                             # hub mode only
    secretRef:
      name: prod-us-east-kubeconfig
  postBuild:
    substitute:
      cluster_name: prod-us-east
      region: us-east-1

For canary traffic shifting inside a single cluster, the usual companions are Argo Rollouts and Flagger. The enterprise deployment bottlenecks guide covers the approval and release practices around this flow.

6. Governance, security and compliance

GitOps gives auditors stronger evidence than pipeline-pushed deployments. Every production change is a reviewed pull request with an approver, a timestamp and a diff, and the controller records what it applied. The table maps common control objectives to the mechanism in each tool. Map the rows to your own framework, for example SOC 2 change management (CC8.1), ISO/IEC 27001:2022 control 8.32 or PCI DSS requirement 6.5.

Control objectiveGitOps mechanismArgo CDFlux
Change approvalBranch protection, required reviews, CODEOWNERS per pathEnforced in the Git platformEnforced in the Git platform
Separation of dutiesAuthors cannot approve their own production change; no standing human write access to production clustersAppProjects restrict repositories, clusters and namespaces; SSO groups map to rolesPer-tenant service accounts through serviceAccountName; Kubernetes RBAC
Integrity of what is deployedSigned commits and artifacts, verified before applySource Integrity commit verification (3.5)GitRepository signature verification; cosign verification for OCI
Audit trailGit history plus controller events forwarded to the SIEMNotifications controller and API audit logsNotification controller alerts and events
Drift detectionUnapproved changes detected and revertedOutOfSync status and selfHealContinuous reconciliation; SSA field ignore rules (2.9)
Secrets handlingNo plaintext secrets in GitExternal Secrets Operator or a secrets pluginNative SOPS decryption, or External Secrets Operator
Break-glassDocumented emergency path, then reconcile the fix back into GitDisable automated sync for the Application; skip reconciliation for a cluster (3.4 and later)flux suspend the Kustomization, then resume
Control plane hardeningThe GitOps system is itself a privileged workloadInternal mTLS (3.5), SSO, HA installationMulti-tenancy lockdown and network policies through Flux Operator

Credential concentration is the largest security difference between patterns. A hub that stores a kubeconfig for every cluster is one breach away from the whole fleet. Per-cluster Flux and the Argo CD agent keep credentials local; prefer them for regulated or internet-exposed fleets. Our DevSecOps consulting team can map these controls to your audit scope.

7. Operating model

Most GitOps programs that stall do so on ownership rather than technology. Agree who is responsible (R), accountable (A), consulted (C) and informed (I) for each activity before the pilot starts.

ActivityPlatform teamProduct teamsSecuritySRE / on-call
Run, upgrade and back up the GitOps control planeA/RICC
Templates, golden paths and cluster onboardingA/RCCI
Application manifests and promotion requestsCA/RII
Admission policy definition and exceptionsRCAI
Triage sync failures and drift alertsCRIA/R
Approve break-glass and reconcile afterwardCRIA

The platform team's role here mirrors the product-team model described in platform engineering golden paths: templates and paved paths are the product, and product teams are the customers.

8. Adoption roadmap

A phased rollout proves the model with willing teams before it becomes mandatory. Durations are indicative for an estate of a few dozen clusters, and each phase ends only when its exit criteria are met.

Figure 05 Adoption roadmapProve it with willing teams, then make it the standard
  1. Phase 0 · Weeks 1–2

    Assess

    • Inventory clusters, network zones, teams and compliance scope
    • Choose a pattern per trust zone
    • Baseline DORA metrics

    Exit criteriaAn approved architecture decision record for each trust zone

  2. Phase 1 · Weeks 3–8

    Foundation

    • Bootstrap control planes with Terraform
    • Repository layout, CODEOWNERS, SSO and RBAC
    • Secrets integration; admission policy in audit mode
    • Controller metrics and alerts

    Exit criteriaNon-production control plane rebuilt from Git in a test; platform add-ons delivered by GitOps

  3. Phase 2 · Weeks 9–12

    Pilot

    • One or two willing product teams
    • Promotion by pull request from dev to production
    • Break-glass drill

    Exit criteriaPilot services reach production only through Git for four consecutive weeks

  4. Phase 3 · Weeks 13–24

    Scale

    • Self-service onboarding template
    • Fleet templating and progressive rollout across regions
    • Admission policy switched to enforce
    • Remove standing human write access to production

    Exit criteriaAgreed share of services onboarded; no standing cluster-admin in production

  5. Phase 4 · Ongoing

    Operate

    • Quarterly tool upgrades through Git, non-production first
    • Control plane rebuild game days
    • Monthly metrics review

    Exit criteriaUpgrades and rebuilds are routine work, not projects

Indicative timeline for an estate of a few dozen clusters. Adjust durations to your fleet, team capacity and audit calendar.

9. Scale and resilience of the control plane

The GitOps control plane becomes tier-0 infrastructure. Design and operate it accordingly.

  • Argo CD scaling. Use the HA installation, shard the application controller across clusters and scale repository servers for manifest rendering. Argo CD 3.5 also adds concurrency to how the ApplicationSet controller manages Applications.
  • Flux scaling. Per-cluster Flux scales with the fleet by design. A Flux hub can be sharded across several instances, and the Flux Operator configures sharding and vertical scaling.
  • Failure behavior. When the control plane is down, running workloads keep running; only new changes and drift correction stop. State this explicitly in your recovery objectives.
  • Disaster recovery. The desired state is already in Git. Back up what is not: cluster credentials, SSO configuration and secret references. Rehearse a full control plane rebuild from Terraform and Git at least twice a year. The cloud disaster recovery guide covers setting RTO and RPO for tier-0 services.
  • Upgrades. Read the upgrade notes for every minor version. For example, Argo CD 3.3 required the ServerSideApply=true sync option on the Application that manages Argo CD itself. Roll upgrades through non-production first.

10. Measuring success

Report outcomes leadership already tracks, plus a few GitOps-specific indicators taken from the controllers' Prometheus metrics and your Git platform. The DevOps and SRE operating model guide explains consistent DORA definitions.

MetricDefinitionDirection to aim for
Lead time for changesPull request merged to running in productionFalling
Deployment frequencyProduction promotions per service per weekRising or stable
Change failure rateShare of promotions that cause a rollback or incidentFalling
Time to restoreIncident start to service restored, including Git revertsFalling
Changes delivered through GitProduction changes made by the controller rather than by people100% outside break-glass
Drift eventsOut-of-sync events caused by manual changesNear zero
Reconcile success rateSuccessful syncs over total, per clusterStable; investigate dips

11. Anti-patterns and how to avoid them

Anti-patternWhy it hurtsDo this instead
A branch per environmentBranches drift through cherry-picks; nobody can see what runs whereOne branch, a folder per environment, promotion by pull request
One control plane for production and non-productionA non-production compromise or bad upgrade reaches productionA control plane per trust zone
A hub holding admin credentials for regulated clustersOne breach exposes the fleetPull inside the regulated zone, with scoped service accounts
Two GitOps tools managing the same objectsControllers overwrite each other and resources flapOne owner per namespace; label ownership and enforce it with policy
Automated sync with prune and weak branch protectionA bad merge can delete production resourcesRequired reviews and CODEOWNERS; Argo CD sync windows or Flux prune annotations for critical resources
Plaintext secrets in GitLeaks persist in historyExternal Secrets Operator or SOPS
No break-glass procedureEngineers patch by hand during incidents and self-heal reverts the fixSuspend, fix, commit the fix to Git, resume
Treating the GitOps tool as set-and-forgetIt falls outside supported versions and upgrades become migrationsQuarterly upgrades through Git, tested in non-production first

12. Argo CD vs Flux comparison

AreaArgo CD 3.5Flux 2.9
ArchitectureServer with API, UI, repository server and application controllerIndependent controllers configured by custom resources
Default multi-cluster modelOne instance pushes to many clusters; agent mode for pullFlux in every cluster pulling; hub and spoke optional
Web UIBuilt in, central for the fleetFlux Operator web UI; CLI first
Fleet templatingApplicationSet generators: cluster, Git, list, matrix, pull requestKustomization with postBuild substitution; Flux Operator ResourceSets
Rollout across clustersApplicationSet progressive syncsdependsOn chains with health-aware wait
Access controlAppProjects and SSO-mapped rolesKubernetes RBAC and per-tenant service accounts
SourcesGit, Helm repositories, OCIGit, Helm, OCI, S3-compatible buckets
Signature verificationSource Integrity, new in 3.5Git signature verification; cosign for OCI
Image update automationArgo CD Image Updater, a separate projectBuilt in, generally available since 2.7
Rendered manifestsSource Hydrator, beta in 3.5Render in CI or publish rendered OCI artifacts
In-cluster progressive deliveryArgo RolloutsFlagger
Commercial supportSeveral vendors and managed offeringsControlPlane enterprise distribution; Azure managed extension

13. Decision helper

Answer five questions about your estate. Example answers are pre-selected; change them to match your environment.

Figure 06 Decision helperFive questions to a starting recommendation
Do product teams want a web UI to see and sync their own applications?
Can a central cluster reach every cluster's Kubernetes API within one trust zone?
Do you want to ship configuration as OCI artifacts instead of reading Git directly?
Do many teams share the GitOps tool and need SSO-based permissions inside it?
Are clusters created by Cluster API or Terraform, with GitOps installed during provisioning?

Example answers are pre-selected. The helper gives a starting point for an architecture decision record, not a substitute for one.

Frequently asked questions

Can an enterprise run Argo CD and Flux together?

Yes. A common split is Flux for platform add-ons installed during cluster provisioning and Argo CD for application teams. Give each tool exclusive ownership of its namespaces and resources so the controllers do not overwrite each other.

Is Flux still maintained after Weaveworks closed?

Yes. Weaveworks shut down in early 2024, and Flux continued as a CNCF graduated project with maintainers from several companies. It has shipped regular releases since, through Flux 2.9 in June 2026.

How many clusters can one Argo CD instance manage?

There is no fixed limit. It depends on resource count, sync frequency and controller sharding. In enterprises the constraint is usually trust zones and blast radius rather than raw scale, which is why production and non-production get separate control planes.

Does GitOps satisfy change-management audit requirements?

It provides strong evidence: every change is a reviewed pull request with an approver, a timestamp and a diff, and the controller records what was applied. Auditors still expect documented policy, branch protection, separation of duties and a break-glass procedure.

Adopt GitOps with CloudForge

CloudForge helps platform and security teams choose the right pattern for each trust zone and stand it up with the controls auditors expect. Engagements are documented and designed for a clean handover to your team.

Planning a GitOps rollout across your clusters? Contact CloudForge to review your estate and agree the pattern for each trust zone.

Sources and further reading

Related expertise

Put this into practice

Platform engineering consulting DevOps consulting CI/CD deployment engineering
Continue learning
Eliminating Deployment Bottlenecks: How Enterprise Engineering Teams Ship Code Faster7 min read AWS EKS Cost Optimization: Karpenter, Spot, Graviton and Pod Rightsizing16 min read Internal Developer Platforms: Navigating the Build Versus Buy Decision for Scaling Teams14 min read

Want this applied to your cloud environment?

Send CloudForge your requirements and the company will identify the highest-impact next step for your cost, delivery or reliability goals.

Contact CloudForge