Argo CD vs Flux: An Enterprise Guide to Multi-Cluster GitOps
Compare Argo CD and Flux for enterprise multi-cluster GitOps: topology patterns, reference architecture, governance controls, operating model and adoption roadmap.
Argo CD and Flux are both mature, CNCF graduated GitOps controllers. Either can reconcile a Kubernetes estate from Git or an OCI registry. For an enterprise running clusters across environments, regions and compliance boundaries, the harder decision is not which tool has the longer feature list. It is where reconciliation runs, which side opens the network connection, and where cluster credentials are stored.
This guide is written for platform leads, architects, CTOs and security reviewers. It compares Argo CD 3.5 and Flux 2.9 as of October 2026, then sets out a reference architecture, governance controls, an operating model and a phased adoption roadmap that apply to either tool.
Executive summary
- Both tools are production-grade. The functional gap between them narrows with each release. Choose on operating model and topology before features.
- Topology matters more than the tool. Decide per trust zone whether a central control plane pushes to clusters or each cluster pulls its own state.
- Choose Argo CD when many product teams need a shared web UI, SSO-mapped permissions and fleet templating from a central control plane.
- Choose Flux when clusters must pull their own state with no inbound access and no central credential store, or when GitOps should be installed with the cluster as plain Kubernetes objects.
- Never share a control plane across trust zones. Separate production from non-production, and give regulated environments their own.
| If your estate looks like this | Recommended pattern | Why |
|---|---|---|
| A few network zones, many product teams, dozens of clusters | Argo CD: one highly available instance per environment tier, with ApplicationSets | Shared UI and SSO permissions; push is acceptable inside one trust zone |
| Regulated workloads, private or edge clusters, no inbound access allowed | Flux in each cluster, or the Argo CD agent where a central UI is mandatory | No central credential store; failures stay local |
| Clusters created by Cluster API or Terraform at scale | Flux installed during provisioning; hub and spoke for the platform layer | GitOps arrives with the cluster as Kubernetes objects |
| A platform team serving many product teams | Flux for platform add-ons, Argo CD for applications, with strict ownership | Each tool covers the layer it suits best |
1. How GitOps reconciliation works
GitOps replaces imperative deploy scripts with a controller that continuously compares two things: the desired state in Git or an OCI registry, and the live state in the cluster. When they differ, the controller applies the difference. When someone changes the cluster by hand, the controller detects the drift and restores the approved state.
For an enterprise this matters for two reasons. The Git history becomes the change record, with an author, reviewer, timestamp and diff for every production change. And unapproved changes do not survive the next reconcile.
- Desired (Git)
- replicas: 3
- Live (cluster)
- replicas: 3
- Image
- checkout@sha256:9f2c…
Synced and healthy
Illustrative example. A manual scale-up creates drift; self-heal (Argo CD selfHeal, or Flux reconciliation) restores the approved state from Git.
Self-heal is a policy decision. Argo CD reverts drift when selfHeal is enabled on an Application; Flux reverts it on its regular reconciliation interval. Both need a documented break-glass path for incidents, covered in the governance section below.
2. Argo CD and Flux in 2026
Argo CD 3.5: an application delivery platform
Argo CD runs an API server, web UI, repository server and application controller. The Application is the unit of delivery, and ApplicationSet generates Applications from cluster inventories, Git directories or pull requests.
Version 3.5, released as a candidate in June 2026, added mutual TLS between internal components, Git commit signature verification through Source Integrity, ApplicationSet management in the UI, Helm 4 support and ApplicationSets in any namespace. It also promoted impersonation and the Source Hydrator to beta. Version 3.4 added a way to register a cluster without the central controller reconciling it, which supports estates that mix a central hub with agents.
Flux 2.9: a toolkit of composable controllers
Flux is a set of controllers for sources, Kustomize, Helm, notifications and image automation, configured entirely through Kubernetes custom resources. There is no required central server.
OCI artifact support reached general availability in Flux 2.6 and image update automation in Flux 2.7. Flux 2.9, released on June 30, 2026, added a CLI plugin system and field-level ignore rules for server-side apply drift detection. The Flux Operator, maintained by ControlPlane, adds a web UI, sharding and multi-tenancy lockdown, and is available through Red Hat OperatorHub.
3. The four multi-cluster patterns
Every multi-cluster design answers three questions: which side opens the network connection, where cluster credentials are stored, and what happens to the fleet when the central cluster fails.
- Who connects
- Argo CD on the hub opens a connection to every spoke's Kubernetes API.
- Credentials
- The hub stores a credential for every cluster.
- If the hub fails
- No cluster receives new syncs. Running workloads keep running.
- Best for
- One trust zone with many product teams that want a single UI.
- Watch for
- Cross-network API traffic and egress cost; the hub becomes a high-value target.
Arrows point in the direction the connection is opened. The table below summarizes all four patterns.
| Pattern | Connection | Credentials | If the central cluster fails | Best fit |
|---|---|---|---|---|
| Argo CD hub push | Hub opens a connection to each spoke's Kubernetes API | Hub stores one per cluster | No syncs anywhere; workloads keep running | One trust zone, many product teams, one UI |
| Argo CD agent pull | Agents on spokes connect out to the hub over gRPC with client certificates | None stored on the hub | Spokes keep current state; UI and new configuration pause | Firewalled or edge clusters that still need a central UI |
| Flux per-cluster pull | Each cluster pulls from Git or OCI; nothing connects inbound | Each cluster holds only its own read credential | No hub to fail; impact stays in one cluster | Regulated, edge and large uniform fleets |
| Flux hub and spoke | A management cluster applies through kubeconfig secrets | Hub stores one per cluster | No reconciliation until the hub returns | Cluster API estates with a management cluster |
Argo CD hub push and Flux hub and spoke share the same risk profile: one cluster holds credentials for the fleet. The difference is operational. Argo CD provides a single UI and permission model; Flux provides plain Kubernetes objects that provisioning tools can create with the cluster, and a hub that can be sharded across several Flux instances.
The agent pattern is Argo CD's answer to the pull model. The argocd-agent project inverts the connection so that spokes connect out to a principal on the hub, authenticate with client certificates and reconcile locally. Its documentation lists high availability for the principal and full multi-tenancy as work in progress, so treat it as maturing for production use.
4. Enterprise reference architecture
The architecture below applies to either tool. It separates control planes by trust zone, keeps a single source of truth with a signed supply chain, and treats regulated environments as pull-only.
The same architecture applies to either tool. Switch the tool to see which components fill each role.
Design principles
- Separate control planes per trust zone. Production, non-production and regulated environments never share a GitOps control plane.
- Pull across untrusted networks. Use hub push only inside one trust zone. Across zones, clusters pull.
- Make Git the change record. No routine kubectl in production. Exceptions follow a documented break-glass procedure.
- Promote immutable artifacts. Production references image digests and versioned configuration, never a moving branch or a
latesttag. - Split platform and application ownership. Platform add-ons and product workloads live in separate paths with separate owners.
- Enforce policy at admission. CI checks are advisory. Kyverno or OPA Gatekeeper in the cluster is the control.
- Keep secrets out of Git. Reference external stores through External Secrets Operator, or encrypt with SOPS.
- Manage the GitOps platform as code. Bootstrap it with Terraform, upgrade it through Git, and rehearse rebuilding it.
The supply chain controls behind principle 4 are covered in the DevSecOps pipeline security guide, and the platform boundaries behind principle 5 in the internal developer platform strategy guide.
5. Repository structure and promotion
Use one branch with a folder per environment, not a branch per environment. Environment branches drift apart through cherry-picks and make it hard to see what is running where. A folder layout makes every environment's state reviewable on main, and CODEOWNERS assigns approval rights per path.
platform-config/
├── clusters/ # what runs where, one folder per cluster
│ ├── nonprod/dev-use1/
│ ├── nonprod/stg-use1/
│ └── prod/prd-use1/ prd-euw1/
├── platform/ # add-ons owned by the platform team
│ ├── base/ # ingress, cert-manager, external-secrets, kyverno
│ └── overlays/nonprod/ prod/
└── apps/ # product teams own their folders via CODEOWNERS
└── payments/
├── base/
└── overlays/dev/ staging/ prod/ # prod pins image digests
Promotion is a pull request that copies a tested digest from the staging overlay to the production overlay. CI can open that pull request automatically after staging passes; a person with production approval rights merges it. Rollout across production regions then proceeds in waves, and an unhealthy wave stops the release.
Dev
ApplicationSet step 1- dev-1v2.14.1
Staging
RollingSync step 2- stg-usv2.14.1
- stg-euv2.14.1
Production
RollingSync step 3 · maxUpdate 25%- prod-us-eastv2.14.1
- prod-us-westv2.14.1
- prod-eu-westv2.14.1
- prod-ap-southv2.14.1
Every cluster is running v2.14.1. Start the release to roll out v2.15.0.
Illustrative rollout of one service across seven clusters. Argo CD uses ApplicationSet progressive syncs; Flux chains Kustomizations with dependsOn and wait: true.
Argo CD expresses waves with ApplicationSet progressive syncs. Each RollingSync step waits for the previous step's Applications to report Healthy, and maxUpdate limits how many production clusters change at once. Progressive syncs must be enabled on the ApplicationSet controller.
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
name: payments
namespace: argocd
spec:
goTemplate: true
generators:
- clusters:
selector:
matchLabels:
fleet: payments
strategy:
type: RollingSync
rollingSync:
steps:
- matchExpressions:
- {key: env, operator: In, values: [staging]}
- matchExpressions:
- {key: env, operator: In, values: [prod]}
maxUpdate: 25% # one production cluster in four at a time
template:
metadata:
name: 'payments-{{.name}}'
labels:
env: '{{index .metadata.labels "env"}}'
spec:
project: payments # AppProject limits repositories and destinations
source:
repoURL: https://github.com/acme/platform-config
targetRevision: main
path: 'apps/payments/overlays/{{index .metadata.labels "env"}}'
destination:
server: '{{.server}}'
namespace: payments
Flux expresses the same ordering with dependsOn between Kustomizations. Setting wait: true means a Kustomization reports Ready only when its workloads are healthy, so the next wave starts only after the previous one succeeds. In hub mode, kubeConfig points at the target cluster; in per-cluster mode it is omitted.
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: payments-prod-us-east
namespace: flux-system
spec:
interval: 10m
retryInterval: 2m
serviceAccountName: payments-deployer # tenant RBAC
sourceRef:
kind: GitRepository
name: platform-config
path: ./apps/payments/overlays/prod
prune: true
wait: true
dependsOn:
- name: payments-staging
kubeConfig: # hub mode only
secretRef:
name: prod-us-east-kubeconfig
postBuild:
substitute:
cluster_name: prod-us-east
region: us-east-1
For canary traffic shifting inside a single cluster, the usual companions are Argo Rollouts and Flagger. The enterprise deployment bottlenecks guide covers the approval and release practices around this flow.
6. Governance, security and compliance
GitOps gives auditors stronger evidence than pipeline-pushed deployments. Every production change is a reviewed pull request with an approver, a timestamp and a diff, and the controller records what it applied. The table maps common control objectives to the mechanism in each tool. Map the rows to your own framework, for example SOC 2 change management (CC8.1), ISO/IEC 27001:2022 control 8.32 or PCI DSS requirement 6.5.
| Control objective | GitOps mechanism | Argo CD | Flux |
|---|---|---|---|
| Change approval | Branch protection, required reviews, CODEOWNERS per path | Enforced in the Git platform | Enforced in the Git platform |
| Separation of duties | Authors cannot approve their own production change; no standing human write access to production clusters | AppProjects restrict repositories, clusters and namespaces; SSO groups map to roles | Per-tenant service accounts through serviceAccountName; Kubernetes RBAC |
| Integrity of what is deployed | Signed commits and artifacts, verified before apply | Source Integrity commit verification (3.5) | GitRepository signature verification; cosign verification for OCI |
| Audit trail | Git history plus controller events forwarded to the SIEM | Notifications controller and API audit logs | Notification controller alerts and events |
| Drift detection | Unapproved changes detected and reverted | OutOfSync status and selfHeal | Continuous reconciliation; SSA field ignore rules (2.9) |
| Secrets handling | No plaintext secrets in Git | External Secrets Operator or a secrets plugin | Native SOPS decryption, or External Secrets Operator |
| Break-glass | Documented emergency path, then reconcile the fix back into Git | Disable automated sync for the Application; skip reconciliation for a cluster (3.4 and later) | flux suspend the Kustomization, then resume |
| Control plane hardening | The GitOps system is itself a privileged workload | Internal mTLS (3.5), SSO, HA installation | Multi-tenancy lockdown and network policies through Flux Operator |
Credential concentration is the largest security difference between patterns. A hub that stores a kubeconfig for every cluster is one breach away from the whole fleet. Per-cluster Flux and the Argo CD agent keep credentials local; prefer them for regulated or internet-exposed fleets. Our DevSecOps consulting team can map these controls to your audit scope.
7. Operating model
Most GitOps programs that stall do so on ownership rather than technology. Agree who is responsible (R), accountable (A), consulted (C) and informed (I) for each activity before the pilot starts.
| Activity | Platform team | Product teams | Security | SRE / on-call |
|---|---|---|---|---|
| Run, upgrade and back up the GitOps control plane | A/R | I | C | C |
| Templates, golden paths and cluster onboarding | A/R | C | C | I |
| Application manifests and promotion requests | C | A/R | I | I |
| Admission policy definition and exceptions | R | C | A | I |
| Triage sync failures and drift alerts | C | R | I | A/R |
| Approve break-glass and reconcile afterward | C | R | I | A |
The platform team's role here mirrors the product-team model described in platform engineering golden paths: templates and paved paths are the product, and product teams are the customers.
8. Adoption roadmap
A phased rollout proves the model with willing teams before it becomes mandatory. Durations are indicative for an estate of a few dozen clusters, and each phase ends only when its exit criteria are met.
- Phase 0 · Weeks 1–2
Assess
- Inventory clusters, network zones, teams and compliance scope
- Choose a pattern per trust zone
- Baseline DORA metrics
Exit criteriaAn approved architecture decision record for each trust zone
- Phase 1 · Weeks 3–8
Foundation
- Bootstrap control planes with Terraform
- Repository layout, CODEOWNERS, SSO and RBAC
- Secrets integration; admission policy in audit mode
- Controller metrics and alerts
Exit criteriaNon-production control plane rebuilt from Git in a test; platform add-ons delivered by GitOps
- Phase 2 · Weeks 9–12
Pilot
- One or two willing product teams
- Promotion by pull request from dev to production
- Break-glass drill
Exit criteriaPilot services reach production only through Git for four consecutive weeks
- Phase 3 · Weeks 13–24
Scale
- Self-service onboarding template
- Fleet templating and progressive rollout across regions
- Admission policy switched to enforce
- Remove standing human write access to production
Exit criteriaAgreed share of services onboarded; no standing cluster-admin in production
- Phase 4 · Ongoing
Operate
- Quarterly tool upgrades through Git, non-production first
- Control plane rebuild game days
- Monthly metrics review
Exit criteriaUpgrades and rebuilds are routine work, not projects
Indicative timeline for an estate of a few dozen clusters. Adjust durations to your fleet, team capacity and audit calendar.
9. Scale and resilience of the control plane
The GitOps control plane becomes tier-0 infrastructure. Design and operate it accordingly.
- Argo CD scaling. Use the HA installation, shard the application controller across clusters and scale repository servers for manifest rendering. Argo CD 3.5 also adds concurrency to how the ApplicationSet controller manages Applications.
- Flux scaling. Per-cluster Flux scales with the fleet by design. A Flux hub can be sharded across several instances, and the Flux Operator configures sharding and vertical scaling.
- Failure behavior. When the control plane is down, running workloads keep running; only new changes and drift correction stop. State this explicitly in your recovery objectives.
- Disaster recovery. The desired state is already in Git. Back up what is not: cluster credentials, SSO configuration and secret references. Rehearse a full control plane rebuild from Terraform and Git at least twice a year. The cloud disaster recovery guide covers setting RTO and RPO for tier-0 services.
- Upgrades. Read the upgrade notes for every minor version. For example, Argo CD 3.3 required the
ServerSideApply=truesync option on the Application that manages Argo CD itself. Roll upgrades through non-production first.
10. Measuring success
Report outcomes leadership already tracks, plus a few GitOps-specific indicators taken from the controllers' Prometheus metrics and your Git platform. The DevOps and SRE operating model guide explains consistent DORA definitions.
| Metric | Definition | Direction to aim for |
|---|---|---|
| Lead time for changes | Pull request merged to running in production | Falling |
| Deployment frequency | Production promotions per service per week | Rising or stable |
| Change failure rate | Share of promotions that cause a rollback or incident | Falling |
| Time to restore | Incident start to service restored, including Git reverts | Falling |
| Changes delivered through Git | Production changes made by the controller rather than by people | 100% outside break-glass |
| Drift events | Out-of-sync events caused by manual changes | Near zero |
| Reconcile success rate | Successful syncs over total, per cluster | Stable; investigate dips |
11. Anti-patterns and how to avoid them
| Anti-pattern | Why it hurts | Do this instead |
|---|---|---|
| A branch per environment | Branches drift through cherry-picks; nobody can see what runs where | One branch, a folder per environment, promotion by pull request |
| One control plane for production and non-production | A non-production compromise or bad upgrade reaches production | A control plane per trust zone |
| A hub holding admin credentials for regulated clusters | One breach exposes the fleet | Pull inside the regulated zone, with scoped service accounts |
| Two GitOps tools managing the same objects | Controllers overwrite each other and resources flap | One owner per namespace; label ownership and enforce it with policy |
| Automated sync with prune and weak branch protection | A bad merge can delete production resources | Required reviews and CODEOWNERS; Argo CD sync windows or Flux prune annotations for critical resources |
| Plaintext secrets in Git | Leaks persist in history | External Secrets Operator or SOPS |
| No break-glass procedure | Engineers patch by hand during incidents and self-heal reverts the fix | Suspend, fix, commit the fix to Git, resume |
| Treating the GitOps tool as set-and-forget | It falls outside supported versions and upgrades become migrations | Quarterly upgrades through Git, tested in non-production first |
12. Argo CD vs Flux comparison
| Area | Argo CD 3.5 | Flux 2.9 |
|---|---|---|
| Architecture | Server with API, UI, repository server and application controller | Independent controllers configured by custom resources |
| Default multi-cluster model | One instance pushes to many clusters; agent mode for pull | Flux in every cluster pulling; hub and spoke optional |
| Web UI | Built in, central for the fleet | Flux Operator web UI; CLI first |
| Fleet templating | ApplicationSet generators: cluster, Git, list, matrix, pull request | Kustomization with postBuild substitution; Flux Operator ResourceSets |
| Rollout across clusters | ApplicationSet progressive syncs | dependsOn chains with health-aware wait |
| Access control | AppProjects and SSO-mapped roles | Kubernetes RBAC and per-tenant service accounts |
| Sources | Git, Helm repositories, OCI | Git, Helm, OCI, S3-compatible buckets |
| Signature verification | Source Integrity, new in 3.5 | Git signature verification; cosign for OCI |
| Image update automation | Argo CD Image Updater, a separate project | Built in, generally available since 2.7 |
| Rendered manifests | Source Hydrator, beta in 3.5 | Render in CI or publish rendered OCI artifacts |
| In-cluster progressive delivery | Argo Rollouts | Flagger |
| Commercial support | Several vendors and managed offerings | ControlPlane enterprise distribution; Azure managed extension |
13. Decision helper
Answer five questions about your estate. Example answers are pre-selected; change them to match your environment.
Example answers are pre-selected. The helper gives a starting point for an architecture decision record, not a substitute for one.
Frequently asked questions
Can an enterprise run Argo CD and Flux together?
Yes. A common split is Flux for platform add-ons installed during cluster provisioning and Argo CD for application teams. Give each tool exclusive ownership of its namespaces and resources so the controllers do not overwrite each other.
Is Flux still maintained after Weaveworks closed?
Yes. Weaveworks shut down in early 2024, and Flux continued as a CNCF graduated project with maintainers from several companies. It has shipped regular releases since, through Flux 2.9 in June 2026.
How many clusters can one Argo CD instance manage?
There is no fixed limit. It depends on resource count, sync frequency and controller sharding. In enterprises the constraint is usually trust zones and blast radius rather than raw scale, which is why production and non-production get separate control planes.
Does GitOps satisfy change-management audit requirements?
It provides strong evidence: every change is a reviewed pull request with an approver, a timestamp and a diff, and the controller records what was applied. Auditors still expect documented policy, branch protection, separation of duties and a break-glass procedure.
Adopt GitOps with CloudForge
CloudForge helps platform and security teams choose the right pattern for each trust zone and stand it up with the controls auditors expect. Engagements are documented and designed for a clean handover to your team.
- Infrastructure Audit: a review of your clusters, pipelines and delivery controls with a prioritized roadmap. See pricing.
- Platform engineering: reference architecture, repository design, bootstrap and pilot through platform engineering consulting.
- Delivery and reliability: DevOps consulting, CI/CD deployment engineering and SRE consulting for scale-out and operations.
Planning a GitOps rollout across your clusters? Contact CloudForge to review your estate and agree the pattern for each trust zone.
Sources and further reading
- InfoQ: Argo CD 3.5 tightens supply chain security with internal mTLS and Source Integrity
- Argo CD v3.5.2 release notes
- Argo CD v3.5.0-rc1 changelog
- argocd-agent: features overview
- argocd-agent: hybrid architecture
- Red Hat: Argo CD Agent architecture overview
- Flux: Announcing Flux 2.9 GA
- Flux v2.9.0 release notes
- Flux: Announcing Flux 2.7 GA
- ControlPlane: Flux architecture and multi-cluster strategies
Want this applied to your cloud environment?
Send CloudForge your requirements and the company will identify the highest-impact next step for your cost, delivery or reliability goals.
Contact CloudForge