kubectl edit during an outage, an ad-hoc Helm upgrade is run from a local laptop, or a rogue controller alters resource requests. Days or weeks later, the next pipeline deployment wipes out the fix, triggering a catastrophic repeat incident. Here is the comprehensive architectural blueprint for eliminating drift entirely using ArgoCD and strict GitOps principles.
In traditional software delivery models, continuous deployment systems push changes into the cluster: a Jenkins runner, a GitHub Actions workflow, or an engineer running kubectl apply -f manifest.yaml from their terminal.
This imperative, push-based model inevitably causes Configuration Drift: the state of the live running resources inside your Kubernetes cluster deviates from the version-controlled manifests declared in your Git repository.
kubectl edit deployment payment-api to increase memory limits from 1Gi to 4Gi or change an environment variable. The fix works, the alert resolves, but nobody ever commits the change back to the Git repository.kubectl scale --replicas=10 deployment checkout or kubectl rollout restart.helm upgrade --install with local, uncommitted values files, bypassing CI/CD checks.replicas: 3, creating an infinite thrashing cycle.GitOps is not merely using Git for storing YAML files. It is an operational model defined by four foundational OpenGitOps principles:
| Layer / Component | Desired State (Git) | Live Cluster State | ArgoCD Self-Healing Action |
|---|---|---|---|
| Deployment Replicas | replicas: 3 (or managed by HPA) |
Dynamic HPA scaling (3 → 15) | Preserved via ignoreDifferences |
| Container Memory Limits | limits.memory: 2Gi |
Ad-hoc kubectl edit (4Gi) |
Reverted to 2Gi (or PR merged) |
| Environment Variables | LOG_LEVEL: info |
Manual change in dashboard | Immediate GitOps overwrite |
| Orphaned ConfigMaps | Removed in latest Git commit | Existing in namespace | Pruned automatically (prune: true) |
ArgoCD implements the GitOps model through three primary control plane components running natively inside Kubernetes:
How does ArgoCD determine whether a resource has drifted? It does not perform a naive text comparison. Instead, it utilizes a three-way merge calculation:
GET /apis/...).kubectl.kubernetes.io/last-applied-configuration annotation or Server-Side Apply managed fields.
If the live state deviates from the desired state in fields not explicitly marked to be ignored, ArgoCD flags the resource status as OutOfSync.
To eliminate configuration drift, an ArgoCD Application must be configured with both prune: true and selfHeal: true.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: payment-service-production
namespace: argocd
finalizers:
- resources-finalizer.argocd.argoproj.io
spec:
project: production-platform
source:
repoURL: https://github.com/naveedkumbhar/production-gitops.git
targetRevision: main
path: environments/production/payment-service
destination:
server: https://kubernetes.default.svc
namespace: payments
syncPolicy:
automated:
prune: true # Deletes resources removed from Git
selfHeal: true # Overwrites manual cluster changes immediately
allowEmpty: false
syncOptions:
- CreateNamespace=true
- ApplyOutOfSyncOnly=true
- PruneLast=true
- ServerSideApply=true
- RespectIgnoreDifferences=true
retry:
limit: 5
backoff:
duration: 5s
factor: 2
maxDuration: 3m
ignoreDifferences:
# 1. HPA Replicas Exception
- group: apps
kind: Deployment
jsonPointers:
- /spec/replicas
# 2. Mutating Webhooks / Certificate Managers
- group: ""
kind: Secret
jsonPointers:
- /data/tls.crt
- /data/tls.key
| Directive | Impact on Configuration Drift | Production Recommendation |
|---|---|---|
prune: true |
Automatically deletes resources in the cluster that were deleted from the Git repository. Prevents zombie resources and dangling ingress rules. | Mandatory. Without this, deleted microservices or old configmaps persist forever. |
selfHeal: true |
Whenever an out-of-band change occurs in Kubernetes (e.g. kubectl edit), ArgoCD instantly overwrites the live state back to Git state. |
Mandatory for Staging & Production. Completely eliminates imperative human drift. |
ServerSideApply=true |
Leverages Kubernetes Server-Side Apply (SSA) instead of client-side kubectl.kubernetes.io/last-applied-configuration. Avoids 256KB annotation limits on large CRDs. |
Highly Recommended. Modern standard for Kubernetes 1.25+. |
ApplyOutOfSyncOnly=true |
Only issues PATCH requests for resources that have drifted, avoiding unnecessary API hammering across hundreds of unchanged resources. | Critical for scale. Dramatically reduces API server CPU load. |
RespectIgnoreDifferences=true |
Ensures that fields configured under ignoreDifferences are NOT overwritten during automated syncs. |
Mandatory when using HPA or dynamic controllers. |
ignoreDifferences
The single most common mistake teams make when enabling selfHeal: true is triggering a reconciliation conflict storm with Kubernetes controllers.
replicas: 3 for your API Deployment.15 replicas.OutOfSync.selfHeal: true enabled, ArgoCD immediately applies the Git manifest, scaling the Deployment down to 3 replicas!
To resolve this, tell ArgoCD to ignore the replica count during state comparison, and use RespectIgnoreDifferences=true to keep it from applying the Git replica value during syncs:
spec:
syncPolicy:
syncOptions:
- RespectIgnoreDifferences=true
ignoreDifferences:
# Option A: JSON Pointer targeting Deployment replicas
- group: apps
kind: Deployment
jsonPointers:
- /spec/replicas
# Option B: Ignore changes made by kube-controller-manager
- group: apps
kind: Deployment
managedFieldsManagers:
- kube-controller-manager
# Option C: JQ Path Expression for dynamic annotations
- group: apps
kind: Deployment
jqPathExpressions:
- .metadata.annotations["deployment.kubernetes.io/revision"]
Additionally, in your Git repository manifests, omit the spec.replicas field entirely from the Deployment YAML once an HPA is attached. Kubernetes will retain the live replica count while HPA manages it dynamically.
Self-healing catches drift, but the root cause of human-induced drift is overly permissive cluster access. If developers have cluster-admin or namespace write privileges in production, configuration drift will persist.
kubectl get, kubectl describe, kubectl logs, but cannot run apply, edit, delete, or scale.# Read-Only RBAC Role for Production Engineers
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: sre-developer-readonly
rules:
- apiGroups: ["", "apps", "batch", "networking.k8s.io"]
resources:
- pods
- pods/log
- services
- deployments
- statefulsets
- cronjobs
- configmaps
- ingresses
verbs: ["get", "list", "watch"]
- apiGroups: [""]
resources: ["pods/exec"] # Blocked in production; use ephemeral debug pods if needed
verbs: []
One of the biggest excuses engineers offer for manually editing Kubernetes clusters is secrets rotation: "We needed to update an API key in a hurry!"
Committing plain text secrets to Git violates compliance and security standards. Managing secrets imperatively with kubectl create secret creates catastrophic drift.
The industry standard solution is the External Secrets Operator (ESO):
ExternalSecret and SecretStore manifests are committed to your Git repository.Secret objects periodically (e.g. refreshInterval: 1h).ExternalSecret object, while ESO handles secret hydration—eliminating both secret drift and security compromises.apiVersion: external-secrets.io/v1beta1
kind: ExternalSecret
metadata:
name: payment-gateway-credentials
namespace: payments
spec:
refreshInterval: "1h"
secretStoreRef:
name: aws-secrets-manager
kind: ClusterSecretStore
target:
name: payment-gateway-secret
creationPolicy: Owner
data:
- secretKey: STRIPE_API_KEY
remoteRef:
key: production/payments/stripe
property: api_key
Do not wait for a failed deployment to discover that someone tampered with cluster resources. Monitor ArgoCD Prometheus metrics to trigger immediate alerts whenever an application enters an OutOfSync state.
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: argocd-drift-alerts
namespace: monitoring
spec:
groups:
- name: ArgoCDDrift
rules:
- alert: ArgoCDApplicationOutOfSync
expr: |
argocd_app_info{sync_status="OutOfSync",project="production-platform"} == 1
for: 10m
labels:
severity: warning
team: platform-sre
annotations:
summary: "Configuration drift detected on {{ $labels.name }}"
description: "Application {{ $labels.name }} in namespace {{ $labels.dest_namespace }} has been OutOfSync for over 10 minutes. Live state differs from Git source."
When an application flags OutOfSync or fails to sync in production, follow this standardized SRE triage workflow:
argocd app diff payment-service-production --local environments/production/payment-service
Identify whether the discrepancy is in user-defined business logic, an unexpected environment variable, or controller-injected metadata.
argocd app get payment-service-production --show-params
Determine which exact sub-resource (Deployment, Service, Ingress) is failing to converge.
kubectl get deployment payment-service -n payments -o yaml | grep -A 20 "managedFields:"
Check if another manager (e.g. kustomize-controller, helm, or an admission webhook) has claimed ownership of conflicting fields.
argocd app sync payment-service-production --force --server-side
ArgoCD continuously runs a three-way reconciliation algorithm (every 180 seconds by default, or immediately when triggered via Git webhooks). It compares the desired state from Git against the live cluster state and the resource's last-applied configuration or managed fields. Any deviation marks the application as OutOfSync.
Auto-sync reconciles changes pushed to Git into the cluster. Self-healing (selfHeal: true) actively watches the cluster: if an engineer runs kubectl edit or an imperative command altering live resources, ArgoCD detects the runtime drift and overwrites the cluster back to the state in Git.
Configure ignoreDifferences for /spec/replicas on Deployments in your ArgoCD Application manifest, and set RespectIgnoreDifferences=true in syncOptions. Additionally, omit the replicas field from your base Deployment manifests in Git once an HPA is attached.
By default, ignoreDifferences only suppresses the visual diff warning in the ArgoCD UI. During an automated sync or self-heal, ArgoCD will still try to enforce the Git value unless RespectIgnoreDifferences=true is explicitly enabled in syncOptions.
OutOfSync indicates a configuration discrepancy between Git and the live cluster state. Degraded indicates that the runtime workload is failing or unhealthy (e.g. pods in CrashLoopBackOff, missing PVCs, or failing readiness probes), regardless of whether manifests match Git.
Building zero-drift, highly resilient GitOps pipelines across multi-tenant AWS EKS clusters requires deep architectural precision. Explore over 970 battle-tested production scenarios or dive into the complete curriculum.