Managing Hybrid Workloads with Anthos: From On-Prem Kubernetes to GCP at Enterprise Scale

Practical guide to running Anthos across on-premises data centers and GCP, covering fleet management, policy enforcement, and service mesh at scale.

#gcp#anthos#hybrid-cloud#kubernetes
Cover image for the article: Managing Hybrid Workloads with Anthos: From On-Prem Kubernetes to GCP at Enterprise Scale

Running Kubernetes in two places is easy. Running it consistently, securely, and observably across on-premises data centers and GCP without doubling your platform team — that's the actual problem Anthos solves. After 18 months managing 23 clusters across three data centers and two GCP regions, here's what works, what doesn't, and what the documentation won't tell you.

The Problem Statement

Our enterprise had a common constraint: regulatory requirements mandating certain workloads stay on-premises, combined with a strategic push toward GCP for everything else. The result was two separate Kubernetes ecosystems with different tooling, policies, and operational procedures.

The pain points before Anthos:

  • Policy drift: Security policies applied in GCP didn't match on-prem configurations
  • Deployment inconsistency: CI/CD pipelines forked into separate paths per environment
  • Observability gaps: No unified view across all clusters
  • Certificate management: Different PKI infrastructure per environment
  • Team overhead: Separate SRE runbooks for each environment

Hybrid Architecture Before Anthos

Anthos Architecture: What You Actually Deploy

Anthos isn't a single product — it's a control plane that layers fleet management, policy enforcement, service mesh, and configuration management across heterogeneous clusters.

Fleet Registration

Every cluster — GKE on GCP, GKE on-prem (VMware or bare metal), or attached third-party clusters — registers to a fleet:

# fleet-membership.yaml
apiVersion: gkehub.googleapis.com/v1
kind: Membership
metadata:
  name: dc-east-prod-01
spec:
  endpoint:
    kubernetesResource:
      membershipCrSpec:
        # Connect agent auto-installed
        connectAgent:
          namespace: gke-connect
      resourceOptions:
        connectVersion: "1.14"
  authority:
    issuer: "https://container.googleapis.com/v1/projects/my-project/locations/us-east1/memberships/dc-east-prod-01"

The Connect Agent establishes an outbound-only tunnel from your on-prem cluster to GCP's fleet management API. No inbound firewall rules required — this was a key security requirement for our network team.

Fleet Topology

Our fleet after migration:

ClusterLocationTypeNodesWorkload Type
dc-east-prod-01Virginia DCBare metal48Regulated financial
dc-east-prod-02Virginia DCBare metal32PII processing
dc-west-prod-01Oregon DCVMware24Internal services
gke-east-produs-east4GKE AutopilotAutoPublic APIs
gke-central-produs-central1GKE Standard64ML workloads
gke-west-produs-west1GKE AutopilotAutoCDN origin

Config Sync: GitOps at Fleet Scale

Config Sync is Anthos's GitOps engine. It watches a Git repository and reconciles cluster state — but unlike vanilla ArgoCD, it operates fleet-wide with namespace-scoped inheritance.

# config-sync/rootsync.yaml
apiVersion: configsync.gke.io/v1beta1
kind: RootSync
metadata:
  name: root-sync
  namespace: config-management-system
spec:
  sourceFormat: unstructured
  git:
    repo: "https://github.com/myorg/fleet-config"
    branch: main
    dir: "clusters"
    auth: gcpserviceaccount
    gcpServiceAccountEmail: config-sync@my-project.iam.gserviceaccount.com
  override:
    reconcileTimeout: 5m
    retryCount: 3

The repository structure uses cluster selectors for targeted configuration:

fleet-config/
├── clusters/
│   ├── cluster-selectors/
│   │   ├── on-prem.yaml          # Selects all on-prem clusters
│   │   ├── gcp.yaml              # Selects all GCP clusters
│   │   └── regulated.yaml        # Selects PII/financial clusters
│   ├── namespaces/
│   │   ├── base/                  # Applied everywhere
│   │   ├── on-prem/              # On-prem specific
│   │   └── gcp/                  # GCP specific
│   └── policies/
│       ├── base/                  # Universal policies
│       └── regulated/            # Extra policies for regulated clusters

Policy Controller: Guardrails That Actually Work

Policy Controller (built on OPA Gatekeeper) enforces constraints across the entire fleet. The critical difference from standalone Gatekeeper: policies are centrally managed and automatically distributed.

# policies/base/require-resource-limits.yaml
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredResources
metadata:
  name: require-cpu-memory-limits
  annotations:
    configsync.gke.io/cluster-name-selector: "*"
spec:
  match:
    kinds:
      - apiGroups: ["apps"]
        kinds: ["Deployment", "StatefulSet", "DaemonSet"]
    excludedNamespaces:
      - kube-system
      - gke-system
      - config-management-system
  parameters:
    requiredResources:
      - cpu
      - memory
    requireLimits: true
    requireRequests: true

For regulated clusters, additional constraints apply:

# policies/regulated/restrict-egress.yaml
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRestrictNetworkEgress
metadata:
  name: restrict-external-egress
  annotations:
    configsync.gke.io/cluster-name-selector: "regulated"
spec:
  match:
    kinds:
      - apiGroups: ["networking.k8s.io"]
        kinds: ["NetworkPolicy"]
  parameters:
    allowedExternalCIDRs:
      - "10.0.0.0/8"       # Internal network
      - "172.16.0.0/12"    # Cross-DC links
    deniedPorts:
      - 22
      - 3389
    requireExplicitEgress: true

Service Mesh: Anthos Service Mesh vs Istio

Anthos Service Mesh (ASM) is managed Istio with fleet-aware features. The killer feature: multi-cluster service discovery without manual ServiceEntry definitions.

Services in any fleet cluster automatically discover services in other clusters:

FeatureVanilla IstioASM
Multi-cluster discoveryManual ServiceEntryAutomatic
Certificate managementManual cert-managerManaged CA (Mesh CA)
Control plane operationsSelf-managedGoogle-managed
Cross-cluster load balancingComplex configLocality-aware automatic
ObservabilityPrometheus/GrafanaCloud Monitoring integrated

The cross-cluster traffic flow for a request from a GKE cluster to an on-prem service:

ASM Cross-Cluster Traffic Flow

Operational Metrics After 18 Months

MetricBefore AnthosAfter AnthosChange
Policy violations detected/monthUnknown340 (blocked)Visibility gained
Configuration drift incidents12/month0-100%
Mean time to deploy (fleet-wide)4 hours18 minutes-92%
Platform team size8 FTEs5 FTEs-37%
Cluster upgrade time2 days/cluster4 hours/cluster-83%
Security audit preparation3 weeks2 days-90%

What Doesn't Work Well

Bare metal provisioning is painful. GKE on Bare Metal requires specific hardware configurations, and the installation process has more failure modes than GKE on VMware. Budget 2-3 weeks for initial bare metal cluster bring-up.

Config Sync debugging is opaque. When reconciliation fails, error messages often point to symptoms rather than root causes. We built a custom dashboard that monitors RootSync and RepoSync status objects to surface issues faster.

ASM version upgrades require coordination. With a managed control plane, Google pushes updates on their schedule. We've had two incidents where ASM upgrades changed default mTLS behavior, briefly breaking cross-cluster communication.

Cost is substantial. Anthos licensing per cluster plus the GKE Enterprise tier adds $2,500-5,000/month per on-prem cluster depending on node count. The ROI comes from reduced operational headcount, not infrastructure savings.

Lessons Learned

  1. Start with Config Sync before ASM. Get GitOps working fleet-wide before adding service mesh complexity.
  2. Use namespace-level isolation, not cluster-level. Running separate clusters for each team doesn't scale. Anthos's multi-tenancy features (namespace quotas, network policies, RBAC) are mature enough for shared clusters.
  3. Invest in fleet-wide observability first. Cloud Monitoring with fleet-scoped metrics gives you the visibility to make confident changes.
  4. Treat Anthos as a platform, not a tool. The value compounds when Config Sync, Policy Controller, and ASM work together. Cherry-picking individual components delivers limited value.

Conclusion

Anthos eliminated configuration drift, reduced our platform team by 37%, and cut fleet-wide deployment time from 4 hours to 18 minutes. The investment is significant — both in licensing and initial setup — but for enterprises with genuine hybrid requirements, it's the only GCP-native solution that treats on-prem and cloud clusters as first-class citizens in a unified control plane.

The alternative is building the same capabilities from ArgoCD + OPA + Istio + custom glue. We estimated 6 months of platform engineering to replicate what Anthos provides out of the box. At our scale, buying beat building.

Comments

    No comments yet. Be the first to share your thoughts.