Managing Hybrid Workloads with Anthos: From On-Prem Kubernetes to GCP at Enterprise Scale
Practical guide to running Anthos across on-premises data centers and GCP, covering fleet management, policy enforcement, and service mesh at scale.

Running Kubernetes in two places is easy. Running it consistently, securely, and observably across on-premises data centers and GCP without doubling your platform team — that's the actual problem Anthos solves. After 18 months managing 23 clusters across three data centers and two GCP regions, here's what works, what doesn't, and what the documentation won't tell you.
The Problem Statement
Our enterprise had a common constraint: regulatory requirements mandating certain workloads stay on-premises, combined with a strategic push toward GCP for everything else. The result was two separate Kubernetes ecosystems with different tooling, policies, and operational procedures.
The pain points before Anthos:
- Policy drift: Security policies applied in GCP didn't match on-prem configurations
- Deployment inconsistency: CI/CD pipelines forked into separate paths per environment
- Observability gaps: No unified view across all clusters
- Certificate management: Different PKI infrastructure per environment
- Team overhead: Separate SRE runbooks for each environment
Anthos Architecture: What You Actually Deploy
Anthos isn't a single product — it's a control plane that layers fleet management, policy enforcement, service mesh, and configuration management across heterogeneous clusters.
Fleet Registration
Every cluster — GKE on GCP, GKE on-prem (VMware or bare metal), or attached third-party clusters — registers to a fleet:
# fleet-membership.yaml
apiVersion: gkehub.googleapis.com/v1
kind: Membership
metadata:
name: dc-east-prod-01
spec:
endpoint:
kubernetesResource:
membershipCrSpec:
# Connect agent auto-installed
connectAgent:
namespace: gke-connect
resourceOptions:
connectVersion: "1.14"
authority:
issuer: "https://container.googleapis.com/v1/projects/my-project/locations/us-east1/memberships/dc-east-prod-01"
The Connect Agent establishes an outbound-only tunnel from your on-prem cluster to GCP's fleet management API. No inbound firewall rules required — this was a key security requirement for our network team.
Fleet Topology
Our fleet after migration:
| Cluster | Location | Type | Nodes | Workload Type |
|---|---|---|---|---|
| dc-east-prod-01 | Virginia DC | Bare metal | 48 | Regulated financial |
| dc-east-prod-02 | Virginia DC | Bare metal | 32 | PII processing |
| dc-west-prod-01 | Oregon DC | VMware | 24 | Internal services |
| gke-east-prod | us-east4 | GKE Autopilot | Auto | Public APIs |
| gke-central-prod | us-central1 | GKE Standard | 64 | ML workloads |
| gke-west-prod | us-west1 | GKE Autopilot | Auto | CDN origin |
Config Sync: GitOps at Fleet Scale
Config Sync is Anthos's GitOps engine. It watches a Git repository and reconciles cluster state — but unlike vanilla ArgoCD, it operates fleet-wide with namespace-scoped inheritance.
# config-sync/rootsync.yaml
apiVersion: configsync.gke.io/v1beta1
kind: RootSync
metadata:
name: root-sync
namespace: config-management-system
spec:
sourceFormat: unstructured
git:
repo: "https://github.com/myorg/fleet-config"
branch: main
dir: "clusters"
auth: gcpserviceaccount
gcpServiceAccountEmail: config-sync@my-project.iam.gserviceaccount.com
override:
reconcileTimeout: 5m
retryCount: 3
The repository structure uses cluster selectors for targeted configuration:
fleet-config/
├── clusters/
│ ├── cluster-selectors/
│ │ ├── on-prem.yaml # Selects all on-prem clusters
│ │ ├── gcp.yaml # Selects all GCP clusters
│ │ └── regulated.yaml # Selects PII/financial clusters
│ ├── namespaces/
│ │ ├── base/ # Applied everywhere
│ │ ├── on-prem/ # On-prem specific
│ │ └── gcp/ # GCP specific
│ └── policies/
│ ├── base/ # Universal policies
│ └── regulated/ # Extra policies for regulated clusters
Policy Controller: Guardrails That Actually Work
Policy Controller (built on OPA Gatekeeper) enforces constraints across the entire fleet. The critical difference from standalone Gatekeeper: policies are centrally managed and automatically distributed.
# policies/base/require-resource-limits.yaml
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredResources
metadata:
name: require-cpu-memory-limits
annotations:
configsync.gke.io/cluster-name-selector: "*"
spec:
match:
kinds:
- apiGroups: ["apps"]
kinds: ["Deployment", "StatefulSet", "DaemonSet"]
excludedNamespaces:
- kube-system
- gke-system
- config-management-system
parameters:
requiredResources:
- cpu
- memory
requireLimits: true
requireRequests: true
For regulated clusters, additional constraints apply:
# policies/regulated/restrict-egress.yaml
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRestrictNetworkEgress
metadata:
name: restrict-external-egress
annotations:
configsync.gke.io/cluster-name-selector: "regulated"
spec:
match:
kinds:
- apiGroups: ["networking.k8s.io"]
kinds: ["NetworkPolicy"]
parameters:
allowedExternalCIDRs:
- "10.0.0.0/8" # Internal network
- "172.16.0.0/12" # Cross-DC links
deniedPorts:
- 22
- 3389
requireExplicitEgress: true
Service Mesh: Anthos Service Mesh vs Istio
Anthos Service Mesh (ASM) is managed Istio with fleet-aware features. The killer feature: multi-cluster service discovery without manual ServiceEntry definitions.
Services in any fleet cluster automatically discover services in other clusters:
| Feature | Vanilla Istio | ASM |
|---|---|---|
| Multi-cluster discovery | Manual ServiceEntry | Automatic |
| Certificate management | Manual cert-manager | Managed CA (Mesh CA) |
| Control plane operations | Self-managed | Google-managed |
| Cross-cluster load balancing | Complex config | Locality-aware automatic |
| Observability | Prometheus/Grafana | Cloud Monitoring integrated |
The cross-cluster traffic flow for a request from a GKE cluster to an on-prem service:
Operational Metrics After 18 Months
| Metric | Before Anthos | After Anthos | Change |
|---|---|---|---|
| Policy violations detected/month | Unknown | 340 (blocked) | Visibility gained |
| Configuration drift incidents | 12/month | 0 | -100% |
| Mean time to deploy (fleet-wide) | 4 hours | 18 minutes | -92% |
| Platform team size | 8 FTEs | 5 FTEs | -37% |
| Cluster upgrade time | 2 days/cluster | 4 hours/cluster | -83% |
| Security audit preparation | 3 weeks | 2 days | -90% |
What Doesn't Work Well
Bare metal provisioning is painful. GKE on Bare Metal requires specific hardware configurations, and the installation process has more failure modes than GKE on VMware. Budget 2-3 weeks for initial bare metal cluster bring-up.
Config Sync debugging is opaque. When reconciliation fails, error messages often point to symptoms rather than root causes. We built a custom dashboard that monitors RootSync and RepoSync status objects to surface issues faster.
ASM version upgrades require coordination. With a managed control plane, Google pushes updates on their schedule. We've had two incidents where ASM upgrades changed default mTLS behavior, briefly breaking cross-cluster communication.
Cost is substantial. Anthos licensing per cluster plus the GKE Enterprise tier adds $2,500-5,000/month per on-prem cluster depending on node count. The ROI comes from reduced operational headcount, not infrastructure savings.
Lessons Learned
- Start with Config Sync before ASM. Get GitOps working fleet-wide before adding service mesh complexity.
- Use namespace-level isolation, not cluster-level. Running separate clusters for each team doesn't scale. Anthos's multi-tenancy features (namespace quotas, network policies, RBAC) are mature enough for shared clusters.
- Invest in fleet-wide observability first. Cloud Monitoring with fleet-scoped metrics gives you the visibility to make confident changes.
- Treat Anthos as a platform, not a tool. The value compounds when Config Sync, Policy Controller, and ASM work together. Cherry-picking individual components delivers limited value.
Conclusion
Anthos eliminated configuration drift, reduced our platform team by 37%, and cut fleet-wide deployment time from 4 hours to 18 minutes. The investment is significant — both in licensing and initial setup — but for enterprises with genuine hybrid requirements, it's the only GCP-native solution that treats on-prem and cloud clusters as first-class citizens in a unified control plane.
The alternative is building the same capabilities from ArgoCD + OPA + Istio + custom glue. We estimated 6 months of platform engineering to replicate what Anthos provides out of the box. At our scale, buying beat building.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.