Kubernetes Resource Quotas and Limit Ranges for Multi-Tenant Clusters
Implementing resource quotas and limit ranges to prevent noisy neighbors, enforce fair sharing, and maintain cluster stability in multi-tenant Kubernetes

Introduction
In multi-tenant Kubernetes clusters, resource management is the difference between stable shared infrastructure and cascading failures. Without quotas and limits, a single runaway deployment can consume all cluster resources, starving other tenants. Kubernetes Resource Quotas and Limit Ranges work together to enforce fair resource distribution: quotas cap total namespace consumption while limit ranges constrain individual pod resources.
Based on operating clusters with 10-50 tenant namespaces, proper quota configuration reduces resource contention incidents by 90% and improves cluster utilization from 35% to 65% by enabling confident overcommitment.
Resource Quotas vs. Limit Ranges
| Feature | Resource Quota | Limit Range |
|---|---|---|
| Scope | Namespace total | Individual pod/container |
| Enforcement | Admission controller | Admission controller |
| What it caps | Total resource consumption | Per-unit resource boundaries |
| Default values | No | Yes (default requests/limits) |
| Effect of violation | Pod creation rejected | Pod creation rejected |
| Applies to | All resources in namespace | New pods only |
Resource Quota Configuration
Compute Resource Quotas
apiVersion: v1
kind: ResourceQuota
metadata:
name: compute-quota
namespace: team-backend
spec:
hard:
# CPU limits
requests.cpu: "20"
limits.cpu: "40"
# Memory limits
requests.memory: 40Gi
limits.memory: 80Gi
# Pod count
pods: "100"
# Persistent storage
requests.storage: 500Gi
persistentvolumeclaims: "20"
# GPU (if applicable)
requests.nvidia.com/gpu: "4"
Object Count Quotas
apiVersion: v1
kind: ResourceQuota
metadata:
name: object-quota
namespace: team-backend
spec:
hard:
configmaps: "50"
secrets: "50"
services: "20"
services.loadbalancers: "2"
services.nodeports: "5"
replicationcontrollers: "20"
resourcequotas: "5"
Storage Class-Specific Quotas
apiVersion: v1
kind: ResourceQuota
metadata:
name: storage-quota
namespace: team-backend
spec:
hard:
# Standard storage: 500Gi max
standard.storageclass.storage.k8s.io/requests.storage: 500Gi
standard.storageclass.storage.k8s.io/persistentvolumeclaims: "15"
# SSD storage: 100Gi max (more expensive)
ssd.storageclass.storage.k8s.io/requests.storage: 100Gi
ssd.storageclass.storage.k8s.io/persistentvolumeclaims: "5"
Limit Range Configuration
Default Container Limits
apiVersion: v1
kind: LimitRange
metadata:
name: container-limits
namespace: team-backend
spec:
limits:
- type: Container
default:
cpu: "500m"
memory: "512Mi"
defaultRequest:
cpu: "100m"
memory: "128Mi"
max:
cpu: "4"
memory: "8Gi"
min:
cpu: "50m"
memory: "64Mi"
maxLimitRequestRatio:
cpu: "10"
memory: "4"
- type: Pod
max:
cpu: "8"
memory: "16Gi"
min:
cpu: "50m"
memory: "64Mi"
- type: PersistentVolumeClaim
max:
storage: "100Gi"
min:
storage: "1Gi"
Understanding maxLimitRequestRatio
The maxLimitRequestRatio controls how much overcommitment is allowed per container:
| Ratio | Meaning | Risk Level | Use Case |
|---|---|---|---|
| 1 | Guaranteed (limit = request) | Lowest | Critical databases |
| 2 | 2x overcommit | Low | Production APIs |
| 4 | 4x overcommit | Medium | Background workers |
| 10 | 10x overcommit | High | Dev/test environments |
# Example: CPU ratio of 4 means if request is 250m, limit can be at most 1000m
# This prevents pods from requesting 100m CPU but setting limits at 8 CPU
Quota Allocation Strategy
Approach 1: Proportional Allocation
Allocate quotas proportional to team size or workload importance:
| Namespace | CPU Request | CPU Limit | Memory Request | Memory Limit | Pods |
|---|---|---|---|---|---|
| team-platform | 40 cores | 80 cores | 80Gi | 160Gi | 200 |
| team-backend | 20 cores | 40 cores | 40Gi | 80Gi | 100 |
| team-frontend | 10 cores | 20 cores | 20Gi | 40Gi | 50 |
| team-data | 30 cores | 60 cores | 120Gi | 240Gi | 80 |
| Cluster Total | 100 cores | 200 cores | 260Gi | 520Gi | 430 |
Approach 2: Tiered Service Classes
# Gold tier: production workloads with guaranteed resources
apiVersion: v1
kind: ResourceQuota
metadata:
name: gold-quota
namespace: payments-prod
spec:
hard:
requests.cpu: "30"
limits.cpu: "30" # No overcommit allowed
requests.memory: 60Gi
limits.memory: 60Gi # No overcommit allowed
scopeSelector:
matchExpressions:
- operator: In
scopeName: PriorityClass
values: ["high-priority"]
---
# Silver tier: standard workloads with moderate overcommit
apiVersion: v1
kind: ResourceQuota
metadata:
name: silver-quota
namespace: api-prod
spec:
hard:
requests.cpu: "20"
limits.cpu: "40" # 2x overcommit
requests.memory: 40Gi
limits.memory: 80Gi # 2x overcommit
Monitoring Quota Usage
Prometheus Metrics
# Current quota usage percentage per namespace
kube_resourcequota{type="used"} / kube_resourcequota{type="hard"} * 100
# Namespaces approaching quota (>80% used)
(
kube_resourcequota{type="used", resource="requests.cpu"}
/
kube_resourcequota{type="hard", resource="requests.cpu"}
) > 0.8
# Pods rejected due to quota
increase(kube_resourcequota_created_total{result="rejected"}[1h])
Alerting Rules
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: quota-alerts
spec:
groups:
- name: resource-quotas
rules:
- alert: NamespaceQuotaNearLimit
expr: |
kube_resourcequota{type="used"}
/ kube_resourcequota{type="hard"} > 0.85
for: 5m
labels:
severity: warning
annotations:
summary: "Namespace {{ $labels.namespace }} is at {{ $value | humanizePercentage }} of {{ $labels.resource }} quota"
- alert: NamespaceQuotaExhausted
expr: |
kube_resourcequota{type="used"}
/ kube_resourcequota{type="hard"} > 0.95
for: 2m
labels:
severity: critical
annotations:
summary: "Namespace {{ $labels.namespace }} has exhausted {{ $labels.resource }} quota"
Common Pitfalls and Solutions
Pitfall 1: Forgetting to Set Requests
When a ResourceQuota is set for requests.cpu, every pod MUST specify CPU requests. Pods without requests are rejected:
# Error: pods "my-pod" is forbidden: failed quota: compute-quota:
# must specify requests.cpu, requests.memory
Solution: Always pair quotas with LimitRanges that set defaults:
# The LimitRange default ensures pods without explicit requests
# still get admitted with sensible defaults
spec:
limits:
- type: Container
defaultRequest:
cpu: "100m"
memory: "128Mi"
Pitfall 2: Over-Constraining Batch Jobs
# Separate quota scope for batch workloads
apiVersion: v1
kind: ResourceQuota
metadata:
name: batch-quota
namespace: team-data
spec:
hard:
requests.cpu: "50"
limits.cpu: "100"
scopeSelector:
matchExpressions:
- operator: In
scopeName: PriorityClass
values: ["batch"]
Automation: Namespace Provisioning
#!/bin/bash
# Provision a new tenant namespace with quotas and limits
NAMESPACE=$1
CPU_REQUESTS=$2
MEMORY_REQUESTS=$3
kubectl create namespace "$NAMESPACE"
kubectl apply -f - <<EOF
apiVersion: v1
kind: ResourceQuota
metadata:
name: compute-quota
namespace: $NAMESPACE
spec:
hard:
requests.cpu: "${CPU_REQUESTS}"
limits.cpu: "$((CPU_REQUESTS * 2))"
requests.memory: "${MEMORY_REQUESTS}Gi"
limits.memory: "$((MEMORY_REQUESTS * 2))Gi"
pods: "100"
---
apiVersion: v1
kind: LimitRange
metadata:
name: default-limits
namespace: $NAMESPACE
spec:
limits:
- type: Container
default:
cpu: "500m"
memory: "512Mi"
defaultRequest:
cpu: "100m"
memory: "128Mi"
max:
cpu: "4"
memory: "8Gi"
min:
cpu: "50m"
memory: "64Mi"
EOF
echo "Provisioned namespace $NAMESPACE with ${CPU_REQUESTS} CPU, ${MEMORY_REQUESTS}Gi memory quota"
Key Takeaways
- Always deploy LimitRanges alongside ResourceQuotas to provide default requests and limits; otherwise pods without explicit resource specifications will be rejected.
- Use maxLimitRequestRatio to control overcommitment at the container level, with ratios of 2-4x for production and up to 10x for development namespaces.
- Monitor quota utilization with Prometheus and alert at 85% to give teams time to request increases or optimize before hitting hard limits.
- Separate quotas by priority class to ensure critical workloads have guaranteed capacity while batch and development workloads share remaining resources.
- Set storage quotas per storage class to prevent expensive SSD storage from being consumed by workloads that could use standard storage.
- Proper quota implementation improves cluster utilization from 35% to 65% by enabling confident overcommitment with guardrails.
- Automate namespace provisioning with standardized quota templates to ensure consistent governance as teams are onboarded.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.