Production GraphQL with AWS AppSync: Caching, Auth, and Real-Time Subscriptions
Lessons from running AppSync at scale — implementing multi-layer caching, fine-grained authorization, and WebSocket subscriptions serving 50K concurrent users.

Why AppSync Over Self-Hosted GraphQL
Every team running GraphQL on AWS faces the same fork in the road: self-host with Apollo Server on ECS/Lambda, or go managed with AppSync. After operating both patterns at scale, I'll make the case that AppSync wins for 80% of production use cases — not because it's simpler (it's not, initially), but because it eliminates entire categories of infrastructure you'd otherwise build yourself.
The problem with self-hosted GraphQL isn't the happy path — it's everything else. Connection pooling, subscription fan-out, resolver-level caching, per-field authorization, DDoS protection on the WebSocket endpoint. AppSync handles all of these as managed capabilities.
Architecture: Multi-Layer Caching
AppSync offers two caching layers that most teams under-utilize:
- Server-side caching — A managed ElastiCache (Memcached) cluster behind your API that caches resolver responses.
- Client-side caching — Response headers that instruct clients to cache results locally.
The key architectural decision is cache key design. AppSync lets you configure cache keys at the resolver level using the request context:
// VTL resolver for getCachedProduct
// Cache key: product:{id}:{userRole}
// Different roles see different pricing data
export function request(ctx) {
const cacheKey = `product:${ctx.args.id}:${ctx.identity.claims['custom:role']}`;
return {
operation: 'GetItem',
key: util.dynamodb.toMapValues({ PK: `PRODUCT#${ctx.args.id}` }),
consistentRead: false,
};
}
export function response(ctx) {
if (ctx.error) {
util.error(ctx.error.message, ctx.error.type);
}
const product = ctx.result;
// Strip wholesale pricing for non-dealer roles
if (ctx.identity.claims['custom:role'] !== 'dealer') {
delete product.wholesalePrice;
delete product.dealerMargin;
}
return product;
}
For our product catalog API, server-side caching reduced DynamoDB read costs by 73% while keeping p99 latency under 45ms.
Cache Invalidation Strategy
AppSync doesn't have built-in cache invalidation hooks, so we implemented a pattern using DynamoDB Streams → Lambda → AppSync cache flush:
import { AppSyncClient, FlushApiCacheCommand } from '@aws-sdk/client-appsync';
const appSync = new AppSyncClient({});
export async function handler(event: DynamoDBStreamEvent) {
const modifiedProducts = event.Records
.filter(r => r.eventName === 'MODIFY' || r.eventName === 'REMOVE')
.filter(r => r.dynamodb?.Keys?.PK?.S?.startsWith('PRODUCT#'))
.map(r => r.dynamodb?.Keys?.PK?.S?.replace('PRODUCT#', ''));
if (modifiedProducts.length > 0) {
// For targeted invalidation, use type-level flush
await appSync.send(new FlushApiCacheCommand({
apiId: process.env.APPSYNC_API_ID!,
}));
console.log(`Cache flushed for ${modifiedProducts.length} product updates`);
}
}
The trade-off: AppSync only supports full cache flush or type-level flush — no individual key invalidation. For high-write scenarios, we use short TTLs (60-120s) instead of event-driven invalidation to avoid flush storms.
Fine-Grained Authorization
AppSync supports five authorization modes simultaneously. Our production pattern uses Cognito for user identity combined with Lambda authorizers for service-to-service calls:
// Lambda authorizer for AppSync
// Validates service JWTs and returns field-level permissions
import { verify } from 'jsonwebtoken';
interface AppSyncAuthEvent {
authorizationToken: string;
requestContext: {
apiId: string;
accountId: string;
queryString: string;
variables: Record<string, unknown>;
};
}
export async function handler(event: AppSyncAuthEvent) {
const token = event.authorizationToken.replace('Bearer ', '');
try {
const decoded = verify(token, process.env.SERVICE_JWT_SECRET!) as ServiceToken;
return {
isAuthorized: true,
resolverContext: {
serviceId: decoded.serviceId,
permissions: decoded.permissions.join(','),
},
deniedFields: getDeniedFields(decoded.permissions),
ttlOverride: 300, // Cache auth decision for 5 minutes
};
} catch (error) {
return { isAuthorized: false };
}
}
function getDeniedFields(permissions: string[]): string[] {
const allSensitiveFields = [
'arn:aws:appsync:us-east-1:123456789:apis/xxx/types/User/fields/email',
'arn:aws:appsync:us-east-1:123456789:apis/xxx/types/User/fields/phoneNumber',
'arn:aws:appsync:us-east-1:123456789:apis/xxx/types/Order/fields/paymentMethod',
];
if (permissions.includes('pii:read')) {
return []; // Full access
}
return allSensitiveFields;
}
The deniedFields array is AppSync's most powerful authorization primitive — it lets you strip sensitive fields from responses without modifying resolver logic. The resolver always fetches the full object; AppSync removes fields before they reach the client.
Real-Time Subscriptions at Scale
AppSync WebSocket subscriptions handle the connection management, fan-out, and authentication that you'd otherwise build on API Gateway WebSocket + DynamoDB + Lambda. But there are production gotchas:
Connection Limits and Fan-Out
AppSync supports 100 subscriptions per connection and scales horizontally for concurrent connections. Our dashboard serves 50K concurrent users with the following subscription schema:
type Subscription {
onOrderStatusChanged(customerId: ID!): OrderStatusEvent
@aws_subscribe(mutations: ["updateOrderStatus"])
onInventoryAlert(warehouseId: ID!): InventoryAlert
@aws_subscribe(mutations: ["createInventoryAlert"])
onMetricUpdate(dashboardId: ID!): MetricDataPoint
@aws_subscribe(mutations: ["publishMetric"])
}
The critical pattern: always filter subscriptions by a user-scoped argument. Without the customerId filter, every subscriber receives every mutation event — a recipe for bandwidth explosions.
Benchmarks: AppSync vs. Self-Hosted Apollo on Lambda
We ran a parallel comparison over 30 days with equivalent query patterns:
| Metric | Self-Hosted Apollo (Lambda) | AppSync | Winner |
|---|---|---|---|
| p50 Latency | 34ms | 28ms | AppSync |
| p99 Latency | 890ms (cold starts) | 62ms | AppSync |
| Monthly cost (2M queries/day) | $2,340 | $1,680 | AppSync (-28%) |
| WebSocket monthly cost (50K concurrent) | $4,100 (API GW + DDB + Lambda) | $890 | AppSync (-78%) |
| Time to implement auth | 3 weeks | 3 days | AppSync |
| Schema-first DX | Excellent | Good (VTL learning curve) | Apollo |
The p99 improvement is the killer: AppSync resolvers don't have cold starts. Your data sources might (if they're Lambda), but the GraphQL layer itself is always warm.
Production Pitfalls and Mitigations
After 18 months in production, these are the issues that bit us:
- VTL resolver debugging is painful — Use JavaScript resolvers (APPSYNC_JS runtime) instead. The DX is dramatically better and performance is equivalent.
- Subscription filter limitations — Enhanced subscription filters (released 2024) finally support complex filtering, but the syntax is non-obvious.
- Resolver pipeline limits — Max 10 functions per pipeline resolver. If you hit this, your resolver is doing too much.
- Schema size limits — 1MB schema limit. Use schema stitching via CDK if you approach this.
Conclusion
AppSync isn't the right choice for every GraphQL API — if you need full control over the execution engine, federation across multiple subgraphs, or have a team deeply invested in the Apollo ecosystem, self-hosting makes sense.
But for teams that want managed WebSocket infrastructure, built-in caching, per-field authorization, and zero cold-start resolver execution — AppSync eliminates months of undifferentiated infrastructure work. The 28% cost reduction and 78% savings on real-time infrastructure aren't trivial either.
Start with JavaScript resolvers (not VTL), invest in cache key design early, and always scope your subscriptions. AppSync rewards deliberate architecture with operational simplicity.
Recommended reading

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Comments
No comments yet. Be the first to share your thoughts.