Every resilience conversation I have with customers ends in the same place. The fear is not the failure. The fear is the blast radius. A bad deploy, a gray failure in one AZ, a noisy tenant, an operator typo at 2 AM: the incident itself is survivable. What keeps architects up is that it hits everyone at once.
So I spent time with the best public material on cell-based architecture: Airbnb's zonal migration talk from KubeCon EU, Slack's engineering write-up of their AZ drain, New Relic's cell fleet program, Booking.com's vCluster cells, Neon's hard lessons from going cellular under pressure, and WSO2's take on cells for AI agents. I also worked through the AWS reference implementation for cell-based architecture on EKS.
##
My read, after all of it: cells are the mechanism that turns an incident's blast radius from "all of our customers" into "one Nth of our customers." A cell is a full vertical slice of the system, fixed in size, fully independent: its own services, data stores, networking, and security boundary. You scale by adding cells, never by enlarging them. The pattern has three parts: the cell, a thin stateless router that maps a partition key to a cell, and a control plane that stays out of the request path entirely.
This guide is what I would hand a customer architect who asks how to build that on EKS: when the investment is justified, the exact reference shape, and the mistakes to avoid.When cells are worth it, and when they are not
Airbnb’s lead architect opens his KubeCon talk with the rule: 99 times out of 100, do regional. A regional cluster is easier to reason about: one control plane, one deployment with HPA and pod topology spread. Cells are the exception, not the default.
Build cells when at least two of these are true:
A single-tenant-impacting failure is a business-level event (regulated data, SLA penalties, revenue on the line)
You have crossed, or are on a growth curve toward, concrete platform ceilings: ~10,000 active pods per cluster, the EKS etcd 8 GB ceiling, Kubernetes API rate limiting (Neon hit all of these)
Your incidents are dominated by gray failures, bad deploys, or operator errors where recovery speed matters more than prevention (Airbnb’s thesis: regional clusters “don’t protect us from ourselves”)
You have platform-engineering maturity to automate provisioning, wave deployments, and fleet health, because per-cell operational overhead multiplies
Do not build cells when:
Your regional clusters are nowhere near their ceilings and your incidents are ordinary (Booking.com’s warning: cells are “comparatively costlier,” and she uses them only for payment workloads)
Your data model cannot be partitioned without synchronous cross-cell calls (a cell with sync cross-cell dependencies is a distributed monolith with extra cost)
You need it this quarter. Neon cellularized under pressure and had more incidents in two months than in the entire prior year. My take: cells are a proactive investment or an emergency liability. There is no middle option.The reference architecture on EKS
The AWS-sanctioned shape comes from the AWS Solutions Architects’ reference implementation (guidance-for-cell-based-architecture-for-amazon-eks, MIT-0 Terraform).
Cell = one EKS cluster confined to a single Availability Zone. Not a namespace, not a node group. The cluster boundary is the blast-radius boundary: its own control plane, compute, ALBs, IAM roles, and data stores. Multiple cells per AZ are allowed; cells spread across AZs so an AZ failure takes out only its cells.
[DIAGRAM 1]
Five hard rules inside each cell:
Cross-zone load balancing stays off. The AWS Load Balancer Controller configures per-cell ALBs with cross-zone LB disabled. This preserves AZ isolation and removes inter-AZ data transfer for chatty workloads, which is also a cost win.
Pods are pinned to the cell’s AZ. Node selectors (or AZ-scoped Karpenter provisioners) keep scheduling inside the cell. Pods must never drift across the cell boundar``
Traffic enters through Route 53 weighted routing. Weighted records (reference split 33/33/34 across three cells) with health checks for automatic failover, plus Amazon Application Recovery Controller zonal shift as the break-glass. Per-cell DNS records (cell1/cell2/cell3.example.com) allow direct cell access for debugging.
Karpenter handles node scaling. Managed node groups for baseline and system pods (CoreDNS, LB controller), Karpenter for dynamic scaling beyond. Karpenter also fixes the insufficient-capacity class of failures that static node groups leave you exposed to.
Least-privilege IAM per cell, with EKS Pod Identity. Each cell gets its own roles; never share a role across cells. The trust policy sketch for a per-cell pod role:```
{
“Version”: “2012-10-17”,
“Statement”: [{
“Effect”: “Allow”,
“Principal”: { “Service”: “pods.eks.amazonaws.com” },
“Action”: [”sts:AssumeRole”, “sts:TagSession”]
}]
}Then scope the role’s permissions to resources inside that cell’s account or VPC only. A cell’s compromised credential must not reach another cell’s resources.
The stack you can copy directly: AWS’s three-layer isolation template (AZ → Region → Account), the thin stateless router (hash function → DynamoDB mapping → Route 53 entry from a networking shared-services account), and HPA-over-VPA as the hyperscale autoscaler, because adding replicas also adds resiliency. Full working Terraform is in the guidance repo above; treat it as the starting point, not the finished product.
The reference’s one compromise to fix: the README permits multi-AZ RDS/DynamoDB/ElastiCache shared across cells for persistence. That is shared data fate wearing a high-availability costume. Prefer per-cell data stores (see Data isolation below); use the shared variant only with an explicit, time-boxed migration plan to per-cell storage.The router: the one component allowed to be shared
Every source treats the router as the single shared component, and invests its paranoia there. If the router fails, N cells become one logical cell.
[DIAGRAM 2]
Design rules, no exceptions:
Thin and stateless. It maps the partition key to a cell and nothing else. No auth, no rate limiting, no business logic (Hyperforce’s rule: the router stays “dumb” or it becomes the new single point of failure).
In a separate failure domain from the cells. Slack’s drains are driven from edge load balancers in other regions with a regionally replicated control plane, so draining works when the target AZ is completely offline. Your Route 53 layer and mapping store must survive the loss of any cell.
Static stability: it works with the control plane down. Cache the routing map locally at the router; use DNS for cell discovery rather than database lookups on the hot path. Test this explicitly: kill the mapping store and confirm requests still route.
Modeled on Slack’s pattern: the drain is reweighting, not reconfiguration. Two stock mechanisms do the whole job: weighted target clusters plus dynamic weight assignment. The reweight propagates in seconds, completes in-flight requests gracefully, and gives 1% granularity so operators can experimentally drain during an incident and undrain if it doesn’t help.
For EKS, the concrete recipe is: Ro
{
“Name”: “api.example.com”,
“Type”: “A”,
“SetIdentifier”: “cell-1”,
“Weight”: “33”,
“AliasTarget”: {
“DNSName”: “k8s-cell1-us-east-1a.elb.amazonaws.com”,
“EvaluateTargetHealth”: true
}
}eat with SetIdentifier cell-2 / Weight 33 and cell-3 / Weight 34. A drain is a weight change to 0 on the failing cell’s record, propagated by DNS.
4. Partitioning: the decision you will not get to redo
Every source agrees: the partition key is the most critical decision and the hardest to change. Tenant ID, user geography, or user ID are the usual candidates. It determines data gravity, latency, and load balance for the life of the system. Choose it with the data architects in the room, not after the cluster design is done.
Two cell models, pick one deliberately:
Per-AZ cells (AWS guidance, Slack, Airbnb): the cell boundary is the availability zone. Best when the threat is infrastructure failure and gray failures. Simplest to reason about, aligns with AWS’s own failure domains.
Per-tenant (or per-shard) cells (New Relic, Notion-style, GitHub-style): the cell boundary is the customer set. Best when the threat is noisy neighbors and blast radius must be expressed in customers, not zones. Harder routing, stronger isolation guarantees per customer.
Shuffle sharding: the zero-cost noisy-neighbor layer. On top of cells, assign each tenant a small pseudo-random combination of cells from the pool instead of one cell. The math: 8 cells in virtual shards of 2 gives C(8,2) = 28 unique combinations, so a misbehaving tenant degrades 1/28th of the fleet instead of a whole cell. Route 53 runs this in production: 2,048 virtual name servers, customers assigned shuffle shards of 4, roughly 730 billion possible shards, no two customer domains sharing more than 2 servers. Effectiveness improves with scale, which makes it the rare technique that gets better as you grow. Reserve it for multi-tenant SaaS where tenant behavior, not cell failure, is the threat; it adds router complexity you do not need otherwise.
[DIAGRAM 3]
5. Deployments: waves, or a global single point of failure
Deploying to all cells at once turns your deployment pipeline into the blast radius you built cells to avoid. The wave discipline is unanimous across every source.
The required wave schedule:
Canary cell first (lowest-traffic cell), automated health checks, soak 30 to 60 minutes
10% of cells, with canary analysis on error rate and latency
25% of cells
100%
Automated rollback is non-negotiable at every staData isolation: enforced by tooling, not review boards
Rule: every cell owns its data stores. Databases, caches, queues, all cell-local. Cross-cell needs travel through asynchronous replication or messaging (Kafka or equivalent), never synchronous calls on the request hot path. This is unanimous across sources, and the failure mode has a name: the leaky abstraction, a “quick” sync cross-cell API call, a shared read replica for analytics, a central cache. Each one erodes fault isolation until you are back to a distributed monolith, paying cell prices for regional reliability.
[DIAGRAM 6]
Enforce it mechanically:
Kubernetes NetworkPolicies that deny inter-cell egress by default in every cell:
[YAML CODE BLOCK]
Then add explicit allow rules only for the async messaging endpoints and the cell’s own stores. A new dependency requires a policy change, which is reviewable and auditable.
CI static analysis that flags cross-cell database connections before merge.
Paved roads over review boards: the platform team supplies managed async messaging, Terraform cell-stamping modules, and CI/CD templates with waves baked in. Compliance must be the path of least resistance, not a ticket queue.
The one legitimate exception is genuinely global data (system configuration, identity), which needs async replication with explicit consistency models (CQRS/event sourcing), never synchronous sharing. Name the exception, bound it, and keep it out of the request hot path.8. Observability: per cell, or you are flying blind
Global aggregates hide cell-level degradation. A fleet-wide p99 of 120 ms can mask one cell at 400 ms. Every source that ran cells in production converged on the same rule: each cell gets its own dashboards, its own alerts, and its own error-budget tracking, with a fleet-level aggregated view on top.
Minimum per-cell instrumentation:
Request rate, error rate, and latency percentiles by cell
Cell router weights and traffic distribution (so a stuck drain is visible)
Per-cell saturation signals: CPU/memory, pod churn, Karpenter provisioning latency, etcd metrics
Per-cell error budgets with alerts that fire on the cell, not the fleet
Airbnb rebuilt their developer tooling to operate at whole-unit or per-cell granularity: scale by cell, logs per cell, metrics and alerting by cell and by AZ instead of region-wide. If your current dashboards cannot answer “which cell is unhealthy right now,” the observability work is part of the cell migration, not an afterthought.
9. Failure-mode checklist and acceptance tests
The acceptance test for the whole architecture: deliberately kill a cell in production (with approvals) and verify that only 1/N of users are affected, with automatic evacuation. If the failure escapes the cell boundary, your isolation is fictional. Hyperforce runs exactly this test; so should you.
Checklist of failure modes to design against:
Bad deploy: contained by wave schedule. Verify: a canary-cell failure stops the wave automatically.
Gray failure in one AZ: contained by the drain. Verify: quarterly production drain drills.
Noisy tenant: contained by shuffle sharding. Verify: load-test a poisoned tenant and measure impact at 1/28th of the fleet, not a whole cell.
Operator error: contained by per-cell IAM and per-cell tooling scope. Verify: a cell-scoped role cannot touch another cell’s resources.
Leaky abstraction: contained by default-deny NetworkPolicy and CI checks. Verify: chaos experiments that attempt cross-cell sync calls fail closed.
Router failure: contained by static stability. Verify: kill the mapping store; requests keep routing.
Cell failure: contained by design; verify with the kill-a-cell test above.
Ceilings to respect:
Cells do not make Kubernetes ceilings disappear, they keep one ceiling from being global. Neon measured degradation beyond 10,000 concurrent databases per cluster from the EKS etcd 8 GB ceiling, network confiCell drain runbook: the payoff of the whole architecture
Slack’s drain button is the pattern to copy. Its four design goals are your acceptance criteria:
Remove traffic from a cell within 5 minutes. Their 99.99% SLA allows under an hour of downtime per year; the drain must be fast enough to spend almost none of it.
Zero user-visible errors during the drain. In-flight requests complete gracefully. The drain is a generic mitigation operators reach for before root cause is known, so it must be safe to use experimentally.
Incremental in both directions. Drain and undrain at fine granularity (Slack: 1% on restore) so recovery can be verified before full traffic returns.
Operable from outside the cell. The drain mechanism must not depend on any resource inside the drained AZ. If your drain needs the AZ to be healthy, it is not a drain.
[DIAGRAM 5]
Rehearse it in production. Slack exercises the drain regularly against real services, which is chaos engineering applied to the mitigation tool itself. A drain that has only ever run in staging will surprise you in the incident it was built for. Schedule quarterly drain drills; the drill result (minutes to drain, errors observed) is a reliability metric, reported like any other.. Migration path: go cellular proactively
Neon’s story is the warning label: implementing cells under incident pressure, while keeping the service running, produced more incidents in two months than the entire previous year. The investment case is made while the system is healthy, keyed to the cost of one total outage.
Airbnb’s automation lessons, in order:
PoC first, and let it embarrass you. Two services in three months extrapolated to “1,000 migrations in 125 years.” That number is what forced the automation investment. Budget for it.
Make the migration a normal deployment. Their platform refactor (hardcoded cluster context replaced with a cellset field) meant each service migration surfaced to the owning team as one PR review and one normal deploy, about half a day’s work. The platform team absorbed the complexity; service teams did not.
Automate the batch. A batch job cloned each service’s source, ran the config-migration tooling, looked up the owning team and on-call in the service catalog, created the PR, and opened a Slack thread tagging the on-call. They did 10 to 20 critical services in parallel, 30 to 40 for less critical, roughly 3,000 PRs total.
Phase per service: non-prod, then canary, then production. Non-prod caught sticky sessions and hardcoded cluster assumptions before production ever saw them.
Expect the support burden to be underestimated anyway. Their words, despite the self-service design.
Set expectations honestly:
New Relic’s cellular program is a multi-year effort (started 2020, still evolving cell types in 2025), and their hardest ongoing pain is Cluster API version skew across cloud providers. If you go multi-cloud, that pain is yours too. The control-plane cell (their “C&C cell” with cell-modeling CRDs) is the pattern to copy for fleet lifecycle, but budget real engineering for it.
11. Cost notes
The AWS reference implementation prices a 3-cell deployment at roughly $800 to $1,000 per month: fixed costs around $722 (3 EKS clusters at $219, baseline nodes, 3 ALBs, NAT gateway, EBS, Route 53 hosted zone) plus $63 to $315+ variable (ALB LCU processing, NAT data processing, CloudWatch, DNS queries). Levers: Karpenter provisioner tuning, right-sizing baseline node groups, Spot via Karpenter, Savings Plans for the baseline.
The honest framing for the business case: duplication raises baseline cost, always. The model pays off at scale through linear cost growth instead of step-function re-architecture, avoided global outages, and reduced need for expensive vertical scaling. Hyperforce’s framing is the one to use with leadership: cells are not a best practice, they are a scalin## 12. Sources
- Slack engineering, “Slack’s Migration to a Cellular Architecture” (note: the promised follow-up post never shipped) → https://slack.engineering/slacks-migration-to-a-cellular-architecture/
- Slack engineering, “Traffic 101: Packets Mostly Flow” → https://slack.engineering/traffic-101-packets-mostly-flow/
- Airbnb, KubeCon EU 2026, “1000 Services, 1 Year, 0 Downtime: Airbnb’s Zonal Cluster Migration” (Sunny Beatteay) →
- AWS Architecture Blog, “Journey to Cloud-Native Architecture Series #7: Using Containers and Cell-Based Design” → https://aws.amazon.com/blogs/architecture/journey-to-cloud-native-architecture-series-7-using-containers-and-cell-based-design-for-higher-resiliency-and-efficiency/
- AWS Guidance for Cell-Based Architecture on Amazon EKS (Terraform reference, arsa-codes) → https://github.com/arsa-codes/guidance-for-cell-based-architecture-for-amazon-eks
- New Relic, KubeCon EU 2025, “Resilient Multi-Cloud Strategies” (Cluster API cell program) →
- Booking.com + Loft Labs, KubeCon India 2024, “Cell-Based Kubernetes - The Secret to Scalable, Repeatable and Resilient Cloud Architecture” →
- WSO2, KubeCon EU 2026, “Zero Trust for Autonomous Agents: Isolating AI Workloads on Kubernetes” (cell-based architecture applied to agent isolation) →
- University of Edinburgh + WSO2, KubeCon NA 2025, “People-first Path To Cell-based Architecture” → sched page https://kccncna2025.sched.com/event/27FY1/people-first-path-to-cell-based-architecture-martin-jones-university-of-edinburgh-asanka-abeysinghe-wso2
- Omnistrate, “Moving to Cell-Based Architecture: Lessons from Neon” → https://blog.omnistrate.com/posts/moving-to-cell-based-architecture-lessons-from-neon
- AWS Builders’ Library, “Workload isolation using shuffle-sharding” → https://aws.amazon.com/builders-library/workload-isolation-using-shuffle-sharding/
- Salesforce Hyperforce cellular analysis (Mayank Raj) → https://github.com/rajmayank/mayankraj.com/blob/HEAD/content/blog/cell-based-architecture-blast-radius-containment.md
- developers.dev, “The Architect’s Playbook for Cell-Based Architecture” → https://www.developers.dev/tech-talk/the-architect-s-playbook-for-cell-based-architecture-from-principles-to-production.html
- vCluster KubeCon India 2024 recap (vCluster cell implementation) → https://www.vcluster.com/blog/kubecon-india-2024-recap
The shape keeps getting rediscovered: full-stack isolation, a thin router, waves everywhere. The teams that get it right treat the drain button as the product, not the cell. If you can remove a third of your fleet from traffic in five minutes with zero user-visible errors, the rest of the architecture is details. That is the test I would run first.

