Session
Weighted Sharding: Tackling Data Hotspots in Large-Scale Systems
Sharding is one of the most widely used strategies to scale databases. The assumption is simple: evenly distribute data across shards, and the workload balances itself. But in practice, workloads are rarely uniform. Some shards attract more traffic, some datasets are heavier, and “hotspots” emerge. The result? Certain shards run hot while others remain underutilized, driving inefficiencies and unpredictable performance.
In our production environment, we faced this very problem. Conventional sharding left us with overloaded shards that slowed down queries and ingestion, while others idled. Instead of brute-forcing with more infrastructure, we developed a weighted sharding scheme. Each shard was assigned a weight reflecting its capacity and traffic profile. Events and queries were routed accordingly, allowing us to balance load dynamically across the cluster.
This talk walks through the full journey: how we identified hotspots through monitoring, the design of our weighted sharding prototype, and the measurable improvements in throughput and cost efficiency. Beyond the technical implementation, we’ll share lessons about data modeling choices, observability, and keeping the solution simple enough to operate at scale.
By the end of this session, attendees will see how weighted sharding can transform uneven workloads into predictable, efficient systems and how shard-aware design choices can make infrastructure scale smarter.
Ashish Kattamuri
Staff Software Engineer, Proofpoint
Denver, Colorado, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top