#SystemsAtScale
"Move fast with stable infra" rather than "Move fast and break things" as the motto for the infrastructure group at Facebook. #SystemsAtScale
January 13, 2025 at 1:24 PM
Architecture at Facebook had to evolve from PHP-Memcache-MySQL to HHVM/HACK, TAO, and MyRocks. #SystemsAtScale
January 13, 2025 at 1:24 PM
Conference introductory keynote now happening at @at_scale_events #SystemsAtScale. Jay Parikh of @fb_engineering is describing Facebook's efforts to scale since its inception, and the infrastructure teams' work there.
January 13, 2025 at 1:23 PM
Nobody livetweeted or liveblogged it, but that's okay! There now is video of me giving my talk at #SystemsAtScale about #o11y and how singular metrics and graphs need the ability to mash them up with their context to be successful! https://t.co/SZYXr7vp2W
Systems @Scale 2018 - Resolving Outages Faster with Better Debugging Strategies | At Scale Conferences
Liz Fong-Jones, Staff Site Reliability Engineer at Google, explains why building more dashboards isn’t the solution — using dynamic query evaluation and integrating tracing is.
atscaleconference.com
January 13, 2025 at 1:44 PM
And #SystemsAtScale hosted by @fb_engineering is a wrap! Hope people enjoyed the livetweeting and that I held down the fort okay; few other folks here were heavy Twitter users, for obvious reasons ;)
January 13, 2025 at 1:32 PM
Q: why not k8s? A: "yes, we do read." but K8s won't scale to the number of machines, and trying to force scaling changes in may not meet the k8s community's needs. (same reason that Netflix released its own orchestrator instead of using k8s or mesos) #SystemsAtScale
January 13, 2025 at 1:32 PM
But the work for now is just to get everything standardized and out of private pools to remove human element from fleet operations. [fin] #SystemsAtScale
January 13, 2025 at 1:32 PM
Also have a desire to create in far future update domains that allow evenly spreading out services evenly across a standardized set of failure domains. #SystemsAtScale
January 13, 2025 at 1:32 PM
Don't accept maintenance events until schedulers are confident that things are sufficiently out of the way that they won't suffer impact; however, potentially years out from being able to do that fully automatically. #SystemsAtScale
January 13, 2025 at 1:32 PM
New datacenter tooling to reason about and schedule maintenances to not conflict with each other and not cause undue impact on services. #SystemsAtScale
January 13, 2025 at 1:32 PM
Resource allowance system to be created to avoid humans being involved in horse-trading individual racks and maintaining overall capacity limits; need a fleet ledger to know what's assigned to whom.

Shard by failure domains, stay consistent for each shard #SystemsAtScale
January 13, 2025 at 1:32 PM
but multiple schedulers allowed as long as allocation is centrally handled. Most will still use Tupperware but others may want to vary. #SystemsAtScale
January 13, 2025 at 1:32 PM
Goal to have a container allocation frontend that lets you request any size of container (whether whole machine or less than machine). #SystemsAtScale
January 13, 2025 at 1:32 PM
Binpacking multiple unrelated containers onto a single machine is not a priority due to concerns about noisy neighbors; many teams are using entire single machines and optimized for it. #SystemsAtScale
January 13, 2025 at 1:31 PM
Showing an incomplete architecture diagram, pieces of which *will* be built over time. [ed: interested to see how this does turn out over time; this is definitely an unvarnished look into FB's state of infra] #SystemsAtScale
January 13, 2025 at 1:31 PM
Bare metal, bit twiddling, and no virtualization starting to run head into the problem of trying to migrate. #SystemsAtScale
January 13, 2025 at 1:31 PM
Phases: moving everything tupperware into the twshared pool, then migrating the large customized scheduling solutions into tupperware. #SystemsAtScale
January 13, 2025 at 1:31 PM
Good news is that Tupperware has existed 7+ years. Most services by count but not by fleet size are using it, but ~45% of it is all different private pools, ~45% is non-tupperware; only ~10% is on the "twshared" shared pool. #SystemsAtScale
January 13, 2025 at 1:31 PM
Compute as a service needs to bridge between service and batch scheduling on one side, and maintenance/network on the other.

Storage regarded as out of scope for now. #SystemsAtScale
January 13, 2025 at 1:31 PM
It shouldn't matter to services why they're being moved, as long as they can be moved elsewhere with intent-based automation. #SystemsAtScale
January 13, 2025 at 1:31 PM
Goal is to avoid many:many interactions between constraints, and having a single set of schedulers and compute and maintenance as a service. #SystemsAtScale
January 13, 2025 at 1:31 PM
There's no automated long-term migration scheme for longer-term removal of capacity. #SystemsAtScale
January 13, 2025 at 1:31 PM
Still need an orchestration layer on top to avoid creating SPOFs and tripping over them. #SystemsAtScale
January 13, 2025 at 1:31 PM