amar phanishayee

Intro | Publications | Misc.

Hi !

I work at NVIDIA. Earlier, for most of my "professional" career, I worked at the creative haven of Microsoft Research in Redmond . There, among other things, I started and led Project Fiddle, focusing on systems for fast, efficient, and fault-tolerant large-scale DNN training and serving.

My work in systems for machine learning so far has taken a broad view of DNN training and serving: from a single GPU, all the way to multi-datacenter-scale systems with 10s to 100s of thousands of accelerators. Our innovations in this context cut across the systems stack; they target memory management, structuring parallel computation across GPUs and machines (e.g., proposing pipeline parallel training), synthesized throughput-optimal collectives, data loading, checkpointing, disaggregated inference, fault tolerance, and schedulers for multi-tenant clusters. I was also fortunate, working with a group of fearless engineers, to be one of a small group of software-systems architects for Microsoft's early forays in 1P ("First Party") AI Supercomputing efforts.

Earlier still, many moons ago, I got my PhD in Computer Science at Carnegie Mellon University where I worked with Dave Andersen. The core of my work, with fellow collaborators then, anchored around: (i) FAWN (energy-efficient distributed systems, especially for large-scale random-access IO-bound workloads), and (ii) Incast (catastrophic TCP throughput collapse in datacenter-scale workloads). Incast has since become an adjective and Wimpy is used in straight-faced technical conversations; who'd a thunk it?

In general, the goal of my research is to enable the creation of high performance and efficient systems for large-scale data-intensive computing. To this end, my work follows two broad approaches:
  1. Radically rethinking datacenter architecture

  2. Build robust distributed systems and protocols


Publications

Misc.

Intro | Publications | Misc.