i currently work on
- dynamo - a datacenter scale inference orchestration framework
- sglang - a blazing fast llm inference engine.
My focus is primarily on efficient LLM routing policies and optimizing both the orchestration and engine for RL and agentic inference. You can read more about this here
I am also the author and maintainer of srt-slurm which provides a k8s style deployment experience on SLURM. It is used extensively for benchmarking inside and outside of NVIDIA.
previously
- brev.dev (ai lead -> acq. by nvidia)
- agora labs (founder -> acq. by brev.dev)
- columbia university (dropped out for agora)
- sparkcognition (deep learning for industrial optimization)
- federal reserve (time series modeling)
more of me