Skip to main content

Command Palette

Search for a command to run...

About Me

I'm a systems engineer interested in the boundaries between ML inference, runtimes, compilers, and accelerator hardware.

My work has taken me across several layers of the stack, from ML compiler and hardware/software co-design to low-level runtime and virtualization systems, and more recently large-scale accelerator infrastructure. That path has made me especially interested in the places where those layers stop being cleanly separable: when a compiler decision needs runtime information, when a scheduling decision becomes a network problem, or when something is important enough to justify dedicated hardware.

Outside of work, I've been building mini-vllm-rs, a small LLM inference engine in Rust. I started it as a way to learn modern inference systems by implementing them myself, but it has also become a useful place to test ideas around continuous batching, KV-cache management, speculative decoding, heterogeneous execution, and performance trade-offs.

ML Infra System Design is where I write about questions that come out of that work and my broader experience. I don't expect all of them to have clean answers. The goal is to connect ideas that are often discussed separately, use real systems and experiments where possible, and share what I learn along the way.

If you're interested in inference systems, distributed ML infrastructure, compilers, or hardware/software co-design, stay tuned.