Software engineer based in Mumbai. I work on inference infrastructure — the scheduling, batching, and memory management that determines whether your LLM deployment is economical or not.
Before getting pulled into ML systems, I spent a lot of time on lower layers: writing compilers, implementing databases, building message queues. Not because I needed to, but because I have a hard time using something I can't at least partially explain. That compulsion has produced a lot of redundant software and most of my blog posts.
The posts on this site are usually me working through something I should have understood earlier — FFI and calling conventions, how inference batching actually works under memory pressure, that kind of thing. Writing it down is how I find out what I actually know vs. what I just think I know.
Right now I'm thinking about KV cache architectures, scheduling for mixed workloads, and how much of the “optimize the ML compute graph” problem is actually solved by compiler passes vs. still requiring hand-tuning. Forming opinions slowly.
GMT+5:30 · subham.k.ojha@gmail.com