Story 007 - Making LLM Inference Faster Led Me Down the Performance Stack
Published:
From nvtop to NVIDIA Nsight: Going Beneath GPU Utilization.
Published:
From nvtop to NVIDIA Nsight: Going Beneath GPU Utilization.
Published:
I implemented concurrency on intuition. Then performance engineering made me question my life choices.
Published:
The QA Assignment That Made Me Design My Own Replacement
Published:
I was asked to implement MCP. I fell in love with microservices and it’s a toxic relationship.
Published:
A break-glass temporary recovery action to bring the service back online.
Published:
How do you keep a production chatbot alive when your required GPU hardware disappears?
Published:
Infrastructure vs. Platform: Setting Up a GPU for Production
Published:
The Chatbot That Earned Its Own GPU