Sitemap
A list of all the posts and pages found on the site. For you robots out there is an XML version available for digesting as well.
Pages
Posts
Future Blog Post
Published:
This post will show up by default. To disable scheduling of future posts, edit config.yml and set future: false.
Blog Post number 4
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
Blog Post number 3
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
Blog Post number 2
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
Blog Post number 1
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
engineering_war_stories
Story 000 - Bootstrapping an AI Development Environment
Published:
The Chatbot That Earned Its Own GPU
Story 001 - First production deployment (feat. RHEL EC2)
Published:
Infrastructure vs. Platform: Setting Up a GPU for Production
Story 002 - The Curious Case of Missing G5 Capacity
Published:
How do you keep a production chatbot alive when your required GPU hardware disappears?
Story 003 - When an OS patching broke production
Published:
A break-glass temporary recovery action to bring the service back online.
Story 004 - The MCP Detour That Taught Me Microservices
Published:
I was asked to implement MCP. I fell in love with microservices and it’s a toxic relationship.
Story 005 - I Tried to Give My Job to a Machine
Published:
The QA Assignment That Made Me Design My Own Replacement
Story 006 - The Parallelism That Made Things Slower
Published:
I implemented concurrency on intuition. Then performance engineering made me question my life choices.
Story 007 - Making LLM Inference Faster Led Me Down the Performance Stack
Published:
From nvtop to NVIDIA Nsight: Going Beneath GPU Utilization.
portfolio
32-bit RISC Processor
32-bit RISC Processor design using Harvard Architecture in Verilog.
Parking Forecasting System & Dataset
An IoT-based Smart Parking System with slot availability prediction models trained on parking slot usage data collected on the ThingSpeak IoT cloud platform.
Multi-core Sentiment Engine
Sentimental Analysis from text: A text analysis algorithm based on ETL approach with multiprocessing for parallel processing of multiple text files.
RACECAR LiDAR Datasets
Extraction and Pre-processing of LiDAR sensor data from the full-scale and high-speed autonomous racing open-source Multi-modal sensor dataset- RACECAR dataset. 
Reverse Engineering a Knowledge Worker
A systems-engineering approach for composing existing technologies into a distributed, stateful, and artifact-centric workflow architecture. An autonomous GUI-based QA workflow serves as the reference implementation. 
publications
Analyzing Internal and External Parallelism in Multi-Index FAISS Retrieval: Threads, Processes, OpenMP and the Cost of Concurrency
Parallel Computing, Performance Engineering, Profiling, Concurrency, Workload Granularity
Recommended citation: Suwesh Prasad Sah. (2026). Analyzing Internal and External Parallelism in Multi-Index FAISS Retrieval: Threads, Processes, OpenMP and the Cost of Concurrency (Version 1.0). Zenodo. https://doi.org/10.5281/zenodo.22115697
View Paper
Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference
Speculative Decoding, LLM Inference, GPU Performance Analysis, GPU Kernels, CUDA Graphs, NVIDIA Nsight, PyTorch Profiler
Recommended citation: Suwesh Prasad Sah. (2026). Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference (Version 1.0). arXiv. https://arxiv.org/abs/2609.35188
View Paper
Parallel Neural Computing for Scene Understanding from LiDAR Perception in Autonomous Racing
Published in 12th International Conference on Intelligent Systems and Embedded Design (ISED). IEEE, 2024
Autonomous Racing, Computer Vision, Image Segmentation, Deep Learning, LiDAR Perception, Accelerated Computing
Recommended citation: S. P. Sah and S. Chinara, "Parallel Neural Computing for Scene Understanding from LiDAR Perception in Autonomous Racing," 2024 12th International Conference on Intelligent Systems and Embedded Design (ISED), Rourkela, India, 2024, pp. 01-06, doi: 10.1109/ISED63599.2024.10956572.
View Paper
talks
Talk 1 on Relevant Topic in Your Field
Published:
This is a description of your talk, which is a markdown files that can be all markdown-ified like any other post. Yay markdown!
Conference Proceeding talk 3 on Relevant Topic in Your Field
Published:
This is a description of your conference proceedings talk, note the different field in type. You can put anything in this field.
teaching
Teaching experience 1
Undergraduate course, University 1, Department, 2014
This is a description of a teaching experience. You can use markdown like any other post.
Teaching experience 2
Workshop, University 1, Department, 2015
This is a description of a teaching experience. You can use markdown like any other post.
