Story 000 - Bootstrapping an AI Development Environment
Published:
The Chatbot That Earned Its Own GPU
Published:
The Chatbot That Earned Its Own GPU
Published:
Infrastructure vs. Platform: Setting Up a GPU for Production
Published:
How do you keep a production chatbot alive when your required GPU hardware disappears?
Published:
A break-glass temporary recovery action to bring the service back online.
Published:
I was asked to implement MCP. I fell in love with microservices and itβs a toxic relationship.
Published:
The QA Assignment That Made Me Design My Own Replacement
Published:
I implemented concurrency on intuition. Then performance engineering made me question my life choices.
Published:
From nvtop to NVIDIA Nsight: Going Beneath GPU Utilization.
32-bit RISC Processor design using Harvard Architecture in Verilog.
An IoT-based Smart Parking System with slot availability prediction models trained on parking slot usage data collected on the ThingSpeak IoT cloud platform.
Sentimental Analysis from text: A text analysis algorithm based on ETL approach with multiprocessing for parallel processing of multiple text files.
Extraction and Pre-processing of LiDAR sensor data from the full-scale and high-speed autonomous racing open-source Multi-modal sensor dataset- RACECAR dataset. 
A systems-engineering approach for composing existing technologies into a distributed, stateful, and artifact-centric workflow architecture. An autonomous GUI-based QA workflow serves as the reference implementation. 
Parallel Computing, Performance Engineering, Profiling, Concurrency, Workload Granularity
Recommended citation: Suwesh Prasad Sah. (2026). Analyzing Internal and External Parallelism in Multi-Index FAISS Retrieval: Threads, Processes, OpenMP and the Cost of Concurrency (Version 1.0). Zenodo. https://doi.org/10.5281/zenodo.22115697
View Paper
Speculative Decoding, LLM Inference, GPU Performance Analysis, GPU Kernels, CUDA Graphs, NVIDIA Nsight, PyTorch Profiler
Recommended citation: Suwesh Prasad Sah. (2026). Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference (Version 1.0). arXiv. https://arxiv.org/abs/2609.35188
View Paper
Published in 12th International Conference on Intelligent Systems and Embedded Design (ISED). IEEE, 2024
Autonomous Racing, Computer Vision, Image Segmentation, Deep Learning, LiDAR Perception, Accelerated Computing
Recommended citation: S. P. Sah and S. Chinara, "Parallel Neural Computing for Scene Understanding from LiDAR Perception in Autonomous Racing," 2024 12th International Conference on Intelligent Systems and Embedded Design (ISED), Rourkela, India, 2024, pp. 01-06, doi: 10.1109/ISED63599.2024.10956572.
View Paper
Published:
This is a description of your talk, which is a markdown files that can be all markdown-ified like any other post. Yay markdown!
Published:
This is a description of your conference proceedings talk, note the different field in type. You can put anything in this field.
Undergraduate course, University 1, Department, 2014
This is a description of a teaching experience. You can use markdown like any other post.
Workshop, University 1, Department, 2015
This is a description of a teaching experience. You can use markdown like any other post.