Computer science at NUST · Islamabad, Pakistan
Bilal Rana
I study how to make large language model inference faster, more memory-efficient, and more predictable on constrained GPU systems.
My LLM Inference Systems Lab is a collection of reproducible GPU studies: each one starts with a concrete systems question and publishes the method, measurements, limitations, code, and raw data behind the answer.
Research Interests
- GPU computing and efficient inference for large language models
- KV-cache management and pruning
- Request scheduling and serving in engines such as vLLM
- Quantization and sparse or compressed tensor formats
- Energy-efficient inference
- GPU kernel and roofline analysis
- ML compilers
- Distributed inference
- Hardware-software codesign
Achievements
- 3× High Achiever, NUST SEECS
- Winner, CUST Hackathon (Bideez)
- Winner, FAST Hackathon (InterviewAI)