I work on parallel programming models and runtime systems — making multicore, heterogeneous, and power-constrained hardware easier to program and more energy-efficient. I lead the HiPeC Lab at IIIT‑Delhi.

Vivek Kumar is an Assistant Professor at IIIT‑Delhi, where he leads the HiPeC Lab. His research focuses on parallel programming models and runtime systems — work-stealing schedulers, NUMA- and power-aware task scheduling, and resource management for co-running applications on multicore, multiprocessor, and heterogeneous hardware. His OOPSLA 2012 paper, “Work-Stealing Without The Baggage,” was selected for the SIGPLAN Research Highlights Papers, and his group's HiPC 2025 paper, “Energy-Aware Runtime Resource Harmonizer for Co-running Applications,” won the Distinguished Paper Award.
He earned his Ph.D. in Computer Science from The Australian National University in January 2014, under the supervision of Prof. Steve Blackburn, and holds a Bachelor's degree in Mechanical Engineering from VTU, India. Before the Ph.D., he spent nearly six years in HPC research and development roles at technology firms in Bangalore, including IBM Systems and Technology Labs and C-DAC. Afterward, he spent nearly three years as a Research Scientist at Rice University, in Prof. Vivek Sarkar's Habanero Extreme Scale Software Research Group, before joining IIIT‑Delhi in December 2016.
At IIIT‑Delhi, he has received the Excellence in Teaching Award. Students he has mentored have also been recognized, including a Google PhD Fellowship in Systems and Networking.
S. Kumar, V. Kumar, and S. Bhalachandra, “Power Scheduling for Maximizing Throughput and Fairness in Co-running Applications,” ACM Transactions on Architecture and Code Optimization (TACO), 2026.
Due to the significant cost associated with power consumption, hardware overprovisioning is widely used to cap processor power consumption and improve the average power utilization of servers in data centers and nodes in HPC clusters. Existing solutions primarily focus on cluster-level power management, making limited use of power scheduling within multi-socket servers.
This paper introduces Fulcrum, a novel power management library for co-running parallel applications on multi-socket, multi-core servers, independent of the underlying parallel programming model. Our results show that Fulcrum improves system throughput (geometric mean) by 26.3% under low power caps and by up to 5.3% under higher power caps, with power efficiency improvements of 27.7–8.4%.
V. Jain, V. Parashar, V. Kumar, and C. Sur, “Energy-Aware Runtime Resource Harmonizer for Co-running Applications,” in 32nd IEEE International Conference on High Performance Computing, Data, and Analytics, Hyderabad, India, December 2025.
Modern multiprocessor systems, with their abundance of cores and sockets, make simultaneous execution of multiple applications essential for maximizing utilization. But performance and energy efficiency of co-executing applications are highly sensitive to thread placement, core allocation, and core/uncore frequency settings.
This paper presents Harmonizer, a dynamic resource optimization library for co-running applications that reduces energy consumption by 8.8–35% (20.5% geometric mean) and improves throughput by 4.8% (geometric mean) over the default Linux scheduler, and up to 28.6% energy savings and 15.8% higher throughput over two state-of-the-art approaches.
V. Kumar, “Teaching Task-Based Parallel Programming with a Runtime Systems-Aware Perspective,” SC25-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, St. Louis, MO, USA, November 2025.
Task-based models simplify parallel programming using runtimes that handle scheduling and resource management. This paper presents the structure and experience of teaching Parallel Runtimes for Modern Processors (PRMP) at IIIT‑Delhi, where students implement a basic async–finish task model and gradually improve both the model and the runtime, with a qualitative and quantitative evaluation across three offerings.
S. Kumar and V. Kumar, “KarmaPM: Reward-Driven Power Manager,” in 31st International European Conference on Parallel and Distributed Computing, LNCS, Springer, Dresden, Germany, August 2025.
Hardware overprovisioning is widely used to improve average power utilization, but a uniform power cap across sockets can hurt co-running applications with variable workloads. This paper introduces KarmaPM, a power management library based on application power-donation "karmas" that redistributes power bidirectionally across sockets, improving system throughput by 13.2% at lower power caps and 6.6% at higher caps, with improvements of 12.5% and 4.4% over state-of-the-art approaches.
M. Bhatt, T. Chauhan, R. Agrawal, M. Kumar, V. Kumar, and S. Sircar, “Elastoinertial Stability Analysis and Structure Formation in Viscoelastic Subdiffusive Pipe Flow,” Physics of Fluids, 2024.
H. Saini, V. Kumar, and T. Chakraborty, “Energy Efficient Permanence-based Community Detection Algorithm,” Concurrency and Computation: Practice and Experience, 2024.
Introduces Amoeba, a task-parallel implementation of a permanence-based community detection algorithm for multicore processors. It uses dynamic tasking for irregular computation and can adapt the thread count for energy efficiency, achieving a geometric mean speedup of 15.3× over its sequential version and 12.4% energy savings over a non-adaptive implementation.
V. Kumar, “Teaching High Productivity and High Performance in an Introductory Parallel Programming Course,” in Proc. 28th IEEE HiPC Workshop (HiPCW), Bangalore, India, December 2021.
Presents the structure and experience of teaching the Foundations of Parallel Programming course (FPP) at IIIT‑Delhi using a task-based model, Habanero C/C++ Library (HClib), where students start with traditional models, discover their limitations, and build runtime solutions to achieve high performance.
S. Kumar, A. Gupta, V. Kumar, and S. Bhalachandra, “Cuttlefish: Library for Achieving Energy Efficiency in Multicore Parallel Programs,” in Proc. Intl. Conf. for High Performance Computing, Networking, Storage and Analysis (SC'21), St. Louis, MO, USA, November 2021.
Proposes Cuttlefish, a programming-model-oblivious C/C++ library for energy efficiency in multicore parallel programs on Intel processors. An online profiler discovers memory access patterns from model-specific registers, and Cuttlefish combines DVFS and UFS to adapt core/uncore frequencies — achieving geometric mean energy savings of 19.4% with only 3.6% slowdown.
V. Kumar, “PufferFish: NUMA-Aware Work-stealing Library using Elastic Tasks,” in 27th IEEE HiPC, Pune, India, December 2020.
Presents PufferFish, an async–finish parallel programming model and work-stealing runtime for NUMA systems, coupling data-affinity hints with Hierarchical Place Trees (HPTs) in HClib. Its Hierarchical Elastic Tasks shrink to one worker or expand across workers depending on imbalance, achieving 1.5× speedup over HPT and 1.9× over random work-stealing on a 32-core NUMA AMD EPYC processor.
V. Kumar, A. Tiwari, and G. Mitra, “HetroOMP: OpenMP for Hybrid Load Balancing Across Heterogeneous Processors,” in 15th International Workshop on OpenMP, LNCS, Springer, Auckland, New Zealand, September 2019.
Proposes HetroOMP, an OpenMP accelerator model extension with a new hetro clause enabling simultaneous execution across host and accelerator devices. A proof-of-concept runtime for the TI Keystone-II MPSoC (4 ARM CPUs + 8 DSPs) achieves a 3.6× geometric mean speedup over the OpenMP accelerator model alone.
V. Kumar, “Featherlight Speculative Task Parallelism,” in 25th International European Conference on Parallel and Distributed Computing, LNCS, Springer, Göttingen, Germany, August 2019.
Proposes Featherlight, a programming model for speculative task parallelism satisfying the serial elision property without task cancellation checks. A classroom study shows productivity gains, and the cancellation technique achieves a 1.6× geometric mean speedup over Java Fork/Join on a 20-core machine.
M. Grossman, V. Kumar, N. Vrvilo, Z. Budimlic, and V. Sarkar, “A Pluggable Framework for Composable HPC Scheduling Libraries,” in Proc. IEEE IPDPS Workshops, Orlando, FL, USA, May 2017.
Presents HiPER, a pluggable API framework atop a generalized work-stealing runtime for composing communication, accelerator, and other HPC libraries within a single process — demonstrating programmability and performance gains through unified, asynchronous scheduling.
V. Kumar, K. Murthy, V. Sarkar, and Y. Zheng, “Optimized Distributed Work-Stealing,” in 6th Intl. Workshop on Irregular Applications (IA3), Salt Lake City, UT, USA, November 2016.
Presents SuccessOnlyWS, a load-aware distributed work-stealing strategy in HabaneroUPC++ that overcomes failed steal attempts with a new busy-to-idle work-movement policy — up to 7% faster than the baseline on 12,288 cores of the Edison (CRAY-XC30) supercomputer.
V. Kumar, J. Dolby, and S. M. Blackburn, “Integrating Asynchronous Task Parallelism and Data-centric Atomicity,” at 13th Intl. Conf. on Principles and Practices of Programming on the Java Platform (PPPJ), Lugano, Switzerland, August 2016.
Presents five Java annotations for asynchronous task parallelism and data-centric concurrency control, backed by an efficient work-stealing scheduler. Refactoring classes from existing multithreaded open-source projects showed reduced programming effort and up to 30% performance improvement.
M. Grossman, V. Kumar, Z. Budimlic, and V. Sarkar, “Integrating Asynchronous Task Parallelism with OpenSHMEM,” at 3rd Workshop on OpenSHMEM and Related Technologies, Baltimore, MD, USA, August 2016.
Introduces AsyncSHMEM, integrating OpenSHMEM with a thread-pool-based work-stealing runtime to hide data-transfer latency, interoperate with tasks, and improve load balancing and locality. On the Titan supercomputer, AsyncSHMEM is competitive on regular workloads and significantly outperforms OpenSHMEM+OpenMP on event-driven applications.
V. Kumar, A. Sbirlea, Z. Budimlic, D. Majeti, and V. Sarkar, “Heterogeneous Work-stealing across CPU and DSP cores,” at 19th Intl. Conf. on High Performance Extreme Computing (HPEC), Waltham, MA, USA, September 2015.
Presents HC-K2H, a programming model and work-stealing runtime for the TI Keystone II Hawking platform (4 ARM CPUs + 8 DSPs), allowing tasks to run and synchronize seamlessly on either processor type, with excellent scaling over sequential single-ARM implementations.
V. Kumar, Y. Zheng, V. Cave, Z. Budimlic, and V. Sarkar, “HabaneroUPC++: a Compiler-free PGAS Library,” at 8th Intl. Conf. on Partitioned Global Address Space Programming Models, Eugene, OR, October 2014.
Introduces Habanero-UPC++, a compiler-free PGAS library built on UPC++ and HClib using C++11 lambdas, tightly integrating intra-place and inter-place parallelism. Evaluated on two benchmarks scaled up to 6,000 cores.
V. Kumar, S. M. Blackburn, and D. Grove, “Friendly Barriers: Efficient Work-Stealing With Return Barriers,” at 10th ACM SIGPLAN/SIGOPS Intl. Conf. on Virtual Execution Environments (VEE), Salt Lake City, UT, March 2014.
Addresses dynamic work-stealing overheads dominated by stack introspection on a steal. Low-overhead return barriers cut this overhead roughly in half, improving total performance by up to 20% and the scalability of work-stealing applications.
V. Kumar, D. Frampton, S. M. Blackburn, D. Grove, and O. Tardieu, “Work-Stealing Without The Baggage,” in Proc. ACM SIGPLAN Conf. on Object-Oriented Programming Systems, Languages & Applications (OOPSLA), Tucson, AZ, October 2012. Selected for SIGPLAN Research Highlights Papers, 2013.
Identifies key sources of overhead in work-stealing schedulers and presents two refinements. Our fastest design has just 15% overhead over sequential Java, versus 2.3× for fork-join and 4.1× for the prior system — encouraging further exploitation of hardware parallelism via work-stealing.
V. Kumar and S. M. Blackburn, “Faster Work-Stealing With Return Barriers,” at 6th Workshop on Virtual Machines and Intermediate Languages (VMIL), Tucson, AZ, October 2012.
Identifies the overhead of managing work-stealing information on a victim's execution stack, and uses return barriers to reduce it — up to 58% lower overhead than our prior design on classical work-stealing benchmarks.
V. Kumar, D. Frampton, D. Grove, O. Tardieu, and S. M. Blackburn, “Work-Stealing by Stealing States from Live Stack Frames of a Running Application,” at X10'11 Workshop, collocated with PLDI, San Jose, CA, June 2011.
Proposes letting thieves extract state directly from a producer's stack frames using compiler-provided state maps, avoiding stack-allocated state objects. Discusses design and preliminary findings inside the X10 work-stealing runtime and Jikes RVM.
V. Kumar and S. Kumar, “Method and System for Bidirectional Power Scheduling in Multiprocessor Environments,” Indian Patent Application No. 202511096593, published December 2025.
Describes a bidirectional power scheduling method for a multiprocessor: assigning a predefined power cap to each socket, periodically monitoring CPU utilization, determining whether a socket requires a change to its cap, and dynamically redistributing the assigned power cap bidirectionally among sockets in response.
S. Kumar, V. Kumar, and S. Bhalachandra, “Energy-Efficient Execution of Multicore Parallel Programs under Limited Power Budget,” poster, 29th IEEE HiPC Student Research Symposium, Bangalore, India, December 2022.
V. Kumar, M. Grossman, H. Shan, and V. Sarkar, “Scaling HabaneroUPC++ on Heterogeneous Supercomputers,” poster with extended abstract, 9th Intl. Conf. on Partitioned Global Address Space Programming Models (PGAS), Washington, D.C., September 2015.
V. Kumar, “Rule The Next Generation Supercomputers With X10,” ANU CECS HDR student poster presentation — first prize, 2010.
V. Kumar, “Achieving High Performance and High Productivity in Next Generation Parallel Programming Languages,” Ph.D. thesis, The Australian National University, Canberra, Australia, degree awarded May 2015.
Covers task, data, and loop parallelism and the runtime systems that schedule and manage parallel work — work-stealing schedulers, NUMA-aware runtimes, memory consistency, cache coherency, SIMD vectorization, GPU computing, and power management. Course page ↗
Core second-year course covering OS design principles and internals, paired with hands-on implementation of core OS functionality. Course page ↗
Introduces the design and implementation of parallel runtime systems and explores the challenges of achieving performance and energy efficiency on modern many-core, wide-vector, and accelerator-attached processors. Course page ↗
A faculty-development module for Computer Science teachers, covering the OOP paradigm, reusable code design, defensive programming, unit testing, design patterns, and multithreading. Program page ↗
Introduces the fundamentals of parallel programming, both traditional approaches and newer advancements, with hands-on experience writing parallel programs across several programming models. Course page ↗
Prepares students to build large-scale, multi-component programs in Java — object orientation, reusable code design, test-driven development, and pattern-oriented program design.
Room B-506, R&D Block, 5th Floor
IIIT‑Delhi, Near Okhla Phase 3
New Delhi–110020, India
Email: vivekk [at] iiitd.ac.in