Assistant Professor, IIIT‑Delhi

Vivek Kumar

I work on parallel programming models and runtime systems — making multicore, heterogeneous, and power-constrained hardware easier to program and more energy-efficient. I lead the HiPeC Lab at IIIT‑Delhi.

Vivek Kumar
About

Biography

Vivek Kumar is an Assistant Professor at IIIT‑Delhi, where he leads the HiPeC Lab. His research focuses on parallel programming models and runtime systems — work-stealing schedulers, NUMA- and power-aware task scheduling, and resource management for co-running applications on multicore, multiprocessor, and heterogeneous hardware. His OOPSLA 2012 paper, “Work-Stealing Without The Baggage,” was selected for the SIGPLAN Research Highlights Papers, and his group's HiPC 2025 paper, “Energy-Aware Runtime Resource Harmonizer for Co-running Applications,” won the Distinguished Paper Award.

He earned his Ph.D. in Computer Science from The Australian National University in January 2014, under the supervision of Prof. Steve Blackburn, and holds a Bachelor's degree in Mechanical Engineering from VTU, India. Before the Ph.D., he spent nearly six years in HPC research and development roles at technology firms in Bangalore, including IBM Systems and Technology Labs and C-DAC. Afterward, he spent nearly three years as a Research Scientist at Rice University, in Prof. Vivek Sarkar's Habanero Extreme Scale Software Research Group, before joining IIIT‑Delhi in December 2016.

At IIIT‑Delhi, he has received the Excellence in Teaching Award. Students he has mentored have also been recognized, including a Google PhD Fellowship in Systems and Networking.

Parallel programming models Runtime systems Work-stealing Energy efficiency Heterogeneous computing
Latest

News

2026
Sunil Kumar submitted his Ph.D. thesis (July).
Poonam Dubey joined the lab as a Ph.D. student, co-supervised with Dr. Piyus Kedia.
Patent published in the Indian Patent Office.
Sunil's paper accepted for publication in ACM TACO as a Journal First paper.
Spoorthi S. joined the lab as a Ph.D. student.
2025
Vanshika Jain's HiPC'25 paper received the Distinguished Paper Award.
Vanshika received IEEE TCPP and Murthy Trust travel grants for HiPC'25.
Paper accepted in an SC'25 workshop.
Sunil received a student scholarship from EuroPar'25.
Sunil's paper and artifact accepted at EuroPar'25.
2024
Excellence in Teaching Award, IIIT‑Delhi.
Sunil's poster accepted at HiPC SRS.
Varun's B.Tech. thesis shortlisted for a best-thesis award.
Hardik's B.Tech. thesis work accepted in the Wiley CCPE journal.
2023
Vanshika Jain joined the lab as a Ph.D. student.
Research project proposal accepted by Shell R&D.
Taught Operating Systems in the Monsoon 2023 semester.
2022
Poster accepted at the HiPC conference.
Started teaching a new course — Parallel Runtimes for Modern Processors.
Sunil Kumar selected for a Google PhD Fellowship in Systems and Networking.
2021
Paper accepted in the EduHiPC workshop.
Our SC'21 paper named one of five Best Reproducibility Advancement Finalist papers.
Member of the AICTE Curriculum Revision Committee for Advanced Programming.
Co-organizing the PPEE workshop at HiPC.
Sunil Kumar joined the lab as a Ph.D. student.
Paper accepted at SC'21.
2020
Paper accepted at HiPC'20.
Excellence in Teaching award for CSE502.
2019
Papers accepted at EuroPar'19 and IWOMP'19.
Invited talk at IndoSys'19.
Excellence in Teaching award for CSE502 and CSE201.
2018
Excellence in Teaching award for CSE502 and CSE201.
Abhiprayah Tiwari successfully defended his M.Tech. thesis.
2017
Excellence in Teaching award at IIITD for CSE502, Spring 2017.
Paper accepted at AsHES 2017.
Joined IIIT‑Delhi as an Assistant Professor (December).
2016
Papers accepted at OpenSHMEM, PPPJ, and IA³ 2016.
2015
Papers accepted at HPEC 2015 and PGAS 2015.
2014
Joined Rice University as a Research Scientist (March).
Submitted Ph.D. thesis. Papers accepted at VEE 2014 and PGAS 2014.
2013
OOPSLA 2012 paper selected for SIGPLAN Research Highlights Papers.
Ph.D. research accepted for the Dissertation Research Showcase track at SC13.
2012
Paper accepted at OOPSLA 2012.
2010
Joined Ph.D. studies at ANU.
Publications

Research Papers

TACO 2026

S. Kumar, V. Kumar, and S. Bhalachandra, “Power Scheduling for Maximizing Throughput and Fairness in Co-running Applications,” ACM Transactions on Architecture and Code Optimization (TACO), 2026.

· PDF
power managementenergy efficiencymulticoreco-running

Due to the significant cost associated with power consumption, hardware overprovisioning is widely used to cap processor power consumption and improve the average power utilization of servers in data centers and nodes in HPC clusters. Existing solutions primarily focus on cluster-level power management, making limited use of power scheduling within multi-socket servers.

This paper introduces Fulcrum, a novel power management library for co-running parallel applications on multi-socket, multi-core servers, independent of the underlying parallel programming model. Our results show that Fulcrum improves system throughput (geometric mean) by 26.3% under low power caps and by up to 5.3% under higher power caps, with power efficiency improvements of 27.7–8.4%.

HiPC 2025Distinguished Paper Award

V. Jain, V. Parashar, V. Kumar, and C. Sur, “Energy-Aware Runtime Resource Harmonizer for Co-running Applications,” in 32nd IEEE International Conference on High Performance Computing, Data, and Analytics, Hyderabad, India, December 2025.

· PDF· Slides· Video
energy efficiencyresource managementco-runningmulticore

Modern multiprocessor systems, with their abundance of cores and sockets, make simultaneous execution of multiple applications essential for maximizing utilization. But performance and energy efficiency of co-executing applications are highly sensitive to thread placement, core allocation, and core/uncore frequency settings.

This paper presents Harmonizer, a dynamic resource optimization library for co-running applications that reduces energy consumption by 8.8–35% (20.5% geometric mean) and improves throughput by 4.8% (geometric mean) over the default Linux scheduler, and up to 28.6% energy savings and 15.8% higher throughput over two state-of-the-art approaches.

SC Workshop 2025

V. Kumar, “Teaching Task-Based Parallel Programming with a Runtime Systems-Aware Perspective,” SC25-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, St. Louis, MO, USA, November 2025.

· PDF· Slides· Audio
educationtask parallelismruntime systems

Task-based models simplify parallel programming using runtimes that handle scheduling and resource management. This paper presents the structure and experience of teaching Parallel Runtimes for Modern Processors (PRMP) at IIIT‑Delhi, where students implement a basic async–finish task model and gradually improve both the model and the runtime, with a qualitative and quantitative evaluation across three offerings.

EuroPar 2025

S. Kumar and V. Kumar, “KarmaPM: Reward-Driven Power Manager,” in 31st International European Conference on Parallel and Distributed Computing, LNCS, Springer, Dresden, Germany, August 2025.

· PDF· Slides· Audio
power managementenergy efficiencyco-runningmulticore

Hardware overprovisioning is widely used to improve average power utilization, but a uniform power cap across sockets can hurt co-running applications with variable workloads. This paper introduces KarmaPM, a power management library based on application power-donation "karmas" that redistributes power bidirectionally across sockets, improving system throughput by 13.2% at lower power caps and 6.6% at higher caps, with improvements of 12.5% and 4.4% over state-of-the-art approaches.

Physics of Fluids 2024

M. Bhatt, T. Chauhan, R. Agrawal, M. Kumar, V. Kumar, and S. Sircar, “Elastoinertial Stability Analysis and Structure Formation in Viscoelastic Subdiffusive Pipe Flow,” Physics of Fluids, 2024.

fluid dynamicssimulationinterdisciplinary
CCPE 2024

H. Saini, V. Kumar, and T. Chakraborty, “Energy Efficient Permanence-based Community Detection Algorithm,” Concurrency and Computation: Practice and Experience, 2024.

· PDF
energy efficiencytask parallelismgraph algorithmsmulticore

Introduces Amoeba, a task-parallel implementation of a permanence-based community detection algorithm for multicore processors. It uses dynamic tasking for irregular computation and can adapt the thread count for energy efficiency, achieving a geometric mean speedup of 15.3× over its sequential version and 12.4% energy savings over a non-adaptive implementation.

EduHiPC 2021

V. Kumar, “Teaching High Productivity and High Performance in an Introductory Parallel Programming Course,” in Proc. 28th IEEE HiPC Workshop (HiPCW), Bangalore, India, December 2021.

· PDF· Video
educationtask parallelismparallel programming

Presents the structure and experience of teaching the Foundations of Parallel Programming course (FPP) at IIIT‑Delhi using a task-based model, Habanero C/C++ Library (HClib), where students start with traditional models, discover their limitations, and build runtime solutions to achieve high performance.

SC 2021Best Reproducibility Advancement Finalist

S. Kumar, A. Gupta, V. Kumar, and S. Bhalachandra, “Cuttlefish: Library for Achieving Energy Efficiency in Multicore Parallel Programs,” in Proc. Intl. Conf. for High Performance Computing, Networking, Storage and Analysis (SC'21), St. Louis, MO, USA, November 2021.

· PDF· Video
energy efficiencyDVFSruntime systemsmulticore

Proposes Cuttlefish, a programming-model-oblivious C/C++ library for energy efficiency in multicore parallel programs on Intel processors. An online profiler discovers memory access patterns from model-specific registers, and Cuttlefish combines DVFS and UFS to adapt core/uncore frequencies — achieving geometric mean energy savings of 19.4% with only 3.6% slowdown.

HiPC 2020

V. Kumar, “PufferFish: NUMA-Aware Work-stealing Library using Elastic Tasks,” in 27th IEEE HiPC, Pune, India, December 2020.

work-stealingNUMAtask parallelismruntime systems

Presents PufferFish, an async–finish parallel programming model and work-stealing runtime for NUMA systems, coupling data-affinity hints with Hierarchical Place Trees (HPTs) in HClib. Its Hierarchical Elastic Tasks shrink to one worker or expand across workers depending on imbalance, achieving 1.5× speedup over HPT and 1.9× over random work-stealing on a 32-core NUMA AMD EPYC processor.

IWOMP 2019

V. Kumar, A. Tiwari, and G. Mitra, “HetroOMP: OpenMP for Hybrid Load Balancing Across Heterogeneous Processors,” in 15th International Workshop on OpenMP, LNCS, Springer, Auckland, New Zealand, September 2019.

· PDF· Slides
OpenMPheterogeneouswork-stealing

Proposes HetroOMP, an OpenMP accelerator model extension with a new hetro clause enabling simultaneous execution across host and accelerator devices. A proof-of-concept runtime for the TI Keystone-II MPSoC (4 ARM CPUs + 8 DSPs) achieves a 3.6× geometric mean speedup over the OpenMP accelerator model alone.

EuroPar 2019

V. Kumar, “Featherlight Speculative Task Parallelism,” in 25th International European Conference on Parallel and Distributed Computing, LNCS, Springer, Göttingen, Germany, August 2019.

· PDF· Slides· Audio
work-stealingspeculative parallelismtask parallelism

Proposes Featherlight, a programming model for speculative task parallelism satisfying the serial elision property without task cancellation checks. A classroom study shows productivity gains, and the cancellation technique achieves a 1.6× geometric mean speedup over Java Fork/Join on a 20-core machine.

IPDPSW 2017

M. Grossman, V. Kumar, N. Vrvilo, Z. Budimlic, and V. Sarkar, “A Pluggable Framework for Composable HPC Scheduling Libraries,” in Proc. IEEE IPDPS Workshops, Orlando, FL, USA, May 2017.

· PDF· Slides
work-stealingschedulingruntime systems

Presents HiPER, a pluggable API framework atop a generalized work-stealing runtime for composing communication, accelerator, and other HPC libraries within a single process — demonstrating programmability and performance gains through unified, asynchronous scheduling.

IA³ 2016

V. Kumar, K. Murthy, V. Sarkar, and Y. Zheng, “Optimized Distributed Work-Stealing,” in 6th Intl. Workshop on Irregular Applications (IA3), Salt Lake City, UT, USA, November 2016.

· PDF· Slides· Audio
work-stealingdistributedPGAS

Presents SuccessOnlyWS, a load-aware distributed work-stealing strategy in HabaneroUPC++ that overcomes failed steal attempts with a new busy-to-idle work-movement policy — up to 7% faster than the baseline on 12,288 cores of the Edison (CRAY-XC30) supercomputer.

PPPJ 2016

V. Kumar, J. Dolby, and S. M. Blackburn, “Integrating Asynchronous Task Parallelism and Data-centric Atomicity,” at 13th Intl. Conf. on Principles and Practices of Programming on the Java Platform (PPPJ), Lugano, Switzerland, August 2016.

· PDF· Slides· Audio
work-stealingtask parallelismJava

Presents five Java annotations for asynchronous task parallelism and data-centric concurrency control, backed by an efficient work-stealing scheduler. Refactoring classes from existing multithreaded open-source projects showed reduced programming effort and up to 30% performance improvement.

OpenSHMEM 2016

M. Grossman, V. Kumar, Z. Budimlic, and V. Sarkar, “Integrating Asynchronous Task Parallelism with OpenSHMEM,” at 3rd Workshop on OpenSHMEM and Related Technologies, Baltimore, MD, USA, August 2016.

· PDF· Slides
PGASwork-stealingdistributed

Introduces AsyncSHMEM, integrating OpenSHMEM with a thread-pool-based work-stealing runtime to hide data-transfer latency, interoperate with tasks, and improve load balancing and locality. On the Titan supercomputer, AsyncSHMEM is competitive on regular workloads and significantly outperforms OpenSHMEM+OpenMP on event-driven applications.

HPEC 2015

V. Kumar, A. Sbirlea, Z. Budimlic, D. Majeti, and V. Sarkar, “Heterogeneous Work-stealing across CPU and DSP cores,” at 19th Intl. Conf. on High Performance Extreme Computing (HPEC), Waltham, MA, USA, September 2015.

· PDF· Slides· Audio
work-stealingheterogeneousDSP

Presents HC-K2H, a programming model and work-stealing runtime for the TI Keystone II Hawking platform (4 ARM CPUs + 8 DSPs), allowing tasks to run and synchronize seamlessly on either processor type, with excellent scaling over sequential single-ARM implementations.

PGAS 2014

V. Kumar, Y. Zheng, V. Cave, Z. Budimlic, and V. Sarkar, “HabaneroUPC++: a Compiler-free PGAS Library,” at 8th Intl. Conf. on Partitioned Global Address Space Programming Models, Eugene, OR, October 2014.

· PDF· Slides· Audio
PGASwork-stealingdistributed

Introduces Habanero-UPC++, a compiler-free PGAS library built on UPC++ and HClib using C++11 lambdas, tightly integrating intra-place and inter-place parallelism. Evaluated on two benchmarks scaled up to 6,000 cores.

VEE 2014

V. Kumar, S. M. Blackburn, and D. Grove, “Friendly Barriers: Efficient Work-Stealing With Return Barriers,” at 10th ACM SIGPLAN/SIGOPS Intl. Conf. on Virtual Execution Environments (VEE), Salt Lake City, UT, March 2014.

· PDF· Slides
work-stealingruntime systemsJVM

Addresses dynamic work-stealing overheads dominated by stack introspection on a steal. Low-overhead return barriers cut this overhead roughly in half, improving total performance by up to 20% and the scalability of work-stealing applications.

OOPSLA 2012SIGPLAN Research Highlight

V. Kumar, D. Frampton, S. M. Blackburn, D. Grove, and O. Tardieu, “Work-Stealing Without The Baggage,” in Proc. ACM SIGPLAN Conf. on Object-Oriented Programming Systems, Languages & Applications (OOPSLA), Tucson, AZ, October 2012. Selected for SIGPLAN Research Highlights Papers, 2013.

· PDF· Slides· Audio
work-stealingruntime systemsJVM

Identifies key sources of overhead in work-stealing schedulers and presents two refinements. Our fastest design has just 15% overhead over sequential Java, versus 2.3× for fork-join and 4.1× for the prior system — encouraging further exploitation of hardware parallelism via work-stealing.

VMIL 2012

V. Kumar and S. M. Blackburn, “Faster Work-Stealing With Return Barriers,” at 6th Workshop on Virtual Machines and Intermediate Languages (VMIL), Tucson, AZ, October 2012.

· PDF· Slides
work-stealingruntime systemsJVM

Identifies the overhead of managing work-stealing information on a victim's execution stack, and uses return barriers to reduce it — up to 58% lower overhead than our prior design on classical work-stealing benchmarks.

X10 Workshop 2011

V. Kumar, D. Frampton, D. Grove, O. Tardieu, and S. M. Blackburn, “Work-Stealing by Stealing States from Live Stack Frames of a Running Application,” at X10'11 Workshop, collocated with PLDI, San Jose, CA, June 2011.

· PDF· Slides
work-stealingruntime systemsX10

Proposes letting thieves extract state directly from a producer's stack frames using compiler-provided state maps, avoiding stack-allocated state objects. Discusses design and preliminary findings inside the X10 work-stealing runtime and Jikes RVM.

Patent (IPO) 2025

V. Kumar and S. Kumar, “Method and System for Bidirectional Power Scheduling in Multiprocessor Environments,” Indian Patent Application No. 202511096593, published December 2025.

power managementenergy efficiencymultiprocessorco-running

Describes a bidirectional power scheduling method for a multiprocessor: assigning a predefined power cap to each socket, periodically monitoring CPU utilization, determining whether a socket requires a change to its cap, and dynamically redistributing the assigned power cap bidirectionally among sockets in response.

HiPC SRS 2022 · Poster

S. Kumar, V. Kumar, and S. Bhalachandra, “Energy-Efficient Execution of Multicore Parallel Programs under Limited Power Budget,” poster, 29th IEEE HiPC Student Research Symposium, Bangalore, India, December 2022.

energy efficiencypower managementmulticore
PGAS 2015 · Poster

V. Kumar, M. Grossman, H. Shan, and V. Sarkar, “Scaling HabaneroUPC++ on Heterogeneous Supercomputers,” poster with extended abstract, 9th Intl. Conf. on Partitioned Global Address Space Programming Models (PGAS), Washington, D.C., September 2015.

PGASheterogeneousdistributed
ANU CECS HDR 2010 · Poster

V. Kumar, “Rule The Next Generation Supercomputers With X10,” ANU CECS HDR student poster presentation — first prize, 2010.

X10runtime systems
Ph.D. Thesis · 2015

V. Kumar, “Achieving High Performance and High Productivity in Next Generation Parallel Programming Languages,” Ph.D. thesis, The Australian National University, Canberra, Australia, degree awarded May 2015.

work-stealingruntime systemsJVM
Recognition

Awards & Honors

Teaching

Courses

Multicore Parallel Programming and Runtime Systems (CSE595)
Spring 2026

Covers task, data, and loop parallelism and the runtime systems that schedule and manage parallel work — work-stealing schedulers, NUMA-aware runtimes, memory consistency, cache coherency, SIMD vectorization, GPU computing, and power management. Course page ↗

Operating Systems (CSE213)
Fall 2023–2026

Core second-year course covering OS design principles and internals, paired with hands-on implementation of core OS functionality. Course page ↗

Parallel Runtimes for Modern Processors (CSE513)
Fall 2022, 2024, 2025

Introduces the design and implementation of parallel runtime systems and explores the challenges of achieving performance and energy efficiency on modern many-core, wide-vector, and accelerator-attached processors. Course page ↗

Advanced Programming (ACM CSEDU FDP)
January 2023, July 2023

A faculty-development module for Computer Science teachers, covering the OOP paradigm, reusable code design, defensive programming, unit testing, design patterns, and multithreading. Program page ↗

Foundations of Parallel Programming (CSE502)
Spring 2017–2023

Introduces the fundamentals of parallel programming, both traditional approaches and newer advancements, with hands-on experience writing parallel programs across several programming models. Course page ↗

Advanced Programming (CSE201)
Fall 2017–2021

Prepares students to build large-scale, multi-component programs in Java — object orientation, reusable code design, test-driven development, and pattern-oriented program design.

Mentorship

Students

Full team and lab-wide alumni roster live on the HiPeC Lab members page ↗.
History

Past Affiliations

Reach Out

Contact

Room B-506, R&D Block, 5th Floor
IIIT‑Delhi, Near Okhla Phase 3
New Delhi–110020, India

Email: vivekk [at] iiitd.ac.in