About
Hello! I’m a postdoc in the Analytics and AI Methods at Scale (AAIMS) group at Oak Ridge National Laboratory (ORNL), TN, USA. My current work encompasses scalable ML/AI and federated learning. I’m also involved in the DoE-directed (US Department of Energy) Genesis Mission effort [link1, link2].
I obtained my bachelors in Electrical engineering, masters and doctorate in Intelligent Systems Engineering (ISE). In my Ph.D. studies, I was advised by Prof. Martin Swany in the Luddy School of Informatics, Computing and Engineering at Indiana University Bloomington. There, I worked at the intersection of deep learning, distributed systems and systems for ML/ML for systems. My thesis focused on developing efficient computation and communication models to scale artificial neural networks across edge, cloud and high-performance computing (HPC) environments. My research has been published and presented at leading conference venues and research groups.
Prior to graduate school, I held various software engineering, big data and data science roles in the industry.
Curriculum vitae [last updated 12/2025].
Happy to connect and collaborate. I can be reached at [email].
Research Interests
- Deep Learning systems
- Federated Learning
- Distributed Computing (cloud + HPC)
- ML for Systems/Systems for ML
News
-
07/2026 - Organizing the 3rd Workshop on Federated and Privacy-Preserving AI for HPC (FPAI-HPC’26) in conjunction with IEEE CLUSTER’26 this September in Alexandria, Virginia, USA.
-
06/2026 - Attended 2026 Trillion Parameter Consortium (TPC26) in Baltimore, Maryland, and showcased my work on cross-facility federated learning [link].
-
02/2026 - Tula: Optimizing Performance, Cost and Generalization in Large-Batch Training accepted to IEEE/ACM 26th International Symposium on Cluster, Cloud and Internet Computing (CCGrid’26).
-
12/2025 - Serving on the TPC at the 2026 International Conference for High Performance Computing, Networking, Storage and Analysis, under Data Analytics, Visualization and Storage track (SC’26).
-
11/2025 - Attending SC25 in St. Louis, MO to present my OmniFed work and talk about all things federated learning, privacy and agentic workflows! Please feel free to connect! [pdf]
-
09/2025 — Serving on program committee of the 40th IEEE International Parallel & Distributed Processing Symposium (IPDPS’26).
-
09/2025 — OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC accepted at the Extreme Heterogeneity and AI Convergence in HPC workshop to be held at SC’25.
-
09/2025 — Presented our modular federated learning framework OmniFed as a poster at ORNL Software and Data expo (OSDX’25).
-
08/2025 — Presented my ongoing project at ORNL’s Federated and Collaborative Learning for Exascale Science and Security Workshop, Tennessee, USA.
-
08/2025 — Attended Trillion Parameter Consortium (TPC25) at San Diego, California, USA.
-
07/2025 — Serving on the Artifact Evaluation Board (AEB) of Journal of Systems Research (JSys) for academic year 2025-26.
-
07/2025 — Presented my work “Enabling Large-Batch Training via Learned Gradient Mapping” at ORNL’s ORPA Research Symposium.
-
04/2025 — Released a survey report on trends and advances in computational and communication methods in scalable deep learning [link].
-
03/2025 — OmniLearn: A Framework for Distributed Deep Learning over Heterogeneous Clusters was accepted to IEEE Transactions on Parallel and Distributed Systems (TPDS).
-
01/2025 — Said goodbye to Bloomington and moved to Oak Ridge, Tennessee.
-
12/2024 — Graduated from Indiana University Bloomington [link1], [link2].
-
09/2024 — Successfully passed Ph.D. defense.
Education
- Ph.D., Intelligent Systems Engineering, Indiana University Bloomington, USA.
- Thesis: Towards Building Efficient Computation and Communication Models for Deep Learning Systems [pdf].
- M.S., Computer Engineering, Indiana University Bloomington, USA.
Dev skills
- Distributed computing: Hadoop, Spark, Storm, Kafka, ZeroMQ, Ray, Docker, OpenShift
- ML/AI: PyTorch, TorchTitan, DeepSpeed, Keras, TensorFlow, Hugging Face, NVFlare, FastMCP, Langchain, Microsoft Agent Framework, OpenAI SDK
- Cloud/Cluster computing: Google Cloud Platform (GCP), Amazon Web Services (AWS), Slurm, OpenPBS
Publications
- Posters
- Tyagi, S., & Wang, F. OmniFed: Towards Configurable Cross-Silo Federated Learning. 2026 The Future of Computing, Collaboration Catalyst, ORNL.
- Tyagi, S., Cozma, A., Kotevska, O., & Wang, F. OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC. 2025 ORNL Software and Data Expo (OSDX ‘25) [pdf].
- Tyagi, S., & Swany, M. Accelerating Distributed ML Training via Selective Synchronization (Poster Abstract). 2023 IEEE International Conference on Cluster Computing Workshops (CLUSTER Workshops ‘23), 56-57 [pdf].
- Tyagi, S. Scavenger: A Cloud Service for Optimizing Cost and Performance of DL Training. 2023 IEEE/ACM 23rd International Symposium on Cluster, Cloud and Internet Computing Workshops (CCGridW ‘23), 349-350 [pdf]. (Best Poster Award)
- Workshop proceedings
- Tyagi, S.*, Cozma, A.*, Kotevska, O., & Wang, F. OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC. 2025 Workshop on Extreme Heterogeneity and AI Convergence in HPC at the International Conference for High Performance Computing, Networking, Storage and Analysis (SC’25 workshops). [* Equal contribution] [pdf].
- Conference proceedings
- Tyagi, S., & Wang, F. Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training. IEEE/ACM SIGHPC 26th International Symposium on Cluster, Cloud and Internet Computing (CCGrid ‘26) (Accpt. rate 25%) [pdf].
- Tyagi, S., & Swany, M. Flexible Communication for Optimal Distributed Learning over Unpredictable Networks. 2023 IEEE International Conference on Big Data (BigData ‘23), 925-935 (Accpt. rate 17.5%) [pdf].
- Tyagi, S., & Swany, M. Accelerating Distributed ML Training via Selective Synchronization. 2023 IEEE International Conference on Cluster Computing (CLUSTER ‘23), 1-12 (Accpt. rate 25%) [pdf].
- Tyagi, S., & Swany, M. GraVAC: Adaptive Compression for Communication-Efficient Distributed DL Training. 2023 IEEE 16th International Conference on Cloud Computing (CLOUD ‘23), 319-329 (Accpt. rate 20%) [pdf].
- Tyagi, S., & Sharma, P. Scavenger: A Cloud Service For Optimizing Cost and Performance of ML Training. 2023 IEEE/ACM 23rd International Symposium on Cluster, Cloud and Internet Computing (CCGrid ‘23), 403-413 (Accpt. rate 21%) [pdf].
- Tyagi, S., & Swany, M. ScaDLES: Scalable Deep Learning over Streaming data at the Edge. 2022 IEEE International Conference on Big Data (BigData ‘22), 2113-2122 (Accpt. rate 19.2%) [pdf].
- Tyagi, S., & Sharma, P. Taming Resource Heterogeneity In Distributed ML Training With Dynamic Batching. 2020 IEEE International Conference on Autonomic Computing and Self-Organizing Systems (ACSOS ‘20), 188-194 (Accpt. rate 25%) [pdf].
- Widanage, C., Li, J., Tyagi, S., Teja, R., Peng, B., Kamburugamuve, S., Baum, D., Smith, D., Qiu, J., & Koskey, J. Anomaly Detection over Streaming Data: Indy500 Case Study. 2019 IEEE 12th International Conference on Cloud Computing (CLOUD ‘19), 9-16 (Accpt. rate 20%) [pdf].
- Chaturvedi, S., Tyagi, S., & Simmhan, Y. Collaborative Reuse of Streaming Dataflows in IoT Applications. 2017 IEEE 13th International Conference on e-Science (e-Science ‘17), 403-412 (Accpt. rate 36%) [pdf].
- Journal articles
- Tyagi, S., & Sharma, P. OmniLearn: A Framework for Distributed Deep Learning over Heterogeneous Clusters. IEEE Transactions on Parallel and Distributed Systems (TPDS ‘25) (IF 5.6) [pdf].
- Chaturvedi, S., Tyagi, S., & Simmhan, Y.L. (2019). Cost-Effective Sharing of Streaming Dataflows for IoT Applications. IEEE Transactions on Cloud Computing (TCC ‘19), 9, 1391-1407 (IF 5.3) [pdf].
Employment
- 2025-now: Postdoctoral Research Associate, AAIMS Group, National Center for Computational Sciences (NCCS), Oak Ridge National Laboratory (ORNL), USA.
- 2017-2018: Research Staff Member, DREAM:Lab, Dept. of Computational and Data Sciences (CDS), Indian Institute of Science, India.
- 2016-2017: Data Scientist, RocQ Mobile App Analytics, HT Media Ltd., India.
- 2015-2016: Big Data Engineer, Stayzilla Ltd., India.
- 2014-2015: Software Engineer, Tatras Data Ltd., India.
Teaching
- Associate Instructor (IUB), High-Performance Computing (ENGR-E317/517): Spring 2024
- Associate Instructor (IUB), Computer Networks (ENGR-E318/518, CSCI-P438/538): Fall 2022/2023/2024
- Associate Instructor (IUB), Operating Systems (ENGR-E319/519, CSCI-P436/536): Spring 2023
- Associate Instructor (IUB), Distributed Systems (ENGR-E510, CSCI-B534): Spring 2021/2022
- Associate Instructor (IUB), Cloud Computing (ENGR-E516): Fall 2019/2020/2021
Services
- 2024: Reviewer/Sub-reviewer for IEEE CLUSTER, Journal of Parallel and Distributed Computing (JPDC), USENIX Operating Systems Design and Principles (OSDI) and USENIX Annual Technical Conference (ATC) (Artifact evaluation committee).
- 2025: Reviewer/Sub-reviewer/Program committee for International Joint Conference on Neural Networks (IJCNN), Euro-Par, USENIX Operating Systems Design and Principles (OSDI) (Artifact evaluation committee), Journal of Systems Architecture (JSA), IEEE eScience, Elsevier Neurocomputing, Journal of Systems Research (JSys) Artifact evaluation board, AI4S workshop @Supercomputing (SC25), International Parallel & Distributed Processing Symposium (IPDPS).
- 2026: Reviewer/Sub-reviewer/Program committee for SC26 (Data Analytics, Visualization, & Storage track), Journal of Parallel and Distributed Computing (JPDC), Journal of Systems Research (JSys) Artifact evaluation board, IEEE eScience, International Parallel & Distributed Processing Symposium (IPDPS), Workshop organizer FPAI-HPC’26 @IEEE Cluster’26.
Awards
- NSF Student Grant: To present research at IEEE CLUSTER 2023, Santa Fe, New Mexico.
- Luddy Dean’s Graduate Student Award: In Fall 2023 for outstanding research.
- NSF Travel Award: To present research at IEEE/ACM CCGrid 2023, Bengaluru, India.
- Best early-career research poster award: Awarded at IEEE/ACM CCGrid 2023 [pdf].
- Google Cloud Student Researcher (2021, 2022): Received $2000 Google Cloud credits for research.
- Student Funding: Partially funded in graduate school via National Science Foundation (NSF) grants Data Infrastructure Building Blocks (DiBBS) 17-500 and OAC-2112606.
- NERSC Compute: Received compute hours on NERSC’s Perlmutter system (2026).
Talks
- 06/2026: Invited talk, “OmniFed: Towards Configurable Cross-Silo Federated Learning”, Trillion Parameter Consortium 2026, Baltimore, Maryland, USA [slides].
- 05/2026: Paper presentation, “Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training”, IEEE/ACM SIGHPC 26th International Symposium on Cluster, Cloud and Internet Computing, Sydney, Australia [slides].
- 04/2026: Poster presentation, “OmniFed: Towards Configurable Cross-Silo Federated Learning”, The Future of Computing, Collaboration Catalyst session, ORNL, Oak Ridge, TN, USA.
- 11/2025: Paper presentation, “OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC”, Extreme Heterogeneity and AI Convergence in HPC Workshop, SC25, St. Louis, MO. [slides].
- 08/25: Poster presentation, “OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC” at ORNL Software and Data Expo (OSDX)
- 08/2025: Presented my work on a modular federeated learning framework at ORNL’s Federated and Collaborative Learning for Exascale Science and Security Workshop, Tennessee, USA.
- 07/2025: Research presentation, “Enabling Large-Batch Training via Learned Gradient Mapping” at the 13th Annual ORPA Research Symposium, ORNL, Tennessee, USA.
- 09/2024: Invited talk, “Improving Communication in Federated Learning via Adaptive Gradient Compression”, INRIA, France.
- 09/2024: Invited talk, “Optimizing Compute and Communication in Deep Learning Systems”, Oak Ridge National Laboratory (ORNL), Tennessee, USA.
- 04/2024: Guest lectures, “Parallel Computing with GPUs for Distributed ML Applications”, High-Performance Computing (HPC) course, Indiana University Bloomington, USA [pdf1], [pdf2].
- 12/2023: Paper presentation, “Flexible Communication for Optimal Distributed Learning over Unpredictable Networks.” 2023 IEEE International Conference on Big Data, Sorrento, Italy [slides].
- 11/2023: Paper presentation, “Accelerating DistributedMLTraining via Selective Synchronization.” 2023 IEEE International Conference on Cluster Computing, Santa Fe, New Mexico, USA [slides].
- 11/2023: Poster presentation, “Accelerating Distributed ML Training via Selective Synchronization.” 2023 IEEE International Conference on Cluster Computing, Santa Fe, New Mexico, USA [poster].
- 09/2023: Invited talk, “Towards building efficient computation and communication models for distributed deep learning systems.” Mathematics and Computer Science (MCS) division, Argonne National Laboratory, Illinois, USA.
- 07/2023: Paper presentation, “GraVAC: Adaptive Compression for Communication-Efficient Distributed DL Training.” 2023 IEEE International Conference on Cloud Computing, Chicago, Illinois [slides].
- 05/2023: Paper presentation, “Scavenger: A Cloud Service for Optimizing Cost and Performance of ML Training.” 2023 IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing, Bengaluru, India [slides].
- 05/2023: Poster presentation, “Scavenger: A Cloud Service for Optimizing Cost and Performance of ML Training.” 2023 IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing, Bengaluru, India [poster].
- 12/2022: Paper presentation, “ScaDLES: Scalable Deep Learning over Streaming Data at the Edge.” 2022 IEEE International Conference on Big Data, Osaka, Japan [slides].
- 07/2020: Paper presentation, “Taming Resource Heterogeneity in Distributed ML Training with Dynamic Batching.” 2020 IEEE International Conference on Autonomic Computing and Self-Organizing Systems, virtual [slides].
- 11/2018: “Real-Time Anomaly Detection from Edge to HPC-Cloud”, Intel Speakerships at SC18 (Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis 2018), Dallas, Texas, USA [slides].
