[Selected Papers]   [Full List]   [Google Scholar]   [DBLP]   [arXiv]

Collaborators.

Throughout my journey, I have been incredibly fortunate to work alongside some of the brilliant minds in the field, whose wisdom has been instrumental in shaping my work and enriching my knowledge fundamentally: Chiranjib Bhattacharyya, Avrim Blum, Christos Dimitrikakis, Simon Du, Maryzam Fazel, Vitaly Feldman Pierre Gaillard, Aditya Gopalan, Katja Hofmann Eyke Hüllermeier, Prateek Jain, Tomer Koren, Branislav Kveton, Akshay Krishnamurthy, Haipeng Luo, Shie Mannor, Yishay Mansour, Praneeth Netrapalli, Vianney Perchet, Lev Reyzin, Rob Schapire, Nati Srebro, Michal Valko, Matthew Walter Haifeng Xu, Brian Ziebert (in alphabetical order).

Preprints/ Working Drafts:
  • Variance Adaptive Bandits under Cost Budget
    Amith Bhat, Aadirupa Saha
  • Imitation Beyond Expectation via Second-order Stochastic Dominance
    Syed M. Abbas, Danyal Saeed, Aadirupa Saha, Brian D Ziebart
  • Self-Consistency on a Budget
    Aniket Wadge, Branislav Kveton, Aadirupa Saha
  • Sample Efficient Policy Optimization with Different Feedback Modalities
    Aadirupa Saha, Pierre Gaillard
  • Preference Learning Beyond Confidence: Gradient-based RLHF with Adversarial Preferences
    Aadirupa Saha
  • Optimal Learning Rates with Costly Observations
    Lev Reyzin, Aadirupa Saha (alphabetical)
  • Optimal Rates for Learning Quantum States with Linear Tomography [Arxiv]
    Moise Blanchard, Dmitry Ostrovsky, Aadirupa Saha (alphabetical)
  • Double-Monster: Efficient Min-Max Strategy for Personalized Prediction under General Preferences.
    Aadirupa Saha, Robert Schapire
  • Efficient Predictive Models without Compromising User Privacy [Arxiv version]
    Aadirupa Saha, Hilal Asi
  • Learning to Allocate Resources with Censored Feedback [Arxiv version]
    Giovanni Montanari, Côme Fiegel, Aadirupa Saha, Vianney Perchet
  • COBRA: Contextual Bandit Algorithm for Ensuring Truthful Strategic Agents
    Arun Verma, Indrajit Saha, Aadirupa Saha, Daniela Rus, Makoto Yokoo, Bryan Kian Hsiang Low
  • Best Arm Identification in Linear MNL-Bandits. [Arxiv Version]
    Shubham Gupta, Aadirupa Saha, Sumeet Katariya
Full list of Publications:2026
2025
  • Efficient and Near-Optimal Algorithm for General Contextual Dueling Bandits with Offline Regression Oracles [NeurIPS version]
    Aadirupa Saha, Robert Schapire
    In Neural Information Processing Systems, NeurIPS 2025
  • Imitation Beyond Expectation Using Pluralistic Stochastic Dominance [NeurIPS version]
    Ali Farajzadeh, Danyal Saeed, Syed M Abbas, Rushit N. Shah, Aadirupa Saha, Brian D Ziebart
    In Neural Information Processing Systems, NeurIPS 2025 (*Spotlight*)
  • Source Adaptive Online Learning under Heteroscedastic Noise [OPT-ML version]
    Amith Bhat, Aadirupa Saha, Thomas Kleine Buening, Haipeng Luo
    In OPT for ML Workshop, Neural Information Processing Systems, NeurIPS 2025
  • Efficient Algorithms for Combinatorial-Bandits with Monotonicity. [OPT-ML version]
    Aniket Wadge, Aadirupa Saha
    In OPT for ML Workshop, Neural Information Processing Systems, NeurIPS 2025
  • HPO: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration [Arxiv version] [Workshop version]
    Avinandan Bose, Zhihan Xiong, Aadirupa Saha, Simon Shaolei Du, Maryam Fazel
    In The Next Frontier in Reliable AI Workshop, International Conference on Learning Representations (ICLR), 2025
  • Tracking the Best Expert Privately [Arxiv version]
    Hilal Asi, Vinod Raman, Aadirupa Saha (alphabetical)
    International Conference on Machine Learning, ICML 2025
  • Dueling Convex Optimization for General Preferences: An Unified Framework for Optimal Convergence Rates. [Arxiv Version]
    Aadirupa Saha, Tomer Koren, Yishay Mansour
    International Conference on Machine Learning, ICML 2025
  • Stop Relying on No-Choice and Do not Repeat the Moves: Optimal, Efficient and Practical Algorithms for Assortment Optimization [Arxiv version]
    Aadirupa Saha, Pierre Gaillard
    International Conference on Learning Representations (ICLR), 2025
    1 min pitch! If you've ever encountered MNL Assortment Optimization problem, the go-to approach is to offer the same set of products repeatedly until your customer is really annoyed and selects no item! In fact, it requires a "belief" that no selection is their most preferred choice :-( Oh no!

    But why be so pessimistic? And why annoy your customers repeatedly offering the same items and hoping them to leave (i.e. they decide to choose none of the offered items!)? We got a new idea with no such issues. How? We simply found better concentration tricks! It was a long time wish to resolve this efficiently.
2024
  • Strategic Linear Contextual Bandits. [Arxiv version]
    Thomas Kleine Buening, Aadirupa Saha, Haifeng Xu, Christos Dimitrakakis
    In Neural Information Processing Systems, NeurIPS 2024
  • Dueling in the Dark: An Efficient and Optimal O(√T) Mirror Descent Approach for Competing against Adversarial Preferences. [Arxiv: Coming soon!]
    Aadirupa Saha, Barry-John Theobald, Yonathan Efroni
    In OPT for ML Workshop, Neural Information Processing Systems, NeurIPS 2024
  • A Graph Theoretic Approach for Preference Learning with Feature Information. [Arxiv Version]
    Aadirupa Saha, Arun Rajkumar
    In Uncertainty in Artificial Intelligence, UAI 2024 (*Oral*)
  • Social Welfare for RecSys: Bandits Meet Mechanism Design to Combat Clickbait in Online Recommendation. [Arxiv version]
    Thomas Kleine Buening, Aadirupa Saha, Haifeng Xu, Christos Dimitrakakis
    International Conference on Learning Representations (ICLR), 2024 (*Spotlight*)
  • Only Pay for What Is Uncertain: Variance-Adaptive Thompson Sampling. [Arxiv version]
    Aadirupa Saha, Branislav Kveton
    International Conference on Learning Representations (ICLR), 2024
    1 min pitch! We lay the foundations for Bayesian multi-armed bandits with known and unknown heterogeneous reward variances with Thompson sampling. Our regret analysis shows improved performance with lower reward variances, implying faster learning in low-variance regimes. So why regret if you are already confident - Only Pay for What Is Uncertain!
  • Efficient Private Federated Non-Convex Optimization With Shuffled Model. [Workshop Version]
    Lingxiao Wang, Xingyu Zhou, Kumar Kshitij Patel, Lawrence Tang, Aadirupa Saha
    Privacy Regulation and Protection in ML Workshop, International Conference on Learning Representations (ICLR), 2024
  • Think Before You Duel: Understanding Complexities of Preference Learning under Constrained Resources. [Arxiv version]
    Rohan Deb, Aadirupa Saha
    International Conference on Artificial Intelligence and Statistics, AIStats 2024
  • On the Vulnerability of Fairness Constrained Learning to Malicious Noise. [Arxiv version]
    Avrim Blum, Princewill Okoroafor, Aadirupa Saha, Kevin Stangl (alphabetical)
    International Conference on Artificial Intelligence and Statistics, AIStats 2024
  • Faster Convergence with MultiWay Preferences. [Arxiv version]
    Aadirupa Saha, Vitaly Feldman, Tomer Koren, Yishay Mansour
    International Conference on Artificial Intelligence and Statistics, AIStats 2024
  • Dueling Optimization with a Monotone Adversary
    Avrim Blum, Meghal Gupta, Gene Li, Naren Sarayu Manoj, Aadirupa Saha, Yuanyuan Yang
    Algorithmic Learning Theory, ALT, 2024 (*Outstanding Paper Award*)
2023
  • Eliciting User Preferences for Personalized Multi-Objective Decision Making through Comparative Feedback [Arxiv Version]
    Han Shao, Lee Cohen, Avrim Blum, Yishay Mansour, Aadirupa Saha, Mathew Walter
    In Neural Information Processing Systems, NeurIPS 2023
  • Dueling Optimization with a Monotone Adversary. [Arxiv version]
    Avrim Blum, Meghal Gupta, Gene Li, Naren Sarayu Manoj, Aadirupa Saha, Yuanyuan Yang
    NeurIPS OPT+ML Workshop, NeurIPS, 2023 (Oral)
  • On the Vulnerability of Fairness Constrained Learning to Malicious Noise. [Arxiv version]
    Avrim Blum, Princewill Okoroafor, Aadirupa Saha, Kevin Stangl
    Algorithmic Fairness through the Lens of Time Workshop, NeurIPS, 2023
  • Federated Online and Bandit Convex Optimization [Arxiv Version]
    Kumar Kshitij Patel, Lingxiao Wang, Aadirupa Saha, Nati Srebro
    In the International Conference on Machine Learning, ICML 2023
  • Bandits Meet Mechanism Design to Combat Clickbait in Online Recommendation [Arxiv Version]
    Thomas Kleine Buening, Aadirupa Saha, Haifeng Xu, Christos Dimitrakakis
    Interactive Learning with Implicit Human Feedback Workshop, ICML 2023
  • One Arrow, Two Kills: An Unified Framework for Achieving Optimal Regret Guarantees in Sleeping Bandits [Arxiv Version] [Talk]
    Pierre Gaillard, Aadirupa Saha, Soham Dan
    In International Conference on Artificial Intelligence and Statistics, AIStats 2023
  • 1 min pitch! Sleeping Bandits are as interesting as they sound, but what is the right measure of Sleeping Regret? So many different notions of regrets were studied in the literature --- Sleeping External regret, Ordering regret, Policy regret --- but it is confusing to keep track of the implications of so many different notions, i.e. every combination of stochastic or adversarial losses and availability pairs.

    Can we unify them under a single measure? We found one in this work - Sleeping Internal Regret! One of our main contributions is unifying existing notions of regret in sleeping bandits and exploring their implications for each other. 

    Our proposed algorithm achieves sublinear Internal Regret, even when losses and availabilities are both adversarial, which is the hardest combination of sleeping setup! Further, our results show how a low internal regret leads to both low external regret and low policy regret - One arrow, Two Kills! 

    Our unified notion of sleeping regret also helps to invent a general notion of Sleeping Dueling Bandits that is stronger than the existing regret definitions used in the contemporary dueling bandits literature and overcomes the issue of repeated draws if needed. This is the first bound of this kind in the dueling literature with many potentials!
  • ANACONDA: Improved Dynamic Regret Algorithm for Adaptive Non-Stationary Dueling Bandits [Arxiv Version]
    Thomas Kleine Buening, Aadirupa Saha
    In International Conference on Artificial Intelligence and Statistics, AIStats 2023
  • Dueling RL: Reinforcement Learning with Trajectory Preferences [Arxiv Version]
    Aadirupa Saha*, Aldo Pacchiano*, Jonathan Lee (*Equal contribution)
    In International Conference on Artificial Intelligence and Statistics, AIStats 2023
2022
2021
2020
2019
2018
2015
2014
2013
2011

   [Back to Top]   [Selected Papers]   [Full List]   [Google Scholar]   [DBLP]   [arXiv]