Subhojyoti Mukherjee

I am a research scientist at Adobe Research. My expertise ranges from research and developing algorithms to training machine learning models, Reinforcement Learning, fine-tuning and alignment for LLMs. Download CV
Email: subhomuk [at] adobe [dot] com

Work Experience

Adobe Research (San Jose)
Research Scientist/Engineer
(Mar 2025 - Present)
Product:
  1. Agentic post-training: Express Agent Orchestrator (Adobe Express)
    • Led post-training of the small VLM deployed in production as the orchestrator behind the Express AI Assistant. The model interprets user intent, plans multi-step tool calls, and executes editing actions inside Express (see Adobe MAX 2025 Keynote); patent filed.
    • Built the synthetic data generation pipeline for tool-use and planning trajectories, used to scale training across [N] tools / action types.
    • Aligned automated evaluation with human feedback. Used these custom evals to find failure modes in intent understanding and tool planning, then closed them through targeted data and training-recipe changes validated with controlled ablations.
  2. RL post-training and reward modeling for Firefly Gen5/Gen6/6.5 image/video editing
    • Contributed to RL post-training of Adobe Firefly Gen5 and Gen6 editing models, comparing RL optimization methods ([e.g., GRPO, DPO variants]).
    • Built forward-process RL (AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models) into the Firefly foundation team's RL post-training pipeline for Gen6/6.5. Compared with reverse-process policy gradients (Flow-GRPO), it avoids credit assignment across denoising steps, and it reaches DiffusionNFT's best reward in under half the training time.
    • Built internal reward models for lighting, camera angle, and harmonization, and calibrated them against human preference judgments for editing tasks.
    • This work also motivated Stepwise-Flow-GRPO (CVPR 2026), which adds stepwise credit assignment for GRPO on flow-matching models.
  3. Synthetic data and RL for small on-device LMs (Adobe Document Cloud)
    • Spanish Document Overview model in Acrobat Reader. Led RL post-training for multilingual adaptation. Generated Spanish synthetic data with SmolLM3-3B, scored it with reward models and task-specific rubrics, and post-trained the deployed model on it; patent filed.
    • English Document Overview model in Acrobat Reader (Oct 2025). Led pre-training and post-training of the on-device model. Built part of the pre-training mix by combining Nemotron-CC with synthetic data generated from GPT-OSS-120B; patent filed.

Research Areas: My research focuses on post-training alignment and agentic planning for LLMs and VLMs, developing methods that simultaneously advance deployed products and scientific understanding of reasoning, reward modeling, and sequential decision-making.

Mentoring: Mentoring interns on projects spanning video editing, image editing, and LLM/VLM alignment.

Education

Ph.D.
(Fall 2019 to Feb 2025)
at ECE, University of Wisconsin Madison
advised by Dr. Robert Nowak, Dr. Josiah Hanna, and Dr. Qiaomin Xie

Areas of Research: Reinforcement Learning, Active Learning, incorporating deep active learning strategies for Large Language Models (LLMs), aligning Large Language Models with human feedback (RLHF), and understanding sequential decision-making using transformers (DT).

PhD Thesis: Adaptive Data Collection for Policy Evaluation, Multi-task Learning and LLM Alignment pdf

(Joint) Masters Thesis: Active Sequential Hypothesis Testing with Extension to Active Regression and Multi-armed Bandits pdf
M.S by Research
(2015 to 2018)
at CSE, Indian Institute of Technology (IIT) Madras
advised by Dr. Balaraman Ravindran, and Dr. Nandan Sudarsanam
RISE Lab

Areas of Research: Reinforcement learning, Stochastic and non-stochastic Multi-Armed Bandit settings.

Masters Thesis: Finite-time Analysis of Frequentist Strategies for Multi-armed Bandits pdf
Bachelor of Technology
(2009 to 2013)
at Dept. of Computer Science and Engineering
Meghnad Saha Institute of Technology, Kolkata
under West Bengal University of Technology, India

Research Internships

Amazon AWS AI, Santa Clara, USA
Summer 2024 (full-time)
hosted by Branislav Kveton, Anusha Lalitha
and: Sailik Sengupta, Yifei Ma, Aniket Deshmukh, Gaurush Hiranandani.

Area of Research: Multi-objective alignment for LLMs.
Amazon AWS AI, Santa Clara, USA
Fall 2023 (Part-time)
hosted by Branislav Kveton
and: Yifei Ma, Anusha Lalitha, Kousha Kalantiri, Ge Liu, Aniket Deshmukh, Anoop Deoras.

Area of Research: RLHF with LLMs.
Amazon AWS AI, Santa Clara, USA
Summer 2023 (Full-time)
hosted by Branislav Kveton
and: Yifei Ma, Anusha Lalitha, Ge Liu, Aniket Deshmukh, Anoop Deoras.

Area of Research: Active In-Context Learning with LLMs.
CMU, ECE Dept., Pittsburgh, USA
Summer 2019
hosted by Prof. Gauri Joshi
Area of Research: Structured Bandits.
Adobe Research, San Jose, USA
Spring 2018
hosted by Branislav Kveton
Area of Research: Item recommendation with Ranking and Bandits.
INRIA, SequeL Lab, Lille, France
Fall 2017
hosted by Odalric Maillard
Area of Research: Non-stationary Bandits.

News

2026

2025

2024

2023

2022

2021

2020

2019