|
Joongwon (Daniel) Kim
I am a final-year Ph.D. student in the Natural Language Processing group at the University of Washington, and a visiting researcher at Meta Superintelligence Labs (FAIR). I am thankful to be advised by Hannaneh Hajishirzi.
Previously, I was an undergrad at the University of Pennsylvania, working with Chris Callison-Burch and Mark Yatskar.
My research focuses on language model post-training, with a focus on agents, reinforcement learning and synthetic data. I am also supported by the NSF-GRFP Fellowship.
Update: I am currently on the industry job market and will graduate from my PhD in December 2026. Please feel free to reach out via email or LinkedIn if you are interested!!
News
CV  / 
Email  / 
LinkedIn  / 
Google Scholar  / 
X/Twitter
|
|
Publications / Pre-Prints
|
|
Scaling Test-Time Compute for Agentic Coding
Joongwon Kim, Winnie Yang, Kelvin Niu, Hongming Zhang, Yun Zhu, Eryk Helenowski, Ruan Silva, Zhengxing Chen, Srini Iyer, Manzil Zaheer, Daniel Fried, Hannaneh Hajishirzi, Sanjeev Arora, Gabriel Synnaeve, Ruslan Salakhutdinov, Anirudh Goyal
COLM, 2026
Paper
We introduce a test-time scaling framework for coding agents that represents long rollout trajectories using structured summaries as reusable prior experiences. Recursive Tournament Voting selects promising attempts, while Parallel-Distill-Refine uses their summaries to perform sequential refinement. Our approach improves all five evaluated models on SWE-Bench Verified and Terminal-Bench v2.0, raising Claude 4.5 Opus from 70.9% to 77.6% and from 47.0% to 59.1%, respectively. Work done with the agentic post-training research team in FAIR at Meta.
|
|
Prompt Curriculum Learning for Efficient LLM Post-Training
Zhaolin Gao, Joongwon Kim, Wen Sun, Thorsten Joachims, Sid Wang, Richard Pang, Liang Tan
ICLR, 2026
Paper
We introduce Prompt Curriculum Learning (PCL), which uses a learned value model to select intermediate-difficulty prompts for language model post-training without costly filtering rollouts. PCL allows the policy to sample increasingly difficult prompts in an on-policy manner while maintaining a high effective batch size ratio throughout training, improving both mathematical reasoning performance and training speed. Work done with the Llama post-training team at Meta.
|
|
ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context
Joongwon Kim, Anirudh Goyal, Liang Tan, Hannaneh Hajishirzi, Srinivasan Iyer, Tianlu Wang
Preprint, 2025
Paper
We introduce ASTRO, a framework that teaches language models to reflect, backtrack, and explore alternative reasoning paths. We convert Monte Carlo Tree Search traces into natural-language reasoning examples that capture both successful steps and recovery from mistakes, then improve the models with reinforcement learning using verifiable rewards. ASTRO substantially improves mathematical reasoning in Llama 3 models, particularly on challenging problems that require iterative correction. Work done in collaboration between FAIR and Llama team at Meta.
|
|
|
A Systematic Examination of Preference Learning through the Lens of Instruction-Following
Joongwon Kim, Anirudh Goyal, Aston Zhang, Bo Xiong, Rui Hou, Melanie Kambadur, Dhruv Mahajan, Hannaneh Hajishirzi, Liang Tan
Preprint  
Paper  
We systematically investigate how preference alignment is affected by various attributes of the training set with a focus on instruction-following.
To this end, we employ a novel synthetic data generation pipeline to generate prompts which incorporate multiple verifiable constraints.
We use rejection sampling and MCTS to generate preference pairs, and we perform experiments that investigate the effects of (1) shared prefixes, (2) the contrast and quality of the responses, and (3) the complexity of the training prompts.
Work done in the Llama post-training team at Meta.
|
|
|
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
Joongwon Kim, Bhargavi Paranjape, Tushar Khot, Hannaneh Hajishirzi
Preprint  
Paper  | 
Code  | 
Models  | 
Website
We introduce Husky, an open-source language agent that learns to reason over a unified action space to address multi-step tasks involving numerical, tabular, and knowledge-based reasoning.
Our experiments show that Husky outperforms prior language agents across 14 evaluation sets.
Moreover, we present new evaluation sets that require mixed-tool reasoning and show that Husky matches or even exceeds frontier LMs such as GPT-4 on these tasks despite only using 7B models.
|
|
|
TaskWeb: Selecting Better Source Tasks for Multi-task NLP
Joongwon Kim, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi
Proceedings of EMNLP, 2023 (long)  
Paper  | 
Code  | 
Website  | 
Video  | 
Poster
We introduce TaskWeb, our benchmark of pairwise task transfers between 22 different NLP tasks across three different model types, sizes and adaptation method.
Based on TaskWeb, we propose a new method TaskShop for estimating transferability between source and target tasks with only a small number of target examples.
We demonstrate that selecting helpful source tasks with our method allows us to perform multi-task learning on much smaller training sets and still improve zero-shot performance across various target tasks.
|
|
|
Induce, Edit, Retrieve: Language Grounded Multimodal Schema for Instructional Video Retrieval
Yue Yang, Joongwon Kim, Artemis Panagopolou, Mark Yatskar, Chris Callison-Burch
CVPR 2022 @ ODRUM, 2022 (spotlight talk)  
Paper
We built schemas for goal-oriented tasks by aligning YouTube videos with wikiHow steps. Then, we proposed methods for editing the schemas to handle unseen but related tasks.
Finally, we leveraged our schemas to perform instructional video retrieval on several datasets and demonstrated that our method improves over other retrieval approaches.
|
|
|
BiSECT: Learning to Split and Rephrase Sentences with Bitexts
Joongwon Kim*, Mounica Maddela*, Reno Kriz, Wei Xu, Chris Callison-Burch
Proceedings of EMNLP, 2021 (long)  
Paper  | 
Code
We curated a multilingual corpus for sentence splitting by using machine translation over parallel corpora. Moreover, we developed a
sentence splitter with controllable generation. We showed that our dataset and model outperformed existing methods in both automatic and human evaluations. Work done in collaboration with Georgia Tech.
|
|