I am a postdoctoral researcher in the Multimodal AI Lab at KAIST, working with Professor Joon Son Chung. I completed the integrated Master's/Ph.D. program in the School of Electrical Engineering at KAIST in 2026, advised by Professor Joon Son Chung.
I received my B.S. in Electronics and Information Engineering from Harbin Institute of Technology (HIT), where I worked with Professor Sheng Chang Lan on mmWave radar-based hand gesture recognition.
I interned at Huawei from January to July 2026.
A data-free view of flow-map distillation, followed by a step-by-step derivation of how FreeFlow's regression objective becomes a velocity-alignment update.
We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier-based or classifier-free guidance, MGAudio enables the generative model to guide itself through a dedicated training objective designed for video-conditioned audio generation.
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
Chenshuang Zhang, Kang Zhang, Joon Son Chung, In So Kweon, Junmo Kim, Chengzhi Mao
Neurips, 2025
We propose a self-supervised tracker that leverages motion representations inherently learned by pre-trained video diffusion models during early denoising, enabling robust distinction of visually similar objects and achieving up to 6-point improvements over prior methods on benchmarks and new tests.
We present X-MDPT, a novel diffusion model designed for pose-guided human image generation. X-MDPT distinguishes itself by employing masked diffusion transformers that operate on latent patches, a departure from the commonly-used Unet structures in existing works.
We introduce Physics Informed Distillation (PID), which employs a student model to represent the solution of the ODE system corresponding to the teacher diffusion model, akin to the principles employed in PINNs.
We propose BI-MDRG that bridges the response generation path such that the image history information is utilized for enhanced relevance of text responses to the image content and the consistency of objects in sequential image responses.
To speed up inference and further enhance the performance, our research revisits diffusion models in image super-resolution and proposes a straightforward yet significant diffusion model-based super-resolution method called ACDMSR.
Actual image super-resolution is an extremely challenging task due to complex degradations existing in the image. To solve this problem, two dominant methodologies have emerged: degradation-estimation-based Addressing actual image super-resolution remains a formidable challenge due to the intricate degradations present in images.
To facilitate human computer interaction (HCL) for the community with deafness and hearing loss (D&HL), this article explored the feasibility of recognizing a vocabulary of dynamic Chinese sign language (CSL) based on millimeter-wave (mmWave) radar sensors within the scope of data science.
To further improve the performance and simplify current DPM-based super-resolution methods, we propose a simple but non-trivial DPM-based super-resolution post-process framework i.e. cDPMSR.
This work discards prior practices of directly introducing AT to SSL frameworks and proposed a two-stage framework termed Decoupled Adversarial Contrastive Learning (DeACL).
We point out that InfoNCE loss used in MoCo implicitly attract anchors to their corresponding positive sample with various strength of penalties and identify such inter-anchor hardness-awareness property as a major reason for the necessity of a large dictionary. Our findings motivate us to simplify MoCo v2 via the removal of its dictionary as well as momentum.
Two optimal types of deep neural networks, 3D-CNN and CNN-LSTM are respectively constructed to reveal the temporal gesture motion signatures encoded in multiple adjacent radar chirps.
We investigate the feasibility of using a three-dimensional Doppler-radar array at 24GHz to recognize human gestures with a model consisted of ten classical gestures.
A modified Vivaldi antenna with low self-reflectivity working at 1-4 GHz is designed to improve the radiation gain by opening rectangular grooves and loading parasitic patches on the surface of the radiating patch.