Research Collaborator
I work on vision-language-action models, multimodal generation, and learning from action-free data.
I'm an incoming MS student at ALIN Lab (KAIST), starting Fall 2026, advised by Prof. Jinwoo Shin.
I believe the fusion of modalities—centered around language models—is essential for AI to positively impact society. I'm especially interested in VLA models that can understand and learn from action-label-free datasets, and in the effective representation learning of generative models.
Most recent work on Google Scholar. ‡ indicates equal contribution.
RLDX-1 Technical Report
RLWRLD Technical Report 2026
While Vision-Language-Action models (VLAs) have shown remarkable progress toward human-like generalist robotic policies through the versatile intelligence (broad scene understanding and language-conditioned generalization) inherited from pre-trained Vision-Language Models, they still struggle with complex real-world tasks requiring broader functional capabilities such as motion awareness, long-term memory, and physical sensing. To address this, we introduce RLDX-1, a general-purpose robotic policy for dexterous manipulation built on the Multi-Stream Action Transformer (MSAT), an architecture that unifies these capabilities by integrating heterogeneous modalities through modality-specific streams with cross-modal joint self-attention. RLDX-1 further combines this architecture with system-level design choices, including data synthesis for rare manipulation scenarios, learning procedures specialized for human-like manipulation, and inference optimizations for real-time deployment. Through empirical evaluation, we show that RLDX-1 consistently outperforms recent frontier VLAs across both simulation benchmarks and real-world tasks that require broad functional capabilities beyond general versatility. In particular, RLDX-1 shows superiority in ALLEX humanoid tasks by achieving success rates of 86.8%, highlighting its ability to control a high-DoF humanoid robot under diverse functional demands. Together, these results position RLDX-1 as a promising step toward reliable VLAs for complex, contact-rich, and dynamic real-world dexterous manipulation.
Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?
ACL 2025 Poster
Any-to-any generative models aim to enable seamless interpretation and generation across multiple modalities within a unified framework, yet their ability to preserve relationships across modalities remains uncertain. Do unified models truly achieve cross-modal coherence, or is this coherence merely perceived? To explore this, we introduce ACON, a dataset of 1,000 images (500 newly contributed) paired with captions, editing instructions, and Q&A pairs to evaluate cross-modal transfers rigorously. Using three consistency criteria—cyclic consistency, forward equivariance, and conjugated equivariance—our experiments reveal that any-to-any models do not consistently demonstrate greater cross-modal consistency than specialized models in pointwise evaluations such as cyclic consistency. However, equivariance evaluations uncover weak but observable consistency through structured analyses of the intermediate latent space enabled by multiple editing operations.
VLM Safety Evaluation Must Move from Classification to Causal Controllability
Coming soon...
InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?
KCC 2026 Oral
Recent advances in Vision-Language Models (VLMs) have demonstrated impressive capabilities in perception and reasoning. However, the ability to perform causal inference—a core aspect of human cognition—remains underexplored, particularly in multimodal settings. In this study, we introduce InfoCausalQA, a novel benchmark designed to evaluate causal reasoning grounded in infographics that combine structured visual data with textual context. The benchmark comprises two tasks: quantitative causal reasoning based on inferred numerical trends, and semantic causal reasoning involving five types of causal relations (cause, effect, intervention, counterfactual, and temporal). We manually collected 494 infographic-text pairs from four public sources and used GPT-4o to generate 1,482 high-quality multiple-choice QA pairs, then carefully revised by humans so they cannot be answered from surface-level cues alone but instead require genuine visual grounding. Our results reveal that current VLMs exhibit limited capability in computational reasoning and even more pronounced limitations in semantic causal reasoning, with performance significantly lower than humans. Through InfoCausalQA, we highlight the need for advancing the causal reasoning abilities of multimodal AI systems.
Research Collaborator
Engineering Lead
Collaborative Language Model Engineer
SK AI Summit AI's Got Talent
1st PrizeAwarded with the enhanced version of A-EYE.
SW-Centered University Digital Competition
Grand Prize (2nd Place, Award of the Chief of IITP)Conducted a poster session on 2025 SK AI Summit!
Developed A-EYE, an AI-powered real-time environmental recognition system for the visually impaired, enabling 3-second hazard detection, GPS-guided navigation, OCR/text & face recognition with scene descriptions, emergency alerts, and adjustable 1-10 times audio for safer independent mobility.
BUIDL AI Hackathon
1st Prize (SAGA Track)Developed full-stack Web3 application enabling autonomous, non human-intervention agent interactions within EVM token ecosystems, featured by Metamask integration.
YAICON, 5th
Grand Prize (1st Place)Established a pipeline capable of performing 3D reconstruction of a moving scene using only two shots, followed by editing the reconstructed scene.
The 3rd Entrepreneurship Competition for People with and without Disabilities, 2024
Grand Prize (1st Place, Awarded by the Minister of Education, Republic of Korea.)Developed online demo service of SUNNY Braille with AI integration, containing custom-pretrained math problem detection model (based on ViT)
D-Tech Competition, 2023
Grand Prize (1st Place, Awarded by the Minister of Health and Welfare, Republic of Korea)Developed baseline full stack webpage to serve online math braille service SUNNY Braille.
SK AI Summit AI's Got Talent
1st PrizeAwarded with the enhanced version of A-EYE.
SW-Centered University Digital Competition
Grand Prize (2nd Place, Award of the Chief of IITP)Conducted a poster session on 2025 SK AI Summit!
Developed A-EYE, an AI-powered real-time environmental recognition system for the visually impaired, enabling 3-second hazard detection, GPS-guided navigation, OCR/text & face recognition with scene descriptions, emergency alerts, and adjustable 1-10 times audio for safer independent mobility.
BUIDL AI Hackathon
1st Prize (SAGA Track)Developed full-stack Web3 application enabling autonomous, non human-intervention agent interactions within EVM token ecosystems, featured by Metamask integration.
15th DB Insurance & Finance Contest
Selection PrizeImplemented dual-model function-call system (tool design, tool-call pipeline) into insurance design / analysis pipeline.
YAICON, 5th
Grand Prize (1st Place)Established a pipeline capable of performing 3D reconstruction of a moving scene using only two shots, followed by editing the reconstructed scene.
YAICON, 4th
Siver Prize (2nd Place)Constructed personalized music recommendation system based on MU-LLaMA embeddings.
14th DB Financial Economics Contest, 2024
Honorable MentionDeveloped time series prediction model using LSTM, integrating with prediction of solar power energy generation capacity of certain area.
The 3rd Entrepreneurship Competition for People with and without Disabilities, 2024
Grand Prize (1st Place, Awarded by the Minister of Education, Republic of Korea.)Developed online demo service of SUNNY Braille with AI integration, containing custom-pretrained math problem detection model (based on ViT)
Yonsei AI Entrepreneurship Competition, 2023
3rd prizeIssued by Yonsei University Computing College. Awarded with the enhanced version of SUNNY Braille.
International Capstone Design Fair, 2023
2nd prizeAwarded with the enhanced version of SUNNY Braille.
D-Tech Competition, 2023
Grand Prize (1st Place, Awarded by the Minister of Health and Welfare, Republic of Korea)Developed baseline full stack webpage to serve online math braille service SUNNY Braille.
D-Tech Competition, 2023
EDUTECH PrizeIssued by Korea Edutech Industry Association (KETIA)
ELYPECS Young Leaders
Top-team Prize (Harmony Prize)MLB HRDX Korea
Top 10 SupportersFull resume in PDF (TBU).
MS Student
ALIN Lab (KAIST) — Advised by Prof. Jinwoo Shin
Undergraduate Researcher
ALIN Lab (KAIST) — Advised by Prof. Jinwoo Shin
Independent Graduation Project Student
MMAI Lab (Yonsei Univ.) — Advised by Prof. Jongyoo Kim
Research Collaborator
AIM Intelligence — Discovering robustness of guardrail models.
Engineering Lead
HOWEVER — General SW & HW developer of photo booth system.
Language Model Engineer
GOODGANG Labs — YAI x GoodGang Labs Collaboration Project
Research Intern
MIRLAB (Yonsei Univ.) — Advised by Prof. Youngjae Yu
Member, 13th Cohort
YAI (Yonsei Artificial Intelligence) — Yonsei University AI group
Mandatory Military Service
Military Service — 8th Maneuver Division, Republic of Korea Army
B.S. Student
Yonsei University — Nano Science & Engineering
Electric & Electronic Engineering (Double Major)