Junhyeong Park

I work on vision-language-action models, multimodal generation, and learning from action-free data.

KAIST · ALIN LabMS Student

Junhyeong Park
01

About

I'm an incoming MS student at ALIN Lab (KAIST), starting Fall 2026, advised by Prof. Jinwoo Shin.

I believe the fusion of modalities—centered around language models—is essential for AI to positively impact society. I'm especially interested in VLA models that can understand and learn from action-label-free datasets, and in the effective representation learning of generative models.

02

News

03

Research & Publications

Most recent work on Google Scholar. indicates equal contribution.

RLDX-1 Technical Report

RLDX-1 Technical Report

RLDX Team (Core contributor)

RLWRLD Technical Report 2026

While Vision-Language-Action models (VLAs) have shown remarkable progress toward human-like generalist robotic policies through the versatile intelligence (broad scene understanding and language-conditioned generalization) inherited from pre-trained Vision-Language Models, they still struggle with complex real-world tasks requiring broader functional capabilities such as motion awareness, long-term memory, and physical sensing. To address this, we introduce RLDX-1, a general-purpose robotic policy for dexterous manipulation built on the Multi-Stream Action Transformer (MSAT), an architecture that unifies these capabilities by integrating heterogeneous modalities through modality-specific streams with cross-modal joint self-attention. RLDX-1 further combines this architecture with system-level design choices, including data synthesis for rare manipulation scenarios, learning procedures specialized for human-like manipulation, and inference optimizations for real-time deployment. Through empirical evaluation, we show that RLDX-1 consistently outperforms recent frontier VLAs across both simulation benchmarks and real-world tasks that require broad functional capabilities beyond general versatility. In particular, RLDX-1 shows superiority in ALLEX humanoid tasks by achieving success rates of 86.8%, highlighting its ability to control a high-DoF humanoid robot under diverse functional demands. Together, these results position RLDX-1 as a promising step toward reliable VLAs for complex, contact-rich, and dynamic real-world dexterous manipulation.

Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?

Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?

Jiwan Chung, Janghan Yoon, Junhyeong Park, Sangeyl Lee, Joowon Yang, Sooyeon Park, Youngjae Yu

ACL 2025 Poster

Any-to-any generative models aim to enable seamless interpretation and generation across multiple modalities within a unified framework, yet their ability to preserve relationships across modalities remains uncertain. Do unified models truly achieve cross-modal coherence, or is this coherence merely perceived? To explore this, we introduce ACON, a dataset of 1,000 images (500 newly contributed) paired with captions, editing instructions, and Q&A pairs to evaluate cross-modal transfers rigorously. Using three consistency criteria—cyclic consistency, forward equivariance, and conjugated equivariance—our experiments reveal that any-to-any models do not consistently demonstrate greater cross-modal consistency than specialized models in pointwise evaluations such as cyclic consistency. However, equivariance evaluations uncover weak but observable consistency through structured analyses of the intermediate latent space enabled by multiple editing operations.

VLM Safety Evaluation Must Move from Classification to Causal Controllability

VLM Safety Evaluation Must Move from Classification to Causal Controllability

Junhyeong Park, Hanwool Lee, DongGeon Lee, Dasol Choi, Yejin Son, Haon Park, Youngjae Yu

Coming soon...

InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?

InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?

Keummin Ka, Junhyeong Park, Jaehyun Jeon, Youngjae Yu

KCC 2026 Oral

Recent advances in Vision-Language Models (VLMs) have demonstrated impressive capabilities in perception and reasoning. However, the ability to perform causal inference—a core aspect of human cognition—remains underexplored, particularly in multimodal settings. In this study, we introduce InfoCausalQA, a novel benchmark designed to evaluate causal reasoning grounded in infographics that combine structured visual data with textual context. The benchmark comprises two tasks: quantitative causal reasoning based on inferred numerical trends, and semantic causal reasoning involving five types of causal relations (cause, effect, intervention, counterfactual, and temporal). We manually collected 494 infographic-text pairs from four public sources and used GPT-4o to generate 1,482 high-quality multiple-choice QA pairs, then carefully revised by humans so they cannot be answered from surface-level cues alone but instead require genuine visual grounding. Our results reveal that current VLMs exhibit limited capability in computational reasoning and even more pronounced limitations in semantic causal reasoning, with performance significantly lower than humans. Through InfoCausalQA, we highlight the need for advancing the causal reasoning abilities of multimodal AI systems.

04

Experience

Research Collaborator

AIM Intelligence

Jul. 2025 - Jan. 2026

Engineering Lead

HOWEVER

Dec. 2024 - Apr. 2025

Collaborative Language Model Engineer

GOODGANG Labs

Dec. 2024 - Feb. 2025

Member, 13th Cohort

YAI (Yonsei Artificial Intelligence)

Yonsei University AI group

Dec. 2023 - Jun. 2026
05

Awards

SK AI Summit AI's Got Talent

1st Prize

Awarded with the enhanced version of A-EYE.

Nov. 2025

SW-Centered University Digital Competition

Grand Prize (2nd Place, Award of the Chief of IITP)

Conducted a poster session on 2025 SK AI Summit!
Developed A-EYE, an AI-powered real-time environmental recognition system for the visually impaired, enabling 3-second hazard detection, GPS-guided navigation, OCR/text & face recognition with scene descriptions, emergency alerts, and adjustable 1-10 times audio for safer independent mobility.

Aug. 2025

BUIDL AI Hackathon

1st Prize (SAGA Track)

Developed full-stack Web3 application enabling autonomous, non human-intervention agent interactions within EVM token ecosystems, featured by Metamask integration.

Apr. 2025

YAICON, 5th

Grand Prize (1st Place)

Established a pipeline capable of performing 3D reconstruction of a moving scene using only two shots, followed by editing the reconstructed scene.

Nov. 2024

The 3rd Entrepreneurship Competition for People with and without Disabilities, 2024

Grand Prize (1st Place, Awarded by the Minister of Education, Republic of Korea.)

Developed online demo service of SUNNY Braille with AI integration, containing custom-pretrained math problem detection model (based on ViT)

Feb. 2024

D-Tech Competition, 2023

Grand Prize (1st Place, Awarded by the Minister of Health and Welfare, Republic of Korea)

Developed baseline full stack webpage to serve online math braille service SUNNY Braille.

Sep. 2023

SK AI Summit AI's Got Talent

1st Prize

Awarded with the enhanced version of A-EYE.

Nov. 2025

SW-Centered University Digital Competition

Grand Prize (2nd Place, Award of the Chief of IITP)

Conducted a poster session on 2025 SK AI Summit!
Developed A-EYE, an AI-powered real-time environmental recognition system for the visually impaired, enabling 3-second hazard detection, GPS-guided navigation, OCR/text & face recognition with scene descriptions, emergency alerts, and adjustable 1-10 times audio for safer independent mobility.

Aug. 2025

BUIDL AI Hackathon

1st Prize (SAGA Track)

Developed full-stack Web3 application enabling autonomous, non human-intervention agent interactions within EVM token ecosystems, featured by Metamask integration.

Apr. 2025

15th DB Insurance & Finance Contest

Selection Prize

Implemented dual-model function-call system (tool design, tool-call pipeline) into insurance design / analysis pipeline.

Mar. 2025

YAICON, 5th

Grand Prize (1st Place)

Established a pipeline capable of performing 3D reconstruction of a moving scene using only two shots, followed by editing the reconstructed scene.

Nov. 2024

YAICON, 4th

Siver Prize (2nd Place)

Constructed personalized music recommendation system based on MU-LLaMA embeddings.

Aug. 2024

14th DB Financial Economics Contest, 2024

Honorable Mention

Developed time series prediction model using LSTM, integrating with prediction of solar power energy generation capacity of certain area.

Apr. 2024

The 3rd Entrepreneurship Competition for People with and without Disabilities, 2024

Grand Prize (1st Place, Awarded by the Minister of Education, Republic of Korea.)

Developed online demo service of SUNNY Braille with AI integration, containing custom-pretrained math problem detection model (based on ViT)

Feb. 2024

Yonsei AI Entrepreneurship Competition, 2023

3rd prize

Issued by Yonsei University Computing College. Awarded with the enhanced version of SUNNY Braille.

Dec. 2023

International Capstone Design Fair, 2023

2nd prize

Awarded with the enhanced version of SUNNY Braille.

Oct. 2023

D-Tech Competition, 2023

Grand Prize (1st Place, Awarded by the Minister of Health and Welfare, Republic of Korea)

Developed baseline full stack webpage to serve online math braille service SUNNY Braille.

Sep. 2023

D-Tech Competition, 2023

EDUTECH Prize

Issued by Korea Edutech Industry Association (KETIA)

Sep. 2023

ELYPECS Young Leaders

Top-team Prize (Harmony Prize)

Jan. 2023

MLB HRDX Korea

Top 10 Supporters

Sep. 2022
06

Recent Activity

Selected posts from LinkedIn ↗. Scroll sideways to browse.

07

Vitæ

Full resume in PDF (TBU).

MS Student

ALIN Lab (KAIST) — Advised by Prof. Jinwoo Shin

Aug. 2026

Undergraduate Researcher

ALIN Lab (KAIST) — Advised by Prof. Jinwoo Shin

Sep. 2025 - Aug. 2026

Independent Graduation Project Student

MMAI Lab (Yonsei Univ.) — Advised by Prof. Jongyoo Kim

Jul. 2025 - Nov. 2025

Research Collaborator

AIM Intelligence — Discovering robustness of guardrail models.

Jul. 2025 - Jan. 2026

Engineering Lead

HOWEVER — General SW & HW developer of photo booth system.

Dec. 2024 - Apr. 2025

Language Model Engineer

GOODGANG Labs — YAI x GoodGang Labs Collaboration Project

Dec. 2024 - Feb. 2025

Research Intern

MIRLAB (Yonsei Univ.) — Advised by Prof. Youngjae Yu

Sep. 2024 - Oct. 2025

Member, 13th Cohort

YAI (Yonsei Artificial Intelligence) — Yonsei University AI group

Dec. 2023 - Jun. 2026

Mandatory Military Service

Military Service — 8th Maneuver Division, Republic of Korea Army

Mar. 2021 - Sep. 2022

B.S. Student

Yonsei University — Nano Science & Engineering
Electric & Electronic Engineering (Double Major)

2020 - 2026
08

Stats

Flag Counter