Yixuan Yang

Yixuan Yang

Ph.D. Student in Electrical & Computer Engineering
Duke University

Yixuan Yang standing on a ridge in Death Valley, holding a camera

About Me

Hi there from Yixuan! I am a Ph.D. student in the Department of Electrical and Computer Engineering at Duke University, advised by Professor Rishi Kamaleswaran in the Kamaleswaran Lab. My research focuses on world models and self-supervised representation learning. I study how complex dynamical systems evolve under intervention, and how to build representations of those dynamics that transfer across tasks and domains. My current setting is clinical foundation models. Prior to this, I worked in the General Robotics Lab during my first year under the supervision of Professor Boyuan Chen, exploring Contact-Rich Robotic Manipulation and Embodied AI. My long-term goal is to build foundation models that understand how the world changes: systems that learn predictive world models and make robust decisions in complex, unstructured environments 🧠🤖.

I received my Bachelor's degree in Computer Science from Southern University of Science and Technology, where I was proud to be supervised by Chair Professor Xin Yao. In the Fall 2022 semester, I studied at the University of California, Berkeley as an exchange student. Back in China, I joined Professor Xinlei Chen's research group at Tsinghua University as a visiting student.

Research interests World Models & Latent Dynamics · Self-Supervised Representation Learning · Sequential Decision-Making under Uncertainty · Embodied & Clinical Foundation Models

News

  • [Aug. 2024]
    I officially started my Ph.D. journey in the General Robotics Lab at Duke University! 🚀
  • [May 2024]
    🎉 I was awarded the Guo Xie Bi Rong Fellowship (¥10,000), Class of 2024. (only 4%)
  • [May 2024]
    🎉 I was awarded as an Outstanding Undergraduate Graduate (only 25%), Class of 2024 (SUSTech), as well as a Distinguished Graduate (Top 2 out of 245 students) in the Department of Computer Science and Engineering.
  • [Mar. 2024]
    🎉 The paper "Learning-Based Problem Reduction for Large-Scale Uncapacitated Facility Location Problems" got accepted to the conference CEC 2024 held in Yokohama, Japan.
  • [Dec. 2023]
    🎉 The paper "Poster: Olfactory Sensing in Turbulent Airflow via Collaborative Robots" got accepted to the conference HotMobile '24 held in San Diego, United States.
  • [May 2023]
    I became a visiting student in Professor Xinlei Chen's lab. I worked with several Ph.D. and MS students in the project Gas Source Localization driven by Collaborative Robots.
  • [Jan. 2023]
    I joined Chair Professor Xin Yao (Fellow, IEEE)'s Lab, and started doing research in large-scale UFLP Problem.
  • [Aug. 2022]
    🛫 I arrived at the United States to be an exchange student at UC Berkeley 😎. Hi California!!! (Cal's weather is soooooo fantastic!!! ☀️🥹)
  • [May 2022]
    I attended the Summer Workshop 2022 in the School of Computing at National University of Singapore (NUS). We used Unity to develop a networked 2D Game.

Research Experiences

You can also check my Google Scholar profile.

  • arXiv Clin-JEPA framework diagram: a shared encoder and a latent trajectory predictor over EHR patient state and action text.
    Yixuan Yang, Mehak Arora, Ryan Zhang, Baraa Abed, Junseob Kim, Tilendra Choudhary, Md Hassanuzzaman, Kevin Zhu, Ayman Ali, Chengkun Yang, Alasdair Edward Gent, Victor Moas, Rishikesan Kamaleswaran
    Preprint: arXiv:2605.10840 (May 2026) · under review

    Every JEPA before this one either discards the predictor or trains it on a frozen encoder, so the encoder never learns to support the rollout it will be asked for. Clin-JEPA co-trains them, and the representation spreads out: 55% more effective dimensions, variance no longer piled into a handful of them. Freeze that same encoder, train a predictor the ordinary way, and long-horizon accuracy falls below doing nothing.

  • The Unitree G1's hand closing on a knob mounted to a panel in Isaac Gym. Many copies of the same reaching task training in parallel in Isaac Gym, the goal marked by a yellow wireframe sphere. The physical learning kit, a plywood box carrying knobs, switches, buttons and a readout screen, on a bench in the General Robotics Lab.
    Contact-Rich Humanoid Manipulation for Scientific Laboratory Automation
    Yixuan Yang. Advised by Professor Boyuan Chen
    General Robotics Lab, Duke University (Aug. 2024 – Mar. 2025)

    Laboratory instruments are built for human hands: knobs, switches, small buttons. This project put a Unitree G1 humanoid in front of one, pairing a physical testbed with a dimension-matched replica in Isaac Gym. Most of the work was the task stack itself: an inverse-kinematics control interface, fingertip contact sensing, and the reward and termination design that gets PPO to learn contact-rich motion rather than flail near it.

  • ACM TOSN SniffySquad system figure: multiple ground robots cooperating to localise a gas source in a patchy plume.
    Yuhan Cheng*, Xuecheng Chen*, Yixuan Yang, Haoyang Wang, Jingao Xu, Chaopeng Hong, Susu Xu, Xiao-Ping Zhang, Yunhao Liu, Xinlei Chen
    ACM Transactions on Sensor Networks (Feb 2026) (DOI: 10.1145/3786599)

    Turbulence tears a gas plume into patches, and every patch can look like the source. So SniffySquad stops climbing the gradient and starts sampling: each robot is a Langevin chain on a shared belief map, and its temperature is its role. Hot robots roam, cold robots inspect, and the two trade places by the same acceptance rule parallel tempering uses. Success rate up 20%, path efficiency up 30%.

  • HotMobile
    Yuhan Cheng*, Xuecheng Chen*, Yixuan Yang, Haoyang Wang, Yuxuan Liu, Xinlei Chen
    HotMobile 2024: Proceedings of the 25th International Workshop on Mobile Computing Systems and Applications

    Turbulence scatters a gas plume into patches, so following the concentration gradient walks a robot into a decoy. This early poster gives a robot team heterogeneous roles: some inspect the strongest reading found so far, others keep hunting for new candidates, and the roles are exchanged as the readings change. Search time fell 37% against Surge-Cast and 22% against Infotaxis, which gets trapped in exactly those patches.

  • IEEE CEC Six-stage pipeline: a large facility-location instance, small instances used as training data, cost and rank matrices compressed to four statistics per facility, a network predicting each facility's open probability, the pruned instance, and a faster-converging search on it.
    Shuaixiang Zhang, Yixuan Yang, Hao Tong, Xin Yao
    CEC 2024: IEEE Congress on Evolutionary Computation

    Facility location at scale defeats search heuristics, so this shrinks the problem before solving it. The trick is to look at ranks rather than costs: each facility is described by the shape of its rank distribution across customers, so the same four numbers describe it whether the instance holds fifty sites or five thousand. That scale-free footing is why a model trained on instances a solver can crack exactly transfers to ones it cannot, cutting benchmarks to 13% of their size.

  • Figure titled Learning to Say No, showing five kinds of unsolvable task -- status conflict, item absence, logical contradiction, ambiguity and ethical constraint -- feeding a synthetic data pipeline and a vision-language model that answers
    Final project, Duke ECE 590 / ME 555: Robot Learning (Fall 2024)

    Robots asked to do something impossible tend to try anyway. This builds a synthetic dataset of five ways a task can be unsolvable, from a missing item to an ethical refusal, then fine-tunes a vision-language model to spot them and say why. Refusal accuracy goes from 10% to 78% on the generated images, and 81% in Habitat-Sim, a renderer it never trained on.