Bowei Li

ROBOT LEARNING · REAL-WORLD SYSTEMS

Bowei Li  李博伟

M.S. Student, Department of Electrical and Computer Engineering
Carnegie Mellon University

Robot LearningReal2Sim2RealLong-Horizon ManipulationAgentic Policies

Email Scholar GitHub LinkedIn CV

Scroll to explore

About

I am a Master's student at the Department of Electrical and Computer Engineering, Carnegie Mellon University, advised by Prof. Changliu Liu in the Intelligent Control Lab. My research spans robot learning, scalable robot data generation, long-horizon manipulation, and agentic robot policies.

I received my Bachelor's degree in Telecommunication Engineering jointly from Xidian University and Heriot-Watt University. During my undergraduate studies, I worked at the Digital Governance Engineering Research Center led by Prof. Huailiang Liu on sentiment analysis, and was also remotely advised by Prof. Ran Zhang from UNC Charlotte on reinforcement learning for UAV communications.

Research

Publications & Preprints

* denotes equal contribution. Author names in bold indicate me.

Under review · ICRA

ARSTAG: An Agentic Real2Sim2Real System for Task-Specific Robot Data Generation

Bowei Li, Yuner Zhang, Changliu Liu

arXiv preprint, 2026 · Submitted to ICRA

Adapting visuomotor policies to new manipulation tasks often requires substantial manual engineering or teleoperated data collection. Simulation can provide task-specific data at scale, but constructing the scene, designing expert behavior, and configuring data generation still require significant per-task effort. We present ARSTAG, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data. A hierarchy of language agents constructs a task-scoped simulation scene, generates robot-feasible demonstrations, and expands the training distribution through task-consistent randomization, while a coordinator agent manages cross-stage feedback and recovery. Across seven manipulation tasks spanning grasping, placement, and stacking, the ARSTAG-generated demonstrations enable sim-to-real transfer of three visuomotor policy architectures to a dual-arm robot, with π0.5 achieving an average real-world success rate of 74.6%. Ablations show that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size.
@misc{li2026arstag,
  title         = {ARSTAG: An Agentic Real2Sim2Real System for Task-Specific Robot Data Generation},
  author        = {Li, Bowei and Zhang, Yuner and Liu, Changliu},
  year          = {2026},
  eprint        = {2609.24563},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2609.24563}
}

Under review · RA-L

BrickCraft: Visuomotor Skill Composition with Situated Manual Guidance for Long-Horizon Interlocking Brick Assembly

Jichuan Yu*, Bowei Li*, Zhenran Tang, Guanxing Lu, Chuxiong Hu, Ruixuan Liu, Changliu Liu

arXiv preprint, 2026 · Submitted to IEEE Robotics and Automation Letters (RA-L)

Autonomous robotic assembly of interlocking bricks demands seamless integration of long-horizon task reasoning, spatial grounding, and fine-grained manipulation. This paper presents BrickCraft, a compositional framework designed for long-horizon and generalizable interlocking brick assembly. BrickCraft models the assembly process using a relative formulation, where each step is anchored to a reference brick within the partial structure, thereby decomposing complex tasks into a finite set of reusable primitive skills. BrickCraft bridges the gap between high-level assembly plans and physical execution through situated manuals, which provide explicit spatial guidance for learned visuomotor skills by projecting the assembly intent onto real-time robot observations. Finally, BrickCraft employs a compositional execution pipeline that chains these spatially grounded skills to accomplish long-horizon assembly tasks. Extensive experimental validations demonstrate that BrickCraft acquires proficient assembly skills from a limited set of demonstrations and exhibits strong compositional generalization to unseen structures.
@article{yu2026brickcraft,
  title   = {BrickCraft: Visuomotor Skill Composition with Situated Manual Guidance for Long-Horizon Interlocking Brick Assembly},
  author  = {Yu, Jichuan and Li, Bowei and Tang, Zhenran and Lu, Guanxing and Hu, Chuxiong and Liu, Ruixuan and Liu, Changliu},
  journal = {arXiv preprint arXiv:2605.07605},
  year    = {2026}
}

Under review · ICRA

ManiSkillFormer: Demonstration-Free Compositional Manipulation via Geometric Contracts and Agentic Skill Graph

Peiqi Yu, Mosam Dabhi, Shangtao Li, Bowei Li, Laszlo Jeni, Changliu Liu

arXiv preprint, 2026 · Submitted to ICRA

ManiSkillFormer connects perception and reusable manipulation skills through geometric contracts. LLM agents generate task-specific geometric requirements and motion templates from human-defined skill structures. A perception module grounds the required geometry in the current scene, enabling skill execution and composition without additional robot demonstrations or per-object policy fine-tuning.
@misc{yu2026maniskillformer,
  title         = {ManiSkillFormer: Demonstration-Free Compositional Manipulation via Geometric Contracts and Agentic Skill Graph},
  author        = {Yu, Peiqi and Dabhi, Mosam and Li, Shangtao and Li, Bowei and Jeni, Laszlo and Liu, Changliu},
  year          = {2026},
  eprint        = {2609.16331},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2609.16331}
}
MD-COAS constraint optimization and adaptive diffusion scheduling

IROS 2026

Motion Planning with Model-Based Diffusion via Constraint Optimization and Adaptive Scheduling

Zhilin He, Bowei Li, Jianlin Dou, Yuner Zhang, Changliu Liu

IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2026

MD-COAS integrates soft constraint handling through an inexact Augmented Lagrangian Method with hard projection using Convex Feasible Sets. It adapts constraint enforcement throughout model-based diffusion. Evaluations on non-convex 2D planning benchmarks and a 7-DoF arm avoidance task examine safety, success, convergence, and final trajectory cost.
@inproceedings{he2026mdcoas,
  title     = {Motion Planning with Model-Based Diffusion via Constraint Optimization and Adaptive Scheduling},
  author    = {He, Zhilin and Li, Bowei and Dou, Jianlin and Zhang, Yuner and Liu, Changliu},
  booktitle = {IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
  year      = {2026},
  eprint    = {2607.14455},
  url       = {https://arxiv.org/abs/2607.14455}
}
Robust Pruning

Under review · ACC

Enhancing Certifiable Semantic Robustness via Robust Pruning of Deep Neural Networks

Hanjiang Hu*, Bowei Li*, Ziwei Wang*, Tianhao Wei, Casidhe Hutchison, Eric Sample, Changliu Liu

arXiv preprint, 2025 · Submitted to the American Control Conference (ACC)

Deep neural networks have been widely adopted in many vision and robotics applications with visual inputs. It is essential to verify their robustness against semantic transformation perturbations, such as brightness and contrast. However, current certified training and robustness certification methods face the challenge of over-parameterization, which hinders the tightness and scalability due to over-complicated neural networks. To this end, we first analyze stability and variance of layers and neurons against input perturbation, showing that certifiable robustness can be indicated by a fundamental Unbiased and Smooth Neuron metric (USN). Based on USN, we introduce a novel pruning method that removes neurons with low USN and retains those with high USN, preserving model expressiveness without over-parameterization. To further enhance this pruning, we propose a new Wasserstein distance loss to ensure pruned neurons are more concentrated across layers. We validate our approach on robust keypoint detection under realistic brightness and contrast perturbations, demonstrating superior robustness certification performance and efficiency compared to baselines.
@article{hu2025enhancing,
  title   = {Enhancing Certifiable Semantic Robustness via Robust Pruning of Deep Neural Networks},
  author  = {Hu, Hanjiang and Li, Bowei and Wang, Ziwei and Wei, Tianhao and Hutchison, Casidhe and Sample, Eric and Liu, Changliu},
  journal = {arXiv preprint arXiv:2510.00083},
  year    = {2025}
}

WBCD 1st Prize · ICRA 2025RSS 2025 Workshop

NeSyPack: A Neuro-Symbolic Framework for Bimanual Logistics Packing

Bowei Li*, Peiqi Yu*, Zhenran Tang, Han Zhou, Yifan Sun, Ruixuan Liu, Changliu Liu

RSS 2025 Workshop on Benchmarking Robot Manipulation · First Prize, WBCD Competition @ ICRA 2025

This paper presents NeSyPack, a neuro-symbolic framework for bimanual logistics packing. NeSyPack combines data-driven models and symbolic reasoning to build an explainable hierarchical system that is generalizable, data-efficient, and reliable. It decomposes a task into subtasks via hierarchical reasoning, and further into atomic skills managed by a symbolic skill graph. The graph selects skill parameters, robot configurations, and task-specific control strategies for execution. This modular design enables robustness, adaptability, and efficient reuse — outperforming end-to-end models that require large-scale retraining. Using NeSyPack, our team won the First Prize in the What Bimanuals Can Do (WBCD) competition at the 2025 IEEE International Conference on Robotics and Automation.
@article{li2025nesypack,
  title   = {NeSyPack: A Neuro-Symbolic Framework for Bimanual Logistics Packing},
  author  = {Li, Bowei and Yu, Peiqi and Tang, Zhenran and Zhou, Han and Sun, Yifan and Liu, Ruixuan and Liu, Changliu},
  journal = {arXiv preprint arXiv:2506.06567},
  year    = {2025}
}

IFAC 2025

SPARK: Safe Protective and Assistive Robot Kit

Yifan Sun, Rui Chen, Kai S. Yun, Yikuan Fang, Sebin Jung, Feihan Li, Bowei Li, Weiye Zhao, Changliu Liu

IFAC Symposium on Robotics, 2025

This paper introduces the Safe Protective and Assistive Robot Kit (SPARK), a comprehensive benchmark designed to ensure safety in humanoid autonomy and teleoperation. Humanoid robots pose significant safety risks due to their physical capabilities of interacting with complex environments, and their structures further add complexity to general safety solutions. SPARK is a toolbox that comes with state-of-the-art safe control algorithms in a modular and composable robot control framework. Users can configure safety criteria and sensitivity levels to optimize the balance between safety and performance. SPARK provides simulation benchmarks comparing safety approaches in a variety of environments, tasks, and robot models, and allows quick deployment of synthesized safe controllers on real robots. For hardware deployment, SPARK supports Apple Vision Pro or a Motion Capture System as external sensors, with interfaces for alternative hardware. We demonstrate SPARK's capability with simulation experiments and case studies on a Unitree G1 humanoid.
@inproceedings{sun2025spark,
  title     = {SPARK: Safe Protective and Assistive Robot Kit},
  author    = {Sun, Yifan and Chen, Rui and Yun, Kai S. and Fang, Yikuan and Jung, Sebin and Li, Feihan and Li, Bowei and Zhao, Weiye and Liu, Changliu},
  booktitle = {IFAC Symposium on Robotics},
  year      = {2025}
}
WLMD

IEEE TCCN

When Learning Meets Dynamics: Distributed User Connectivity Maximization in AAV-Based Communication Networks

Bowei Li, Saugat Tripathi, AKM Salman Hosain, Ran Zhang, Miao Wang, Jiang (Linda) Xie

IEEE Transactions on Cognitive Communications and Networking, 12(1):175–188, 2026

Distributed management over UAV-based communication networks (UCNs) has attracted increasing research attention. In this work, we study a distributed user connectivity maximization problem in a UCN, featuring a horizontal study over different levels of information exchange and a consideration of dynamics in UAV set and user distribution. The problem is formulated as a time-coupled mixed-integer non-convex optimization. A heuristic two-stage UAV–user association policy is proposed to determine connectivity faster. To tackle the NP-hard problem at scale, DUCM-1 is proposed under the multi-agent deep Q learning (MA-DQL) framework, evaluating how information-exchange levels impact convergence; DUCM-2 then handles arbitrary join-ins and quits of UAVs in a considered horizon.
@article{li2025when,
  title   = {When Learning Meets Dynamics: Distributed User Connectivity Maximization in AAV-Based Communication Networks},
  author  = {Li, Bowei and Tripathi, Saugat and Hosain, AKM Salman and Zhang, Ran and Wang, Miao and Xie, Jiang},
  journal = {IEEE Transactions on Cognitive Communications and Networking},
  volume  = {12},
  number  = {1},
  pages   = {175--188},
  year    = {2026}
}
LWD

IEEE ComMag

Learning with Dynamics: Autonomous Regulation of UAV-Based Communication Networks with Dynamic UAV Crew

Ran Zhang, Bowei Li, Liyuan Zhang, Jiang (Linda) Xie, Miao Wang

IEEE Communications Magazine, 64(1):72–78, 2026

UAV-based communication networks (UCNs) are a key component in future mobile networking. To handle their dynamic environments, reinforcement learning (RL) has emerged as a promising solution thanks to its strong adaptive decision-making free of environment models. However, most existing RL-based research focuses on control strategies assuming a fixed UAV set, while few works investigate adaptive regulation when the serving UAVs change dynamically. This article discusses RL-based strategy design for adaptive UCN regulation given a dynamic UAV set, addressing both reactive strategies in general UCNs and proactive strategies in solar-powered UCNs, with case studies from our recent works.
@article{zhang2025learning,
  title   = {Learning with Dynamics: Autonomous Regulation of UAV-Based Communication Networks with Dynamic UAV Crew},
  author  = {Zhang, Ran and Li, Bowei and Zhang, Liyuan and Xie, Jiang and Wang, Miao},
  journal = {IEEE Communications Magazine},
  volume  = {64},
  number  = {1},
  pages   = {72--78},
  year    = {2026}
}
USER

ICC 2025

Maximizing User Connectivity in AI-Enabled Multi-UAV Networks: A Distributed Strategy Generalized to Arbitrary User Distributions

Bowei Li, Yang Xu, Ran Zhang, Jiang (Linda) Xie, Miao Wang

IEEE International Conference on Communications (ICC), pp. 530–535, 2025

Deep reinforcement learning (DRL) has been extensively applied to Multi-UAV networks (MUNs) for real-time adaptation in time-varying environments. However, most existing works assume a stationary or predictably dynamic user distribution (UD), which makes UD-specific strategies insufficient when a MUN is deployed in unknown environments. This paper investigates distributed user connectivity maximization in a MUN with generalization to arbitrary UDs. The problem is formulated as a time-coupled combinatorial nonlinear non-convex optimization. A multi-agent CNN-enhanced deep Q learning (MA-CDQL) algorithm is proposed, integrating a ResNet-based CNN that analyzes the input UD in real time. A heatmap algorithm transforms the raw UD into a continuous density map to improve learning efficiency.
@inproceedings{li2025maximizing,
  title     = {Maximizing User Connectivity in AI-Enabled Multi-UAV Networks: A Distributed Strategy Generalized to Arbitrary User Distributions},
  author    = {Li, Bowei and Xu, Yang and Zhang, Ran and Xie, Jiang and Wang, Miao},
  booktitle = {IEEE International Conference on Communications (ICC)},
  pages     = {530--535},
  year      = {2025}
}

Project

PerSEVE

Perception by Segmentation, Extensible Visual Ecosystem

ARM Institute AIDF · Lockheed Martin ATL · Carnegie Mellon University · GE Aerospace · Sep 2025 – Jul 2026

PerSEVE builds manufacturing-specific data and foundation perception models for industrial parts. The project curated high-quality CAD models, generated large-scale labeled synthetic scenes in NVIDIA Isaac Sim, fine-tuned SAM 2.1 for manufacturing segmentation (SAM4MFG), and used geometry-driven hard-example mining to target the model's own failure cases.

8,918curated CAD models
100Ksynthetic scenes
~550Kobject instances
6annotation types

My contributions. Led an image-based Real2Sim pipeline for object reconstruction, editable scene graphs, and relation-constrained Isaac Sim synthesis; developed physics-based synthetic-data generation with domain randomization and dense annotations.

Scene Intelligence · Real2Sim Demo
Photo → segmentation → scene graph → Isaac Sim scene · 2× speed · 1:00
Extension to Real Robots
Generated simulation scenes paired with real-robot execution · 1:14 · see ARSTAG

Competition

WBCD Challenge @ ICRA 2025 · First Prize

What Bimanual Teleoperation and Learning from Demonstration Can Do

ICRA 2025 · Atlanta · Logistics packing track · Team from Carnegie Mellon University

Our CMU team won first place in the bimanual logistics-packing challenge at ICRA 2025, using a neuro-symbolic framework that decomposes packing into reusable skills with failure recovery. The system was later published as NeSyPack at the RSS 2025 Workshop on Benchmarking Robot Manipulation.

My contributions. Developed the neuro-symbolic bimanual task-execution pipeline combining skill graphs, object grounding, 6D pose estimation, and failure recovery; deployed on Unitree G1 and Galaxea R1 Lite.

Competition Highlight
Dual-arm logistics packing at ICRA 2025 · 0:30
Champion Solution
NeSyPack walkthrough · 2:09

Featured by Galaxea Dynamics

  • Galaxea Dynamics — Official partner of the challenge, on top universities tackling dual-arm logistics packing.
    View on X · May 2025
  • Galaxea Dynamics — “The WBCD Champion Team from Carnegie Mellon University has unveiled their solution.”
    View on X · June 2025

Project

ETC · Future Football

Human–Robot Interaction for Future Sports

Carnegie Mellon University · Entertainment Technology Center (ETC)

A collaboration exploring how people and humanoid robots can play football together, through interactive prototypes and real-robot demonstrations on Unitree G1.

My contributions. Trained and deployed reinforcement-learning control policies on Unitree G1, and built ROS/Python interfaces connecting the robot stack to a Unity front end for interactive control and hardware prototyping.

Future Football
Human–robot interaction showcase · 0:53
Robot Football Prototype
Unitree G1 lab demonstration · 0:18

Featured on LinkedIn

ETC project excerpt · Powering the Future of Sport · Video: Carnegie Mellon University
  • Carnegie Mellon University — Powering the Future of Sport: A Draft Week Showcase.
    Watch on LinkedIn · Excerpt from 0:37–0:40 in the full video
  • Derek Ham, ETC Director — Exploring human–robot football and the future of entertainment.
    Watch on LinkedIn

Project

MASCEI · Automated PPE Production

Rapid-response robotic manufacturing in shipping containers

ARM Institute · NIST-funded · Oct 2024 – Sep 2025

An ARM Institute project with AFFOA, Siemens, Yaskawa, Carnegie Mellon University and industry partners, building automated face-mask production lines housed in shipping containers. The lines combine robotic sewing, automated visual inspection, picking and sorting, and end-of-line packing and palletizing, so they can go from set-up to production within days rather than months.

My contributions. Integrated hand–eye, tool, and environment calibration across four industrial robot cells, supporting a unified manufacturing calibration workflow.

MASCEI system demonstration, June 2025 (excerpt) · Video: ARM Institute