On August 11, Prof. Hongliang Ren’s research group from The Chinese University of Hong Kong (CUHK) was invited by the Department of Magnetic Surgery at the First Affiliated Hospital of Xi’an Jiaotong University (XJTU) to visit their specialized magnetic surgery ward and engage in a symposium with Prof. Xiaopeng Yan to explore cross-disciplinary medical-engineering research collaboration.
During the visiting, Prof. Yan presented the development history, landmark surgical cases, and distinct advantages of magnetic surgery compared to conventional surgical techniques. In particular, magnetic recanalization offers a highly effective, minimally invasive solution for severe complications like post-liver transplantation biliary obstructionโoften termed the “Achilles’ heel” of liver transplantationโwhere traditional endoscopic or surgical interventions face major limitations. Prof. Yan also highlighted current clinical pain points and shared valuable insights regarding urgent operational needs.
Prof. Hongliang Ren introduced the group’s latest research achievements in advanced magnetic materials and magnetically controlled surgical robots. Both parties conducted fruitful discussions on the translational application of magnetic technologies in clinical settings, identified key prospective areas for joint research, and established plans for ongoing collaboration.
Special thanks to Prof. Xiaopeng Yan and the Department of Magnetic Surgery at XJTU First Affiliated Hospital for hosting this inspiring exchange!
The award recognizes the work โActively Controlled Continuous-Everting Capsule Robot: Design, Fabrication, and Validation.โ
This research introduces a magnetically controlled capsule robot that uses a continuous-everting outer film to reduce sliding friction against delicate tissue, enabling smoother and safer navigation in complex gastrointestinal environments.
Following the conceptualization of this work, a prototype was manufactured and tested, and the functionality of the everting capsule robot was initially verified.
Figure 1. Max Meng and Botao Lin at the award ceremony
Figure 2. Working concept diagram of the continuous-everting capsule robot. Driven by an external magnetic field, the robot can move by continuously everting its outer film. Because this type of movement is unaffected by environmental friction, it can be applied to lumens of any suitable diameter.
Figure 3. Design and working principle of the continuous-everting capsule robot. (a) Design parameters of the continuous-everting capsule robot. (b) The forward motion cycle of the robot. Three colored markers (blue, red, and yellow) are utilized to track the movement of the outer film. During the traction eversion phase, magnetic attraction drives the inner core forward, which induces the eversion of the outer film via friction between the two components. In the inner core resetting phase, the external magnetic field rotates and translates, causing the inner core to simultaneously rotate and move backward. During this resetting phase, the outer film remains static.
Figure 4. Experimental validations. (a) The capsule robot prototype and the diagram of the actuation control board. (b) Experimental evaluation of the capsule robotโs eversion movement and the inner coreโs resetting capability. During the eversion experiment, external magnetic traction pulled the robot forward approximately 4cm in 4 seconds. In the resetting experiment, the inner core retracted through the robotโs internal cavity in 9 seconds. Throughout this resetting process, the outer film remained entirely stationary. (c) The record of the positions of the robot and the inner core during the experiments. It can be measured that the average velocity of the robot moving and the inner core resetting are 7.5mm/s and 4.4mm/s, respectively.
We are excited to share our latest Comment published in #npj Digital Medicine: โHow can reasoning capability empower the AI copilot robot in endoscopic surgeryโ
Current AI copilots in endoscopic surgery are still largely reactive and vision-driven. While they can detect anatomy, instruments, and scene changes, they often still struggle to truly understand surgical intent, infer hidden tissue dynamics, and respond robustly to uncertaintyโall of which are essential for safe and precise intraoperative assistance.
In this article, we highlight how reasoning capability can become a key enabler for the next generation of AI copilot robots in endoscopic surgery:
๐ง Reasoning beyond perception: Enabling VLA-based surgical robots to go beyond simple visual recognition and translate high-level surgeon intent into precise, context-aware low-level motion goals.
๐ Multimodal and uncertainty-aware intelligence: Fusing endoscopic vision with preoperative imaging, intraoperative sensing, and tracking signals, while dynamically re-weighting information sources under occlusion, bleeding, smoke, and other uncertain conditions.
๐ค Coordinated multi-instrument collaboration: Supporting synchronized control of multiple tools for subtasks such as traction, dissection, and hemostasis, with reasoning-guided adaptation to tissue deformation and workflow variation.
๐ฎ Anticipatory and safer decision-making: Using chain-of-thought-style reasoning to forecast tissue response, evaluate possible action outcomes, and generate more conservative and interpretable assistance under risk.
๐ฅ Surgeon-in-the-loop clinical deployment: Framing the AI copilot robot at LoA 2โ3, where the system assists with task generation, monitoring, and bounded low-level execution under explicit safety constraints and continuous surgeon oversight.
We believe the future of endoscopic robotic assistance lies not only in systems that can see, but in systems that can reason, adapt, and collaborate. With reasoning-enabled VLA models, AI copilot robots may evolve from reactive executors into true cognitive partners in the operating room.
๐ Read the full paper here: https://www.nature.com/articles/s41746-026-02827-8
๐ Kudos to the team: Mr. Guankun Wang, Dr. Long Bai, and Prof. Hongliang Ren.
Thrilled to share our latest Advanced Science work on enabling highly transferable, autonomous navigation for wireless capsule endoscopy (WCE)โusing a lightweight Edge-Contour-Depth Fusion module and deep reinforcement learning (DRL).
WCE has revolutionized GI diagnostics, but its potential is often restricted by incomplete mucosal coverage and the poor ability of existing AI navigation methods to adapt across different patient anatomies. This motivated us to ditch the heavy, brittle, traditional “end-to-end” visual video streams that cause AI models to overfit to a single patient.
๐ง โจ What we developed: A unified, clinically viable framework that features: ๐น Anatomical Landmark Guidance: Operates on stable, low-dimensional coordinates of conserved gastric structures (the fundus and pyloric antrum) rather than high-dimensional raw video. ๐น Lightweight Perception Module: Combines classical Canny edge detection and Hu moments with a compact monocular depth network (DispNet) to run efficiently on low-power clinical hardware. ๐น Robust Sim-to-Real Pipeline: Utilizes a patient-specific digital twin combined with a model-free Adaptive Dynamic Programming (ADP) controller to actively neutralize real-world physical disturbances and actuator latency.
๐ฏ Key Results: โ >97% mucosal coverage achieved within 50 seconds across 8 diverse, patient-derived stomach models in simulation. โ 87% mean coverage stability and a 53% reduction in procedure time during real-world ex-vivo experiments compared to expert manual control. โ Drastically reduced computational overhead, allowing deployment on low-cost processors (<2 TOPS).
๐ก Why it matters: This study establishes a scalable paradigm that conquers the “reality gap” and patient anatomical variability in medical robotics. By decoupling perception from control, it removes the need for expensive, massive patient datasets and high-end GPUs, paving the way for operator-independent, intelligent GI diagnostics.
๐ฑ Whatโs next? We are expanding our training to encompass extreme pathological distortions (like hiatal hernias) and advancing toward fully wireless clinical deployment with dynamic, target-reaching capabilities for intraoperative pathologies.
Thrilled to share our newly accepted paper in IEEE Transactions on Medical Robotics and Bionics, where we introduce a soft, endoscope-deployable microfluidic suction robot that combines multimodal intraluminal locomotion with localized aspiration and sampling for targeted mucus clearance and liquid biopsy.
๐ง โจ What we developed:
A soft intraluminal robotic platform that:
๐น Integrates Locomotion + Sampling: A pneumatically controlled 2ร2 pouch matrix for multimodal actuation, paired with an independent microfluidic suction module for active liquid extraction and sample recovery.
๐น Enables Stable Pitch Control: A balloon-based pitch control mechanism improves controllability for intraluminal operation, with the best overall performance observed at an initial pressure range of 2โ3 kPa.
๐น Balances Compliance and Safety: Single-pouch characterization guided the selection of a 2 mm pouch radius, achieving up to 246.91% maximum deformation; burst tests show a system safety factor โ 5.27 under the reported operating conditions.
๐น Targets Real Clinical Pain Points: Designed for constrained lumens (e.g., distal airway) where conventional airway clearance approaches struggle with reach and effectiveness.
๐ฏ Key Results:
โ Multimodal mobility: Differential actuation achieves 26.9 mm/min forward speed and 4.86ยฐ yaw per drive cycle.
โ Robust suction across viscosities: Efficiently extracts 20โ80% glycerol solutions within 10 s (via parameter tuning).
โ In vivo feasibility: Endoscope-assisted porcine validation confirmed sequential pouch-driven motion and successful recovery of biological samples containing mucus and tissue fragments after saline irrigation.
๐ก Why it matters:
This work demonstrates a compliant, integrated โmove + anchor + suctionโ approach for narrow lumensโsupporting safer localized intervention and sampling, with a path toward distal airway translation.
๐ฑ Whatโs next?
Weโre moving toward more automated closed-loop pneumatic control, improved steerability/navigation, and miniaturization for deeper airway accessโwhile expanding validation in airway-specific models.
๐ค๐ชข Thrilled to share our latest work published on Nature Communication, which redefines actuation for tendon-driven continuum robots โ achieving full 3D omnidirectional motion and body twist using only a single tendon.
๐ง โจ What we developed: A new class of continuum robots that: ๐น Breaks Design Constraints: Eliminates the inherent trade-off between miniaturization and 3D manipulability by replacing multiple tendons with a single eccentric one. ๐น Push-Pull-Twist Actuation: Achieves complex spatial movement through a unique driving mechanism. ๐น High Efficiency: Features an outer diameter of 2.0โ3.5 mm with a hollow ratio exceeding 57% โ doubling the spatial utilization of traditional designs. ๐น Open-Source Support: Includes a derived kinematics model and an open-source simulator for the robotics community.
๐ฏ Key Results:
โ >1,000-fold Improvement: Massive increase in manipulability compared to conventional multi-tendon mechanisms.
โ High Force Retention: Retains at least 70% of tip force across all directions.
โ Versatile Demonstration: Proven success in teleoperation, navigation through tortuous environments, and “chopstick-like” continuum grippers.
๐ก Why it matters: This work proves that miniature robots can maintain high dexterity and power without the bulk of traditional hardware, pointing toward the next generation of surgical actuators.
๐ฑ Whatโs next? We are exploring potential medical applications and the integration of these actuators into complex surgical procedures.
Thrilled to share our latest international collaboration! At NVIDIA GTC 2026 in San Jose, CA, the team led by Professor Hongliang Ren from The Chinese University of Hong Kong (CUHK), in partnership with NVIDIA and 35 leading global institutions, officially released Open-H-Embodiment, the worldโs first and largest open-source dataset for medical robotics, now available on HuggingFace.
During the GTC keynote, Kimberly Powell, NVIDIAโs VP of Healthcare, highlighted this milestone. Our lab is honored to be a primary contributor, filling the critical gap in Embodied AI for medical robotics by providing high-fidelity data for contact dynamics and closed-loop control.
๐ง โจ What we contributed & developed:
This project breaks the “perception-heavy, execution-light” limitation of traditional medical AI. Key highlights include:
๐น 778 Hours of Massive Multimodal Data: The dataset covers 400 complete clinical surgeries and 9 major robotic platforms (e.g., dVRK, CMR Versius, Kuka). It includes 65% clinical data, 23% bench-top experiments, and 12% simulation data.
๐น Three High-Value Specialized Datasets from Our Lab:
Dual-Source Ultrasound Dataset:ย Experts-level trajectories covering in-vivo porcine EUS and human forearm scanning, overcoming complex organ environments and multi-device calibration.
Robotic Surgery Skill Dataset:ย Multi-modal data (RGB/RGB-D + Kinematics) for tissue manipulation and suturing, featuring millisecond-level synchronization and dual-mode control (teleoperation & automation).
Flexible Endoscope Tracking Baseline:ย A standardized dataset addressing hysteresis and deformation in flexible endoscopy, supporting nanosecond-level time synchronization.
๐น Surgical VLA & World Models:
GR00T-H:ย A 3B-parameter Vision-Language-Action model based on NVIDIA Isaac GR00T, capable of long-horizon dexterous tasks like end-to-end suturing.
Cosmos-H-Surgical-Simulator:ย An action-conditioned world model that boosts simulation efficiency by over 70x, bridging the sim-to-real gap.
๐ฏ Key Results: โ Global Standardization: First effort to unify medical robotic data across different devices and institutions under CC-BY-4.0. โ Efficiency Boost: Accelerated surgical simulation (600 sims in 40 mins) to generate high-fidelity video-action pairs. โ Clinical Relevance: Successfully captured nearly 500 hours of real-world clinical data for hernia, gallbladder, and uterine surgeries.
๐ก Why it matters: This initiative provides the foundational “bedrock” for Medical Physical AI. By sharing high-quality, synchronized data for surgery, ultrasound, and endoscopy, we are lowering the barrier for researchers worldwide to develop autonomous surgical agents that are both explainable and adaptive.
๐ฑ Whatโs next? Our lab is continuing to deepen research in: ๐น Reasoning-based autonomous control for surgical robots. ๐น Cross-platform generalization of Medical VLA models. ๐น Clinical translation of Embodied AI to improve patient outcomes.
Thrilled to share our latest #๐ฆ๐ฐ๐ถ๐ฒ๐ป๐ฐ๐ฒ๐๐ฑ๐๐ฎ๐ป๐ฐ๐ฒ work on a ๐ฐ๐ฒ๐ป๐๐ถ๐บ๐ฒ๐๐ฒ๐ฟ-๐๐ฐ๐ฎ๐น๐ฒ, ๐ณ๐๐น๐น๐ ๐๐ฒ๐ฎ๐น๐ฒ๐ฑ, ๐น๐ฒ๐ด๐น๐ฒ๐๐ ๐ฎ๐บ๐ฝ๐ต๐ถ๐ฏ๐ถ๐ผ๐๐ ๐ฟ๐ผ๐ฏ๐ผ๐ that can crawl on sand, jump, and steerably swim – powered by a ๐๐ถ๐ป๐ด๐น๐ฒ ๐๐ฎ๐ฟ๐ถ๐ฎ๐ฏ๐น๐ฒ-๐ผ๐๐๐ฝ๐๐ ๐๐ผ๐ถ๐ฐ๐ฒ-๐ฐ๐ผ๐ถ๐น ๐บ๐ผ๐๐ผ๐ฟ (๐ฉ๐๐ ).
At small scales, reliable ๐ธ๐ข๐ต๐ฆ๐ณ๐ฑ๐ณ๐ฐ๐ฐ๐ง ๐ด๐ฆ๐ข๐ญ๐ช๐ฏ๐จ is tough: transmissions and active mechanisms mean moving parts and dynamic seals that are fragile and ๐ญ๐ฆ๐ข๐ฌ-๐ฑ๐ณ๐ฐ๐ฏ๐ฆ. We wanted an amphibious robot that stays sealed and robust – yet still supports multiple locomotion modes.
An ๐ถ๐ป๐ฒ๐ฟ๐๐ถ๐ฎ-๐ฑ๐ฟ๐ถ๐๐ฒ๐ป ๐ฎ๐ฐ๐๐๐ฎ๐๐ถ๐ผ๐ป + ๐ฝ๐ฎ๐๐๐ถ๐๐ฒ ๐ฝ๐ฟ๐ผ๐ฝ๐๐น๐๐ถ๐ผ๐ป ๐ฑ๐ฒ๐๐ถ๐ด๐ป that:
๐น Uses a variable-output VCM inside a fully sealed rigid shell (no external moving parts)
๐น Switches among three modes: jumping, full-stroke vibration (land), and small-stroke vibration (water)
๐น Uses asymmetric, tilted passive fins for frequency-tuned steering in water (IDMP)
๐น Explains hydrodynamics via aquatic tests, high-speed PIV, and CFD
All of this – ๐ผ๐ป๐ฒ ๐บ๐ฎ๐ถ๐ป ๐น๐ถ๐ป๐ฒ๐ฎ๐ฟ ๐ฎ๐ฐ๐๐๐ฎ๐๐ผ๐ฟ + ๐ฝ๐ฎ๐๐๐ถ๐๐ฒ ๐ณ๐ถ๐ป๐. No exposed legs, gears, or propellers – making sealing and durability much easier at the centimeter scale.
๐ฏ ๐๐ฒ๐ ๐ฅ๐ฒ๐๐๐น๐๐:
โ 24-g prototype (57.5 ร 36 ร 36 mm) with a fully enclosed shell
โ ~1.4 BL/s (~78 mm/s) on dry sand; 41.6 mm/s on flat ground
โ Jump height up to 17.16 mm at 15 V; continuous โtumblerโ jumping
โ Load carrying: 960 g (~40ร body weight) and escape under 5-kg loads
โ In water: min turning radius 5.6 mm; straight swim at 35/44/60 Hz (~32/35/28 mm/s); peak yaw -22ยฐ/s (30 Hz) or 16ยฐ/s (40 Hz)
This work shows how inertia + mode-switchable actuation can bridge the ๐บ๐ผ๐บ๐ฒ๐ป๐๐๐บ-๐ณ๐ฟ๐ฒ๐พ๐๐ฒ๐ป๐ฐ๐ ๐๐ฟ๐ฎ๐ฑ๐ฒ-๐ผ๐ณ๐ณ, enabling jumping and swimming in the same tiny robot. Passive asymmetric fins turn simple reciprocation into steerable thrust – ๐๐ถ๐๐ต ๐ป๐ผ ๐ฎ๐ฑ๐ฑ๐ฒ๐ฑ ๐ฎ๐ฐ๐๐๐ฎ๐๐ผ๐ฟ๐.
๐ฑ ๐ช๐ต๐ฎ๐โ๐ ๐ป๐ฒ๐ ๐?
Future work will focus on improving environmental adaptability via ๐ง๐ฆ๐ฆ๐ฅ๐ฃ๐ข๐ค๐ฌ ๐ค๐ฐ๐ฏ๐ต๐ณ๐ฐ๐ญ and adaptive structures, and optimizing energy efficiency + onboard power ๐ฎ๐ช๐ฏ๐ช๐ข๐ต๐ถ๐ณ๐ช๐ป๐ข๐ต๐ช๐ฐ๐ฏ.
Special shoutout to the team –
Lingqi Tang, Yongzun Yang (co-first authors), Bing Li, Bingfu Zhang, Qiguang He, Hongliang Ren, Yao Li – for making this project possible.
๐ฅ Accurate visualization of subtle vascular dynamics remains a significant challenge in minimally invasive surgery, where dynamic complexities often limit decision-making reliability. Our paper introduces ๐๐ป๐ฑ๐ผ๐๐ผ๐ป๐๐ฟ๐ผ๐น๐ ๐ฎ๐ด, a framework designed to ๐ฒ๐ป๐ต๐ฎ๐ป๐ฐ๐ฒ ๐๐ฎ๐๐ฐ๐๐น๐ฎ๐ฟ ๐บ๐ผ๐๐ถ๐ผ๐ป ๐๐ถ๐๐ถ๐ฏ๐ถ๐น๐ถ๐๐ in endoscopic videos while preserving surrounding tissue structure. The approach integrates ๐ฃ๐ฒ๐ฟ๐ถ๐ผ๐ฑ๐ถ๐ฐ ๐ฅ๐ฒ๐ณ๐ฒ๐ฟ๐ฒ๐ป๐ฐ๐ฒ ๐ฅ๐ฒ๐๐ฒ๐๐๐ถ๐ป๐ด to minimize error accumulation over time and ๐๐ถ๐ฒ๐ฟ๐ฎ๐ฟ๐ฐ๐ต๐ถ๐ฐ๐ฎ๐น ๐ง๐ถ๐๐๐๐ฒ-๐ฎ๐๐ฎ๐ฟ๐ฒ ๐ ๐ฎ๐ด๐ป๐ถ๐ณ๐ถ๐ฐ๐ฎ๐๐ถ๐ผ๐ป for adaptive vessel tracking.
๐ To validate robustness, we constructed ๐๐ป๐ฑ๐ผ๐ฉ๐ ๐ ๐ฎ๐ฐ, a benchmark dataset spanning four surgical specialties and diverse intraoperative scenarios. Quantitative metrics and expert surgeon evaluations indicate improved magnification accuracy and image quality compared to existing methods.
๐ค We extend our sincere gratitude to our collaborators across The Chinese University of Hong Kong (An Wang, Mengya Xu, Yiting Chang, Prof Hongliang Ren), The University of Hong Kong (Rulin Zhou), Southern Medical University (่้พ้ฃ, Prof Hao Chen), The First Affiliated Hospital of Wenzhou Medical University (Yiru Ye), Southern University of Science and
Technology (Prof Jiankun Wang), and Singapore General Hospital (Prof Chwee Ming Lim) for their invaluable contributions to this multidisciplinary work.
The paper is available at https://lnkd.in/ggPEFswF
Thrilled to share our latest work, ๐๐๐จ๐๐๐ง๐, a unified geometry-aware framework for language-guided robotic grasping.
Language-guided grasping is a key capability for intuitive humanโrobot interaction. A robot should not only detect objects but also understand natural instructions such as โpick up the blue cup behind the bowl.โ While recent multimodal models have shown promising results, most existing approaches rely on multi-stage pipelines that loosely couple perception and grasp prediction. These methods often overlook the tight integration of geometry, language, and visual reasoning, making them fragile in cluttered, occluded, or low-texture environments. This motivated us to bridge the gap between semantic language understanding and precise geometric grasp execution.
A novel unified framework for geometry-aware language-guided grasping that includes:
๐น Unified RGB-D Multimodal Representation:
We embed RGB, depth, and language features into a shared representation space, enabling consistent cross-modal semantic alignment for accurate target reasoning.
๐น Depth-Guided Geometric Module (DGGM):
Instead of treating depth as auxiliary input, we explicitly inject geometric priors derived from depth into the attention mechanism, strengthening object discrimination under occlusion and ambiguous visual conditions.
๐น Adaptive Dense Channel Integration (ADCI):
A dynamic multi-layer fusion strategy that balances global semantic cues and fine-grained geometric details for robust grasp prediction.
๐ฏ ๐๐๐ฒ ๐๐๐ฌ๐ฎ๐ฅ๐ญ๐ฌ:
โ GeoLanG significantly outperforms prior multi-stage baselines on OCID-VLG for language-guided grasping.
โ Demonstrates strong robustness in cluttered and heavily occluded scenes.
โ Successfully validated on real robotic hardware, showing reliable sim-to-real transfer.
This work shows that tightly coupling geometric reasoning with multimodal language understanding can significantly enhance robotic grasp reliability. By embedding depth-aware geometric priors directly into attention mechanisms, we reduce ambiguity and improve consistency in grasp decision-making.
GeoLanG provides a pathway toward more intelligent robotic systems that understand not just what object to grasp, but also how to grasp it robustly in complex real-world environments.
๐ฑ ๐๐ก๐๐ญโ๐ฌ ๐ง๐๐ฑ๐ญ?
We are exploring extending this geometry-aware multimodal reasoning toward:
๐น Real-time interactive grasping
๐น Multi-step manipulation tasks
๐น Integration with motion planning and autonomous robotic control