๐ค๐ชข Thrilled to share our latest work published on Nature Communication, which redefines actuation for tendon-driven continuum robots โ achieving full 3D omnidirectional motion and body twist using only a single tendon.
๐ง โจ What we developed: A new class of continuum robots that: ๐น Breaks Design Constraints: Eliminates the inherent trade-off between miniaturization and 3D manipulability by replacing multiple tendons with a single eccentric one. ๐น Push-Pull-Twist Actuation: Achieves complex spatial movement through a unique driving mechanism. ๐น High Efficiency: Features an outer diameter of 2.0โ3.5 mm with a hollow ratio exceeding 57% โ doubling the spatial utilization of traditional designs. ๐น Open-Source Support: Includes a derived kinematics model and an open-source simulator for the robotics community.
๐ฏ Key Results:
โ >1,000-fold Improvement: Massive increase in manipulability compared to conventional multi-tendon mechanisms.
โ High Force Retention: Retains at least 70% of tip force across all directions.
โ Versatile Demonstration: Proven success in teleoperation, navigation through tortuous environments, and “chopstick-like” continuum grippers.
๐ก Why it matters: This work proves that miniature robots can maintain high dexterity and power without the bulk of traditional hardware, pointing toward the next generation of surgical actuators.
๐ฑ Whatโs next? We are exploring potential medical applications and the integration of these actuators into complex surgical procedures.
On the afternoon of April 24, 2026, Professor Hongliang Renโs research team from The Chinese University of Hong Kong visited The Third Affiliated Hospital of Sun Yat-sen University (SYSU Third Hospital) to hold an academic symposium and salon on Embodied Medical Robotics. The event was held in Conference Room 2006 of the Comprehensive Building, aiming to advance academic communication and potential research collaboration in medical robotics and intelligent healthcare.
The meeting was chaired by Dr. Liu Zifeng, Director of the Big Data and Artificial Intelligence Center of SYSU Third Hospital. Vice President Qintai Yang delivered a welcome speech, introducing the hospitalโs clinical strengths and looking forward to close cooperation in medical robot research and translation.
Professor Hongliang Ren delivered a keynote report on minimally invasive flexible robotic systems and embodied intelligence in medicine. Members of his research team then presented recent advances in precision interventional tools, magnetic robotic systems, endoscopic intelligent navigation, autonomous surgical control, and medical image perception, showcasing the teamโs innovative work in clinical-oriented medical robotics.
During the discussion session, clinicians and researchers from multiple departments of SYSU Third Hospital had in-depth exchanges on clinical demands, technical applications, and joint research plans. Both sides reached positive consensus on future cooperation in scientific research, talent development, and clinical translation.
At the end of the meeting, Vice President Qintai Yang delivered a concluding speech and spoke highly of this academic exchange. This symposium effectively bridged engineering research and clinical practice, and laid a solid foundation for the future development and application of embodied medical robotics.
About The Third Affiliated Hospital of Sun Yat-sen University
The Third Affiliated Hospital of Sun Yat-sen University was founded in 1971 and is a comprehensive Grade-A tertiary hospital directly administered by the National Health Commission of China. As a major clinical teaching base of Sun Yat-sen University, the hospital undertakes medical care, education, research, prevention, rehabilitation, and specialist training. It currently operates four campuses: Tianhe, Lingnan, Yuedong, and Zhaoqing. The hospital has developed distinctive strengths in liver disease, brain disorders, and immune-related diseases, supported by national key disciplines and national clinical key specialty programs. It is also recognized as a Guangdong High-level Hospital and serves as an output hospital for the development of national regional medical centers.
Thrilled to share our latest international collaboration! At NVIDIA GTC 2026 in San Jose, CA, the team led by Professor Hongliang Ren from The Chinese University of Hong Kong (CUHK), in partnership with NVIDIA and 35 leading global institutions, officially released Open-H-Embodiment, the worldโs first and largest open-source dataset for medical robotics, now available on HuggingFace.
During the GTC keynote, Kimberly Powell, NVIDIAโs VP of Healthcare, highlighted this milestone. Our lab is honored to be a primary contributor, filling the critical gap in Embodied AI for medical robotics by providing high-fidelity data for contact dynamics and closed-loop control.
๐ง โจ What we contributed & developed:
This project breaks the “perception-heavy, execution-light” limitation of traditional medical AI. Key highlights include:
๐น 778 Hours of Massive Multimodal Data: The dataset covers 400 complete clinical surgeries and 9 major robotic platforms (e.g., dVRK, CMR Versius, Kuka). It includes 65% clinical data, 23% bench-top experiments, and 12% simulation data.
๐น Three High-Value Specialized Datasets from Our Lab:
Dual-Source Ultrasound Dataset:ย Experts-level trajectories covering in-vivo porcine EUS and human forearm scanning, overcoming complex organ environments and multi-device calibration.
Robotic Surgery Skill Dataset:ย Multi-modal data (RGB/RGB-D + Kinematics) for tissue manipulation and suturing, featuring millisecond-level synchronization and dual-mode control (teleoperation & automation).
Flexible Endoscope Tracking Baseline:ย A standardized dataset addressing hysteresis and deformation in flexible endoscopy, supporting nanosecond-level time synchronization.
๐น Surgical VLA & World Models:
GR00T-H:ย A 3B-parameter Vision-Language-Action model based on NVIDIA Isaac GR00T, capable of long-horizon dexterous tasks like end-to-end suturing.
Cosmos-H-Surgical-Simulator:ย An action-conditioned world model that boosts simulation efficiency by over 70x, bridging the sim-to-real gap.
๐ฏ Key Results: โ Global Standardization: First effort to unify medical robotic data across different devices and institutions under CC-BY-4.0. โ Efficiency Boost: Accelerated surgical simulation (600 sims in 40 mins) to generate high-fidelity video-action pairs. โ Clinical Relevance: Successfully captured nearly 500 hours of real-world clinical data for hernia, gallbladder, and uterine surgeries.
๐ก Why it matters: This initiative provides the foundational “bedrock” for Medical Physical AI. By sharing high-quality, synchronized data for surgery, ultrasound, and endoscopy, we are lowering the barrier for researchers worldwide to develop autonomous surgical agents that are both explainable and adaptive.
๐ฑ Whatโs next? Our lab is continuing to deepen research in: ๐น Reasoning-based autonomous control for surgical robots. ๐น Cross-platform generalization of Medical VLA models. ๐น Clinical translation of Embodied AI to improve patient outcomes.
Thrilled to share our latest #๐ฆ๐ฐ๐ถ๐ฒ๐ป๐ฐ๐ฒ๐๐ฑ๐๐ฎ๐ป๐ฐ๐ฒ work on a ๐ฐ๐ฒ๐ป๐๐ถ๐บ๐ฒ๐๐ฒ๐ฟ-๐๐ฐ๐ฎ๐น๐ฒ, ๐ณ๐๐น๐น๐ ๐๐ฒ๐ฎ๐น๐ฒ๐ฑ, ๐น๐ฒ๐ด๐น๐ฒ๐๐ ๐ฎ๐บ๐ฝ๐ต๐ถ๐ฏ๐ถ๐ผ๐๐ ๐ฟ๐ผ๐ฏ๐ผ๐ that can crawl on sand, jump, and steerably swim – powered by a ๐๐ถ๐ป๐ด๐น๐ฒ ๐๐ฎ๐ฟ๐ถ๐ฎ๐ฏ๐น๐ฒ-๐ผ๐๐๐ฝ๐๐ ๐๐ผ๐ถ๐ฐ๐ฒ-๐ฐ๐ผ๐ถ๐น ๐บ๐ผ๐๐ผ๐ฟ (๐ฉ๐๐ ).
At small scales, reliable ๐ธ๐ข๐ต๐ฆ๐ณ๐ฑ๐ณ๐ฐ๐ฐ๐ง ๐ด๐ฆ๐ข๐ญ๐ช๐ฏ๐จ is tough: transmissions and active mechanisms mean moving parts and dynamic seals that are fragile and ๐ญ๐ฆ๐ข๐ฌ-๐ฑ๐ณ๐ฐ๐ฏ๐ฆ. We wanted an amphibious robot that stays sealed and robust – yet still supports multiple locomotion modes.
An ๐ถ๐ป๐ฒ๐ฟ๐๐ถ๐ฎ-๐ฑ๐ฟ๐ถ๐๐ฒ๐ป ๐ฎ๐ฐ๐๐๐ฎ๐๐ถ๐ผ๐ป + ๐ฝ๐ฎ๐๐๐ถ๐๐ฒ ๐ฝ๐ฟ๐ผ๐ฝ๐๐น๐๐ถ๐ผ๐ป ๐ฑ๐ฒ๐๐ถ๐ด๐ป that:
๐น Uses a variable-output VCM inside a fully sealed rigid shell (no external moving parts)
๐น Switches among three modes: jumping, full-stroke vibration (land), and small-stroke vibration (water)
๐น Uses asymmetric, tilted passive fins for frequency-tuned steering in water (IDMP)
๐น Explains hydrodynamics via aquatic tests, high-speed PIV, and CFD
All of this – ๐ผ๐ป๐ฒ ๐บ๐ฎ๐ถ๐ป ๐น๐ถ๐ป๐ฒ๐ฎ๐ฟ ๐ฎ๐ฐ๐๐๐ฎ๐๐ผ๐ฟ + ๐ฝ๐ฎ๐๐๐ถ๐๐ฒ ๐ณ๐ถ๐ป๐. No exposed legs, gears, or propellers – making sealing and durability much easier at the centimeter scale.
๐ฏ ๐๐ฒ๐ ๐ฅ๐ฒ๐๐๐น๐๐:
โ 24-g prototype (57.5 ร 36 ร 36 mm) with a fully enclosed shell
โ ~1.4 BL/s (~78 mm/s) on dry sand; 41.6 mm/s on flat ground
โ Jump height up to 17.16 mm at 15 V; continuous โtumblerโ jumping
โ Load carrying: 960 g (~40ร body weight) and escape under 5-kg loads
โ In water: min turning radius 5.6 mm; straight swim at 35/44/60 Hz (~32/35/28 mm/s); peak yaw -22ยฐ/s (30 Hz) or 16ยฐ/s (40 Hz)
This work shows how inertia + mode-switchable actuation can bridge the ๐บ๐ผ๐บ๐ฒ๐ป๐๐๐บ-๐ณ๐ฟ๐ฒ๐พ๐๐ฒ๐ป๐ฐ๐ ๐๐ฟ๐ฎ๐ฑ๐ฒ-๐ผ๐ณ๐ณ, enabling jumping and swimming in the same tiny robot. Passive asymmetric fins turn simple reciprocation into steerable thrust – ๐๐ถ๐๐ต ๐ป๐ผ ๐ฎ๐ฑ๐ฑ๐ฒ๐ฑ ๐ฎ๐ฐ๐๐๐ฎ๐๐ผ๐ฟ๐.
๐ฑ ๐ช๐ต๐ฎ๐โ๐ ๐ป๐ฒ๐ ๐?
Future work will focus on improving environmental adaptability via ๐ง๐ฆ๐ฆ๐ฅ๐ฃ๐ข๐ค๐ฌ ๐ค๐ฐ๐ฏ๐ต๐ณ๐ฐ๐ญ and adaptive structures, and optimizing energy efficiency + onboard power ๐ฎ๐ช๐ฏ๐ช๐ข๐ต๐ถ๐ณ๐ช๐ป๐ข๐ต๐ช๐ฐ๐ฏ.
Special shoutout to the team –
Lingqi Tang, Yongzun Yang (co-first authors), Bing Li, Bingfu Zhang, Qiguang He, Hongliang Ren, Yao Li – for making this project possible.
๐ฅ Accurate visualization of subtle vascular dynamics remains a significant challenge in minimally invasive surgery, where dynamic complexities often limit decision-making reliability. Our paper introduces ๐๐ป๐ฑ๐ผ๐๐ผ๐ป๐๐ฟ๐ผ๐น๐ ๐ฎ๐ด, a framework designed to ๐ฒ๐ป๐ต๐ฎ๐ป๐ฐ๐ฒ ๐๐ฎ๐๐ฐ๐๐น๐ฎ๐ฟ ๐บ๐ผ๐๐ถ๐ผ๐ป ๐๐ถ๐๐ถ๐ฏ๐ถ๐น๐ถ๐๐ in endoscopic videos while preserving surrounding tissue structure. The approach integrates ๐ฃ๐ฒ๐ฟ๐ถ๐ผ๐ฑ๐ถ๐ฐ ๐ฅ๐ฒ๐ณ๐ฒ๐ฟ๐ฒ๐ป๐ฐ๐ฒ ๐ฅ๐ฒ๐๐ฒ๐๐๐ถ๐ป๐ด to minimize error accumulation over time and ๐๐ถ๐ฒ๐ฟ๐ฎ๐ฟ๐ฐ๐ต๐ถ๐ฐ๐ฎ๐น ๐ง๐ถ๐๐๐๐ฒ-๐ฎ๐๐ฎ๐ฟ๐ฒ ๐ ๐ฎ๐ด๐ป๐ถ๐ณ๐ถ๐ฐ๐ฎ๐๐ถ๐ผ๐ป for adaptive vessel tracking.
๐ To validate robustness, we constructed ๐๐ป๐ฑ๐ผ๐ฉ๐ ๐ ๐ฎ๐ฐ, a benchmark dataset spanning four surgical specialties and diverse intraoperative scenarios. Quantitative metrics and expert surgeon evaluations indicate improved magnification accuracy and image quality compared to existing methods.
๐ค We extend our sincere gratitude to our collaborators across The Chinese University of Hong Kong (An Wang, Mengya Xu, Yiting Chang, Prof Hongliang Ren), The University of Hong Kong (Rulin Zhou), Southern Medical University (่้พ้ฃ, Prof Hao Chen), The First Affiliated Hospital of Wenzhou Medical University (Yiru Ye), Southern University of Science and
Technology (Prof Jiankun Wang), and Singapore General Hospital (Prof Chwee Ming Lim) for their invaluable contributions to this multidisciplinary work.
The paper is available at https://lnkd.in/ggPEFswF
Thrilled to share our latest work, ๐๐๐จ๐๐๐ง๐, a unified geometry-aware framework for language-guided robotic grasping.
Language-guided grasping is a key capability for intuitive humanโrobot interaction. A robot should not only detect objects but also understand natural instructions such as โpick up the blue cup behind the bowl.โ While recent multimodal models have shown promising results, most existing approaches rely on multi-stage pipelines that loosely couple perception and grasp prediction. These methods often overlook the tight integration of geometry, language, and visual reasoning, making them fragile in cluttered, occluded, or low-texture environments. This motivated us to bridge the gap between semantic language understanding and precise geometric grasp execution.
A novel unified framework for geometry-aware language-guided grasping that includes:
๐น Unified RGB-D Multimodal Representation:
We embed RGB, depth, and language features into a shared representation space, enabling consistent cross-modal semantic alignment for accurate target reasoning.
๐น Depth-Guided Geometric Module (DGGM):
Instead of treating depth as auxiliary input, we explicitly inject geometric priors derived from depth into the attention mechanism, strengthening object discrimination under occlusion and ambiguous visual conditions.
๐น Adaptive Dense Channel Integration (ADCI):
A dynamic multi-layer fusion strategy that balances global semantic cues and fine-grained geometric details for robust grasp prediction.
๐ฏ ๐๐๐ฒ ๐๐๐ฌ๐ฎ๐ฅ๐ญ๐ฌ:
โ GeoLanG significantly outperforms prior multi-stage baselines on OCID-VLG for language-guided grasping.
โ Demonstrates strong robustness in cluttered and heavily occluded scenes.
โ Successfully validated on real robotic hardware, showing reliable sim-to-real transfer.
This work shows that tightly coupling geometric reasoning with multimodal language understanding can significantly enhance robotic grasp reliability. By embedding depth-aware geometric priors directly into attention mechanisms, we reduce ambiguity and improve consistency in grasp decision-making.
GeoLanG provides a pathway toward more intelligent robotic systems that understand not just what object to grasp, but also how to grasp it robustly in complex real-world environments.
๐ฑ ๐๐ก๐๐ญโ๐ฌ ๐ง๐๐ฑ๐ญ?
We are exploring extending this geometry-aware multimodal reasoning toward:
๐น Real-time interactive grasping
๐น Multi-step manipulation tasks
๐น Integration with motion planning and autonomous robotic control
Thrilled to share our latest work on enabling robust sparse-to-dense reconstruction for endoscopic surgical robots โ bridging the gap between ๐ฌ๐ฉ๐๐ซ๐ฌ๐ ๐ฌ๐๐ง๐ฌ๐จ๐ซ ๐๐๐ญ๐ ๐๐ง๐ ๐ก๐ข๐ ๐ก-๐ช๐ฎ๐๐ฅ๐ข๐ญ๐ฒ ๐๐ ๐ฆ๐๐ฉ๐ฉ๐ข๐ง๐ using a novel ๐๐ข๐๐๐ฎ๐ฌ๐ข๐จ๐ง-๐๐๐ฌ๐๐ framework.
Fine-tuning foundational models often fails due to a lack of dense ground truth, and self-supervised methods struggle with scale ambiguity, sparse depth sensors offer a reliable geometric prior.
This motivated us to develop EndoDDC, a method that robustly generates dense depth maps by fusing RGB images with sparse depth inputs.
This work demonstrates that diffusion models can effectively solve the “sparse-to-dense” challenge in medical imaging. By providing accurate depth completion despite complex lighting and texture conditions, EndoDDC has the potential to significantly enhance autonomous navigation, procedural safety, and spatial awareness in minimally invasive surgery.
We present ๐๐๐ฎ๐ซ๐จ-๐๐๐, an scenario-aware model designed for the motion control of a parallel continuum neurosurgical robot.
Robotic surgery systems have garnered significant attention for their precision and efficiency, yet achieving autonomous tasks in complex neurosurgical environments remains challenging. Although Vision-Language-Action (VLA) models hold great potential, their development is constrained by the scarcity of data from surgical environments and robotic kinematics. To address this issue, this paper proposes NeuroVLA: a VLA model specifically designed for neurosurgical robotic tumor debulking tasks. Through phantom experiments conducted on a flexible parallel continuum robot, we constructed a dataset and decomposed the debulking task into four skill-based instructions. NeuroVLA utilizes a Vision-Language Model (VLM) as its backbone for scene reasoning, enabling the robot to comprehend the surgical scene and its own state. Experimental results demonstrate that after training on 90 debulking segments, NeuroVLA can infer actions based on images, language instructions, and the robotโs state. It achieved average pixel distance errors of 29.10 pixels and 21.55 pixels for the “alignment” and “transfer” skills, respectively, and success rates of 88.89% and 100% for the “grasping” and “release” skills.
๐ง Technical Framework:
โ End-to-End scenario-aware VLA model
โ Skill-based scenario infer mechanism
โ Debulking task dataset in neurosurgery
๐ฏ Experimental Results:
โ NeuroVLA demonstrates significantly lower pixel distance (PD) errors in the “alignment” and “transfer” skills (29.10 px / 21.55 px), far surpassing the performance of baseline models (such as Octo’s 79.72 px / 65.46 px).
In the “grasping” and “release” skills, NeuroVLA exhibits greater robustness, achieving a grasping success rate of 88.89% and a release success rate of 100%. In contrast, baseline models often misinterpret incomplete forceps closure as task completion, leading to grasping failures.
We present ๐๐ข๐ซ๐ข-๐๐๐ฉ๐ฌ๐ฎ๐ฅ๐, a swallowable kirigami-inspired capsule robot that enables minimally invasive GI biopsyโpushing capsule endoscopy from imaging to tissue sampling.
Wireless capsule endoscopy is comfortable and accessible, but cannot collect biopsy tissue, while histology is still the gold standard. Our work targets safe, depth-controlled, retrievable sampling in a capsule form factor.
๐ง Technical Framework:
โ Kirigami PI skin: flat during locomotion, deploys sharp protrusions when stretched
Thrilled to share our latest work, ๐๐ฎ๐ซ๐ ๐๐ข๐๐๐, the first video-language model specifically designed to address both full and fine-grained surgical video comprehension.
Surgical scene understanding is critical for training and robotic decision-making. While current Multimodal Large Language Models (MLLMs) excel at image analysis, they often overlook the fine-grained temporal reasoning required to capture detailed task execution and specific procedural processes within a surgery. This motivated us to bridge the gap between global video understanding and micro-action analysis.
๐ง โจ What we developed:
A novel framework and resource for surgical video reasoning that includes:
๐น ๐๐ฐ๐จ-๐ฌ๐ญ๐๐ ๐ ๐๐ญ๐๐ ๐๐ ๐จ๐๐ฎ๐ฌ ๐ฆ๐๐๐ก๐๐ง๐ข๐ฌ๐ฆ: The first stage extracts global procedural context, while the second stage performs high-frequency local analysis for fine-grained task execution.
๐น ๐๐ฎ๐ฅ๐ญ๐ข-๐๐ซ๐๐ช๐ฎ๐๐ง๐๐ฒ ๐ ๐ฎ๐ฌ๐ข๐จ๐ง ๐๐ญ๐ญ๐๐ง๐ญ๐ข๐จ๐ง (๐๐ ๐): Effectively integrates low-frequency global features with high-frequency local details to ensure comprehensive scene perception.
๐น ๐๐๐-๐๐๐ ๐๐๐ญ๐๐ฌ๐๐ญ: We constructed a large-scale dataset with over 31,000 video-instruction pairs, featuring hierarchical knowledge representation for enhanced visual reasoning.
๐ฏ Key Results:
โ SurgVidLM significantly outperforms existing models (like Qwen2-VL) in multi-grained surgical video understanding tasks.
โ Capable of inferring anatomical landmarks (e.g., Denonvilliers’ fascia) and providing clinical motivation, moving beyond simple visual description.
โ Demonstrated strong performance on unseen surgical tasks, proving the robustness of our hierarchical training approach.
๐ก Why it matters:
This work shows that by combining global context with localized high-frequency focus, we can significantly reduce “hallucinations” in surgical AI. It provides a pathway toward more intelligent, context-aware surgical assistants that can understand not just what is happening, but how and why specific steps are performed.
๐ฑ Whatโs next?
We are exploring how to extend this multi-grained understanding to real-time intraoperative guidance and integrating it with physical robotic control for autonomous sub-tasks.