Prof. Hongliang Ren Delivers Invited Talk at YAC2026 High-Level Talent Forum III

Changsha, China, May 9, 2026 โ€” Professor Hongliang Ren from the Department of Electronic Engineering, The Chinese University of Hong Kong, was invited to speak at High-Level Talent Forum III during the 41st Youth Academic Annual Conference of Chinese Association of Automation (YAC2026).

At High-Level Talent Forum III, Professor Ren delivered an invited presentation titled โ€œFlexible Robots and Embodied Intelligence for Minimally Invasive Intraluminal Procedures.โ€ His talk focused on recent advances in dexterous robotic motion generation and motion perception for image-guided minimally invasive procedures. He discussed how flexible robotic systems, multi-sensory perception, and embodied intelligence can support more precise, adaptable, and repeatable robotic intervention in confined anatomical environments.

Participants at YAC2026. From left to right: Prof. Guoniu Zhu, Fudan University; Prof. Shumei Yu, Soochow University; Prof. Qin Wan, Changsha University of Science and Technology, moderator of the session; Prof. Hongliang Ren, The Chinese University of Hong Kong; Prof. Ning Tan, Sun Yat-sen University; Prof. Quanquan Liu, Shenzhen MSU-BIT University.

Professor Renโ€™s research spans medical embodied intelligence, biomedical and surgical robotics, intelligent control, medical mechatronics, soft continuum robots, soft sensors, and multi-sensory learning for medical robotics.

Lab members and alumni at YAC2026. From left to right: Prof. Yunkai Lรผ, East China University of Science and Technology; Prof. Yanjie Chen, National University of Defense Technology; Prof. Ning Tan, Sun Yat-sen University; Prof. Shumei Yu, Soochow University; Prof. Hongliang Ren, The Chinese University of Hong Kong; Prof. Quanquan Liu, Shenzhen MSU-BIT University; Prof. Guoniu Zhu, Fudan University; Dr. Min Wang, lab member.

About YAC

YAC is a national annual academic conference organized by the Chinese Association of Automation and its Youth Working Committee. The 2026 conference was held in Changsha from May 8 to 10, 2026, and hosted by Hunan University. It brought together researchers, young scholars, graduate students, and professionals from automation, artificial intelligence, robotics, intelligent systems, and related disciplines. The conference provided a platform for presenting recent theoretical advances, emerging technologies, and interdisciplinary research outcomes.

Prof. Pierre Dupont Delivers Faculty of Engineering Distinguished Lecture on Continuum Robotics

May 6, 2026ย โ€“ The research group of Prof. Hongliang Ren was pleased to host Prof. Pierre Dupont as part of the Faculty of Engineering Distinguished Lecture series.

Prof. Dupont is a Professor of Surgery at Harvard Medical School, Chief of Pediatric Cardiac Bioengineering at Boston Childrenโ€™s Hospital, and an IEEE Fellow. In his talk titled โ€œContinuum Robotics for Minimally Invasive Interventions: Design Tradeoffs for Clinical Applications,โ€ he presented an engineering perspective grounded in realโ€‘world clinical challenges.

Prof. Dupont discussed three complementary continuum robot architectures โ€“ concentric tube robots, magnetic ball chains, and tendonโ€‘actuated systems โ€“ each involving distinct tradeoffs in stiffness, steerability, force capability, and hysteresis.

Key highlights from the lecture:

  • Concentric tube robots enable precise intracardiac catheterization and bimanual neuroendoscopy.
  • Magnetic ball chains offer outstanding steerability for cardiac ablation procedures.
  • Tendonโ€‘actuated systems provide a balanced tradeoff between stiffness and dexterity for transcatheter heart valve repair.

The lecture illustrated how tailoring robotic designs to specific clinical constraints can meaningfully expand the capabilities of minimally invasive interventions.

Short bio of Prof. Pierre Dupont

Prof. Dupont received his B.S., M.S., and Ph.D. in Mechanical Engineering from Rensselaer Polytechnic Institute. He was a Postdoctoral Fellow at Harvard University, later a Professor of Mechanical and Biomedical Engineering at Boston University, and is now Chief of Pediatric Cardiac Bioengineering at Boston Childrenโ€™s Hospital and Professor of Surgery at Harvard Medical School. He is an IEEE Fellow, a former Senior Editor for IEEE Transactions on Robotics, and a member of the Advisory Board for Science Robotics.

Audience engagement

The lecture attracted faculty members and students from engineering and medical backgrounds, followed by a lively Q&A session. Attendees appreciated Prof. Dupontโ€™s clinically driven approach to robotic design, and the discussion continued informally after the talk.

๐Ÿš€ Nature Communication 2026: Single twistable tendon-driven continuum robots

๐Ÿค–๐Ÿชข
Thrilled to share our latest work published on Nature Communication, which redefines actuation for tendon-driven continuum robots โ€” achieving full 3D omnidirectional motion and body twist using only a single tendon.


๐Ÿง โœจ What we developed: A new class of continuum robots that:
๐Ÿ”น Breaks Design Constraints: Eliminates the inherent trade-off between miniaturization and 3D manipulability by replacing multiple tendons with a single eccentric one.
๐Ÿ”น Push-Pull-Twist Actuation: Achieves complex spatial movement through a unique driving mechanism.
๐Ÿ”น High Efficiency: Features an outer diameter of 2.0โ€“3.5 mm with a hollow ratio exceeding 57% โ€” doubling the spatial utilization of traditional designs.
๐Ÿ”น Open-Source Support: Includes a derived kinematics model and an open-source simulator for the robotics community.


๐ŸŽฏ Key Results:

  • โœ… >1,000-fold Improvement: Massive increase in manipulability compared to conventional multi-tendon mechanisms.
  • โœ… High Force Retention: Retains at least 70% of tip force across all directions.
  • โœ… Versatile Demonstration: Proven success in teleoperation, navigation through tortuous environments, and “chopstick-like” continuum grippers.

๐Ÿ’ก Why it matters: This work proves that miniature robots can maintain high dexterity and power without the bulk of traditional hardware, pointing toward the next generation of surgical actuators.


๐ŸŒฑ Whatโ€™s next? We are exploring potential medical applications and the integration of these actuators into complex surgical procedures.

RENLab Visits The Third Affiliated Hospital of Sun Yat-sen University for Embodied Medical Robotics Academic Salon

On the afternoon of April 24, 2026, Professor Hongliang Renโ€™s research team from The Chinese University of Hong Kong visited The Third Affiliated Hospital of Sun Yat-sen University (SYSU Third Hospital) to hold an academic symposium and salon on Embodied Medical Robotics. The event was held in Conference Room 2006 of the Comprehensive Building, aiming to advance academic communication and potential research collaboration in medical robotics and intelligent healthcare.

The meeting was chaired by Dr. Liu Zifeng, Director of the Big Data and Artificial Intelligence Center of SYSU Third Hospital. Vice President Qintai Yang delivered a welcome speech, introducing the hospitalโ€™s clinical strengths and looking forward to close cooperation in medical robot research and translation.

Professor Hongliang Ren delivered a keynote report on minimally invasive flexible robotic systems and embodied intelligence in medicine. Members of his research team then presented recent advances in precision interventional tools, magnetic robotic systems, endoscopic intelligent navigation, autonomous surgical control, and medical image perception, showcasing the teamโ€™s innovative work in clinical-oriented medical robotics.

During the discussion session, clinicians and researchers from multiple departments of SYSU Third Hospital had in-depth exchanges on clinical demands, technical applications, and joint research plans. Both sides reached positive consensus on future cooperation in scientific research, talent development, and clinical translation.

At the end of the meeting, Vice President Qintai Yang delivered a concluding speech and spoke highly of this academic exchange. This symposium effectively bridged engineering research and clinical practice, and laid a solid foundation for the future development and application of embodied medical robotics.

About The Third Affiliated Hospital of Sun Yat-sen University

The Third Affiliated Hospital of Sun Yat-sen University was founded in 1971 and is a comprehensive Grade-A tertiary hospital directly administered by the National Health Commission of China. As a major clinical teaching base of Sun Yat-sen University, the hospital undertakes medical care, education, research, prevention, rehabilitation, and specialist training. It currently operates four campuses: Tianhe, Lingnan, Yuedong, and Zhaoqing. The hospital has developed distinctive strengths in liver disease, brain disorders, and immune-related diseases, supported by national key disciplines and national clinical key specialty programs. It is also recognized as a Guangdong High-level Hospital and serves as an output hospital for the development of national regional medical centers.

๐Ÿš€ย NVIDIA GTC 2026: Open-H-Embodiment โ€” The World’s First and Largest Open-Source Medical Robotics Dataset

Thrilled to share our latest international collaboration! At NVIDIA GTC 2026 in San Jose, CA, the team led by Professor Hongliang Ren from The Chinese University of Hong Kong (CUHK), in partnership with NVIDIA and 35 leading global institutions, officially released Open-H-Embodiment, the worldโ€™s first and largest open-source dataset for medical robotics, now available on HuggingFace.

During the GTC keynote, Kimberly Powell, NVIDIAโ€™s VP of Healthcare, highlighted this milestone. Our lab is honored to be a primary contributor, filling the critical gap in Embodied AI for medical robotics by providing high-fidelity data for contact dynamics and closed-loop control.

๐Ÿง โœจ What we contributed & developed:

This project breaks the “perception-heavy, execution-light” limitation of traditional medical AI. Key highlights include:

๐Ÿ”น 778 Hours of Massive Multimodal Data: The dataset covers 400 complete clinical surgeries and 9 major robotic platforms (e.g., dVRK, CMR Versius, Kuka). It includes 65% clinical data, 23% bench-top experiments, and 12% simulation data.

๐Ÿ”น Three High-Value Specialized Datasets from Our Lab:

  • Dual-Source Ultrasound Dataset:ย Experts-level trajectories covering in-vivo porcine EUS and human forearm scanning, overcoming complex organ environments and multi-device calibration.
  • Robotic Surgery Skill Dataset:ย Multi-modal data (RGB/RGB-D + Kinematics) for tissue manipulation and suturing, featuring millisecond-level synchronization and dual-mode control (teleoperation & automation).
  • Flexible Endoscope Tracking Baseline:ย A standardized dataset addressing hysteresis and deformation in flexible endoscopy, supporting nanosecond-level time synchronization.

๐Ÿ”น Surgical VLA & World Models:

  • GR00T-H:ย A 3B-parameter Vision-Language-Action model based on NVIDIA Isaac GR00T, capable of long-horizon dexterous tasks like end-to-end suturing.
  • Cosmos-H-Surgical-Simulator:ย An action-conditioned world model that boosts simulation efficiency by over 70x, bridging the sim-to-real gap.

๐ŸŽฏ Key Results: โœ… Global Standardization: First effort to unify medical robotic data across different devices and institutions under CC-BY-4.0. โœ… Efficiency Boost: Accelerated surgical simulation (600 sims in 40 mins) to generate high-fidelity video-action pairs. โœ… Clinical Relevance: Successfully captured nearly 500 hours of real-world clinical data for hernia, gallbladder, and uterine surgeries.

๐Ÿ’ก Why it matters: This initiative provides the foundational “bedrock” for Medical Physical AI. By sharing high-quality, synchronized data for surgery, ultrasound, and endoscopy, we are lowering the barrier for researchers worldwide to develop autonomous surgical agents that are both explainable and adaptive.

๐ŸŒฑ Whatโ€™s next? Our lab is continuing to deepen research in: ๐Ÿ”น Reasoning-based autonomous control for surgical robots. ๐Ÿ”น Cross-platform generalization of Medical VLA models. ๐Ÿ”น Clinical translation of Embodied AI to improve patient outcomes.

Datasets address: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Open-H-Embodiment

Project website: https://github.com/open-h

#NVIDIAGTC2026 #MedicalRobotics #EmbodiedAI #HuggingFace #CUHK #OpenSource #HealthcareInnovation

๐Ÿš€ ๐—ฆ๐—ฐ๐—ถ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—”๐—ฑ๐˜ƒ๐—ฎ๐—ป๐—ฐ๐—ฒ๐˜€ 2026: Inertia-Driven Amphibious โ€œLeglessbotโ€ with Asymmetric Microundulatory Fin Arrays ๐Ÿค–๐ŸŒŠ

Thrilled to share our latest #๐—ฆ๐—ฐ๐—ถ๐—ฒ๐—ป๐—ฐ๐—ฒ๐—”๐—ฑ๐˜ƒ๐—ฎ๐—ป๐—ฐ๐—ฒ work on a ๐—ฐ๐—ฒ๐—ป๐˜๐—ถ๐—บ๐—ฒ๐˜๐—ฒ๐—ฟ-๐˜€๐—ฐ๐—ฎ๐—น๐—ฒ, ๐—ณ๐˜‚๐—น๐—น๐˜† ๐˜€๐—ฒ๐—ฎ๐—น๐—ฒ๐—ฑ, ๐—น๐—ฒ๐—ด๐—น๐—ฒ๐˜€๐˜€ ๐—ฎ๐—บ๐—ฝ๐—ต๐—ถ๐—ฏ๐—ถ๐—ผ๐˜‚๐˜€ ๐—ฟ๐—ผ๐—ฏ๐—ผ๐˜ that can crawl on sand, jump, and steerably swim – powered by a ๐˜€๐—ถ๐—ป๐—ด๐—น๐—ฒ ๐˜ƒ๐—ฎ๐—ฟ๐—ถ๐—ฎ๐—ฏ๐—น๐—ฒ-๐—ผ๐˜‚๐˜๐—ฝ๐˜‚๐˜ ๐˜ƒ๐—ผ๐—ถ๐—ฐ๐—ฒ-๐—ฐ๐—ผ๐—ถ๐—น ๐—บ๐—ผ๐˜๐—ผ๐—ฟ (๐—ฉ๐—–๐— ).

At small scales, reliable ๐˜ธ๐˜ข๐˜ต๐˜ฆ๐˜ณ๐˜ฑ๐˜ณ๐˜ฐ๐˜ฐ๐˜ง ๐˜ด๐˜ฆ๐˜ข๐˜ญ๐˜ช๐˜ฏ๐˜จ is tough: transmissions and active mechanisms mean moving parts and dynamic seals that are fragile and ๐˜ญ๐˜ฆ๐˜ข๐˜ฌ-๐˜ฑ๐˜ณ๐˜ฐ๐˜ฏ๐˜ฆ. We wanted an amphibious robot that stays sealed and robust – yet still supports multiple locomotion modes.

๐Ÿง โœจ ๐—ช๐—ต๐—ฎ๐˜ ๐˜„๐—ฒ ๐—ฑ๐—ฒ๐˜ƒ๐—ฒ๐—น๐—ผ๐—ฝ๐—ฒ๐—ฑ:

An ๐—ถ๐—ป๐—ฒ๐—ฟ๐˜๐—ถ๐—ฎ-๐—ฑ๐—ฟ๐—ถ๐˜ƒ๐—ฒ๐—ป ๐—ฎ๐—ฐ๐˜๐˜‚๐—ฎ๐˜๐—ถ๐—ผ๐—ป + ๐—ฝ๐—ฎ๐˜€๐˜€๐—ถ๐˜ƒ๐—ฒ ๐—ฝ๐—ฟ๐—ผ๐—ฝ๐˜‚๐—น๐˜€๐—ถ๐—ผ๐—ป ๐—ฑ๐—ฒ๐˜€๐—ถ๐—ด๐—ป that:

๐Ÿ”น Uses a variable-output VCM inside a fully sealed rigid shell (no external moving parts)

๐Ÿ”น Switches among three modes: jumping, full-stroke vibration (land), and small-stroke vibration (water)

๐Ÿ”น Uses asymmetric, tilted passive fins for frequency-tuned steering in water (IDMP)

๐Ÿ”น Explains hydrodynamics via aquatic tests, high-speed PIV, and CFD

All of this – ๐—ผ๐—ป๐—ฒ ๐—บ๐—ฎ๐—ถ๐—ป ๐—น๐—ถ๐—ป๐—ฒ๐—ฎ๐—ฟ ๐—ฎ๐—ฐ๐˜๐˜‚๐—ฎ๐˜๐—ผ๐—ฟ + ๐—ฝ๐—ฎ๐˜€๐˜€๐—ถ๐˜ƒ๐—ฒ ๐—ณ๐—ถ๐—ป๐˜€. No exposed legs, gears, or propellers – making sealing and durability much easier at the centimeter scale.

๐ŸŽฏ ๐—ž๐—ฒ๐˜† ๐—ฅ๐—ฒ๐˜€๐˜‚๐—น๐˜๐˜€:

โœ… 24-g prototype (57.5 ร— 36 ร— 36 mm) with a fully enclosed shell

โœ… ~1.4 BL/s (~78 mm/s) on dry sand; 41.6 mm/s on flat ground

โœ… Jump height up to 17.16 mm at 15 V; continuous โ€œtumblerโ€ jumping

โœ… Load carrying: 960 g (~40ร— body weight) and escape under 5-kg loads

โœ… In water: min turning radius 5.6 mm; straight swim at 35/44/60 Hz (~32/35/28 mm/s); peak yaw -22ยฐ/s (30 Hz) or 16ยฐ/s (40 Hz)

๐Ÿ’ก ๐—ช๐—ต๐˜† ๐—ถ๐˜ ๐—บ๐—ฎ๐˜๐˜๐—ฒ๐—ฟ๐˜€:

This work shows how inertia + mode-switchable actuation can bridge the ๐—บ๐—ผ๐—บ๐—ฒ๐—ป๐˜๐˜‚๐—บ-๐—ณ๐—ฟ๐—ฒ๐—พ๐˜‚๐—ฒ๐—ป๐—ฐ๐˜† ๐˜๐—ฟ๐—ฎ๐—ฑ๐—ฒ-๐—ผ๐—ณ๐—ณ, enabling jumping and swimming in the same tiny robot. Passive asymmetric fins turn simple reciprocation into steerable thrust – ๐˜„๐—ถ๐˜๐—ต ๐—ป๐—ผ ๐—ฎ๐—ฑ๐—ฑ๐—ฒ๐—ฑ ๐—ฎ๐—ฐ๐˜๐˜‚๐—ฎ๐˜๐—ผ๐—ฟ๐˜€.

๐ŸŒฑ ๐—ช๐—ต๐—ฎ๐˜โ€™๐˜€ ๐—ป๐—ฒ๐˜…๐˜?

Future work will focus on improving environmental adaptability via ๐˜ง๐˜ฆ๐˜ฆ๐˜ฅ๐˜ฃ๐˜ข๐˜ค๐˜ฌ ๐˜ค๐˜ฐ๐˜ฏ๐˜ต๐˜ณ๐˜ฐ๐˜ญ and adaptive structures, and optimizing energy efficiency + onboard power ๐˜ฎ๐˜ช๐˜ฏ๐˜ช๐˜ข๐˜ต๐˜ถ๐˜ณ๐˜ช๐˜ป๐˜ข๐˜ต๐˜ช๐˜ฐ๐˜ฏ.

Special shoutout to the team –

Lingqi Tang, Yongzun Yang (co-first authors), Bing Li, Bingfu Zhang, Qiguang He, Hongliang Ren, Yao Li – for making this project possible.

๐Ÿ”— Paper link: https://lnkd.in/gFDfZYKq

๐Ÿ”– #ScienceAdvances #Robotics #AmphibiousRobots #Microrobots #InertiaDriven #Locomotion #VCM #PIV #CFD

No alternative text description for this image

๐Ÿ”ฌ We are pleased to announce the publication of our latest research, “๐—˜๐—ป๐—ฑ๐—ผ๐—–๐—ผ๐—ป๐˜๐—ฟ๐—ผ๐—น๐— ๐—ฎ๐—ด: ๐—ฅ๐—ผ๐—ฏ๐˜‚๐˜€๐˜ ๐—ฒ๐—ป๐—ฑ๐—ผ๐˜€๐—ฐ๐—ผ๐—ฝ๐—ถ๐—ฐ ๐˜ƒ๐—ฎ๐˜€๐—ฐ๐˜‚๐—น๐—ฎ๐—ฟ ๐—บ๐—ผ๐˜๐—ถ๐—ผ๐—ป ๐—บ๐—ฎ๐—ด๐—ป๐—ถ๐—ณ๐—ถ๐—ฐ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐˜„๐—ถ๐˜๐—ต ๐—ฝ๐—ฒ๐—ฟ๐—ถ๐—ผ๐—ฑ๐—ถ๐—ฐ ๐—ฟ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—ฟ๐—ฒ๐˜€๐—ฒ๐˜๐˜๐—ถ๐—ป๐—ด ๐—ฎ๐—ป๐—ฑ ๐—ต๐—ถ๐—ฒ๐—ฟ๐—ฎ๐—ฟ๐—ฐ๐—ต๐—ถ๐—ฐ๐—ฎ๐—น ๐˜๐—ถ๐˜€๐˜€๐˜‚๐—ฒ-๐—ฎ๐˜„๐—ฎ๐—ฟ๐—ฒ ๐—ฑ๐˜‚๐—ฎ๐—น-๐—บ๐—ฎ๐˜€๐—ธ ๐—ฐ๐—ผ๐—ป๐˜๐—ฟ๐—ผ๐—น” in ๐—”๐—ฑ๐˜ƒ๐—ฎ๐—ป๐—ฐ๐—ฒ๐—ฑ ๐—˜๐—ป๐—ด๐—ถ๐—ป๐—ฒ๐—ฒ๐—ฟ๐—ถ๐—ป๐—ด ๐—œ๐—ป๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐˜๐—ถ๐—ฐ๐˜€!

๐Ÿฅ Accurate visualization of subtle vascular dynamics remains a significant challenge in minimally invasive surgery, where dynamic complexities often limit decision-making reliability. Our paper introduces ๐—˜๐—ป๐—ฑ๐—ผ๐—–๐—ผ๐—ป๐˜๐—ฟ๐—ผ๐—น๐— ๐—ฎ๐—ด, a framework designed to ๐—ฒ๐—ป๐—ต๐—ฎ๐—ป๐—ฐ๐—ฒ ๐˜ƒ๐—ฎ๐˜€๐—ฐ๐˜‚๐—น๐—ฎ๐—ฟ ๐—บ๐—ผ๐˜๐—ถ๐—ผ๐—ป ๐˜ƒ๐—ถ๐˜€๐—ถ๐—ฏ๐—ถ๐—น๐—ถ๐˜๐˜† in endoscopic videos while preserving surrounding tissue structure. The approach integrates ๐—ฃ๐—ฒ๐—ฟ๐—ถ๐—ผ๐—ฑ๐—ถ๐—ฐ ๐—ฅ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—ฅ๐—ฒ๐˜€๐—ฒ๐˜๐˜๐—ถ๐—ป๐—ด to minimize error accumulation over time and ๐—›๐—ถ๐—ฒ๐—ฟ๐—ฎ๐—ฟ๐—ฐ๐—ต๐—ถ๐—ฐ๐—ฎ๐—น ๐—ง๐—ถ๐˜€๐˜€๐˜‚๐—ฒ-๐—ฎ๐˜„๐—ฎ๐—ฟ๐—ฒ ๐— ๐—ฎ๐—ด๐—ป๐—ถ๐—ณ๐—ถ๐—ฐ๐—ฎ๐˜๐—ถ๐—ผ๐—ป for adaptive vessel tracking.

๐Ÿ“Š To validate robustness, we constructed ๐—˜๐—ป๐—ฑ๐—ผ๐—ฉ๐— ๐— ๐Ÿฎ๐Ÿฐ, a benchmark dataset spanning four surgical specialties and diverse intraoperative scenarios. Quantitative metrics and expert surgeon evaluations indicate improved magnification accuracy and image quality compared to existing methods.

๐Ÿค We extend our sincere gratitude to our collaborators across The Chinese University of Hong Kong (An Wang, Mengya Xu, Yiting Chang, Prof Hongliang Ren), The University of Hong Kong (Rulin Zhou), Southern Medical University (่‹Ÿ้พ™้ฃž, Prof Hao Chen), The First Affiliated Hospital of Wenzhou Medical University (Yiru Ye), Southern University of Science and

Technology (Prof Jiankun Wang), and Singapore General Hospital (Prof Chwee Ming Lim) for their invaluable contributions to this multidisciplinary work.

The paper is available at https://lnkd.in/ggPEFswF

#MotionMagnification #SurgicalAI #Endoscopy

No alternative text description for this image

๐Ÿš€ ICRA 2026: ๐‘ฎ๐’†๐’๐‘ณ๐’‚๐’๐‘ฎ: ๐‘ฎ๐’†๐’๐’Ž๐’†๐’•๐’“๐’š-๐‘จ๐’˜๐’‚๐’“๐’† ๐‘ณ๐’‚๐’๐’ˆ๐’–๐’‚๐’ˆ๐’†-๐‘ฎ๐’–๐’Š๐’…๐’†๐’… ๐‘ฎ๐’“๐’‚๐’”๐’‘๐’Š๐’๐’ˆ ๐’˜๐’Š๐’•๐’‰ ๐‘ผ๐’๐’Š๐’‡๐’Š๐’†๐’… ๐‘น๐‘ฎ๐‘ฉ-๐‘ซ ๐‘ด๐’–๐’๐’•๐’Š๐’Ž๐’๐’…๐’‚๐’ ๐‘ณ๐’†๐’‚๐’“๐’๐’Š๐’๐’ˆ

Thrilled to share our latest work, ๐†๐ž๐จ๐‹๐š๐ง๐†, a unified geometry-aware framework for language-guided robotic grasping.

Language-guided grasping is a key capability for intuitive humanโ€“robot interaction. A robot should not only detect objects but also understand natural instructions such as โ€œpick up the blue cup behind the bowl.โ€ While recent multimodal models have shown promising results, most existing approaches rely on multi-stage pipelines that loosely couple perception and grasp prediction. These methods often overlook the tight integration of geometry, language, and visual reasoning, making them fragile in cluttered, occluded, or low-texture environments. This motivated us to bridge the gap between semantic language understanding and precise geometric grasp execution.

๐Ÿง โœจ ๐–๐ก๐š๐ญ ๐ฐ๐ž ๐๐ž๐ฏ๐ž๐ฅ๐จ๐ฉ๐ž๐:

A novel unified framework for geometry-aware language-guided grasping that includes:

๐Ÿ”น Unified RGB-D Multimodal Representation:

 We embed RGB, depth, and language features into a shared representation space, enabling consistent cross-modal semantic alignment for accurate target reasoning.

๐Ÿ”น Depth-Guided Geometric Module (DGGM):

 Instead of treating depth as auxiliary input, we explicitly inject geometric priors derived from depth into the attention mechanism, strengthening object discrimination under occlusion and ambiguous visual conditions.

๐Ÿ”น Adaptive Dense Channel Integration (ADCI):

 A dynamic multi-layer fusion strategy that balances global semantic cues and fine-grained geometric details for robust grasp prediction.

๐ŸŽฏ  ๐Š๐ž๐ฒ ๐‘๐ž๐ฌ๐ฎ๐ฅ๐ญ๐ฌ:

โœ… GeoLanG significantly outperforms prior multi-stage baselines on OCID-VLG for language-guided grasping.

โœ… Demonstrates strong robustness in cluttered and heavily occluded scenes.

โœ… Successfully validated on real robotic hardware, showing reliable sim-to-real transfer.

๐Ÿ’ก ๐–๐ก๐ฒ ๐ข๐ญ ๐ฆ๐š๐ญ๐ญ๐ž๐ซ๐ฌ:

This work shows that tightly coupling geometric reasoning with multimodal language understanding can significantly enhance robotic grasp reliability. By embedding depth-aware geometric priors directly into attention mechanisms, we reduce ambiguity and improve consistency in grasp decision-making.

GeoLanG provides a pathway toward more intelligent robotic systems that understand not just what object to grasp, but also how to grasp it robustly in complex real-world environments.

๐ŸŒฑ ๐–๐ก๐š๐ญโ€™๐ฌ ๐ง๐ž๐ฑ๐ญ?

We are exploring extending this geometry-aware multimodal reasoning toward:

 ๐Ÿ”น Real-time interactive grasping

 ๐Ÿ”น Multi-step manipulation tasks

 ๐Ÿ”น Integration with motion planning and autonomous robotic control

#ICRA2026 #CUHK

No alternative text description for this image
No alternative text description for this image

๐Ÿš€ ICRA 2026: ๐‘ฌ๐’๐’…๐’๐‘ซ๐‘ซ๐‘ช: ๐‘ณ๐’†๐’‚๐’“๐’๐’Š๐’๐’ˆ ๐‘บ๐’‘๐’‚๐’“๐’”๐’† ๐’•๐’ ๐‘ซ๐’†๐’๐’”๐’† ๐‘น๐’†๐’„๐’๐’๐’”๐’•๐’“๐’–๐’„๐’•๐’Š๐’๐’ ๐’‡๐’๐’“ ๐‘ฌ๐’๐’…๐’๐’”๐’„๐’๐’‘๐’Š๐’„ ๐‘น๐’๐’ƒ๐’๐’•๐’Š๐’„ ๐‘ต๐’‚๐’—๐’Š๐’ˆ๐’‚๐’•๐’Š๐’๐’ ๐’—๐’Š๐’‚ ๐‘ซ๐’Š๐’‡๐’‡๐’–๐’”๐’Š๐’๐’ ๐‘ซ๐’†๐’‘๐’•๐’‰ ๐‘ช๐’๐’Ž๐’‘๐’๐’†๐’•๐’Š๐’๐’ ๐Ÿค–

Thrilled to share our latest work on enabling robust sparse-to-dense reconstruction for endoscopic surgical robots โ€” bridging the gap between ๐ฌ๐ฉ๐š๐ซ๐ฌ๐ž ๐ฌ๐ž๐ง๐ฌ๐จ๐ซ ๐๐š๐ญ๐š ๐š๐ง๐ ๐ก๐ข๐ ๐ก-๐ช๐ฎ๐š๐ฅ๐ข๐ญ๐ฒ ๐Ÿ‘๐ƒ ๐ฆ๐š๐ฉ๐ฉ๐ข๐ง๐  using a novel ๐๐ข๐Ÿ๐Ÿ๐ฎ๐ฌ๐ข๐จ๐ง-๐›๐š๐ฌ๐ž๐ framework.

Fine-tuning foundational models often fails due to a lack of dense ground truth, and self-supervised methods struggle with scale ambiguity, sparse depth sensors offer a reliable geometric prior.

This motivated us to develop EndoDDC, a method that robustly generates dense depth maps by fusing RGB images with sparse depth inputs.

๐Ÿง โœจ ๐–๐ก๐š๐ญ ๐ฐ๐ž ๐๐ž๐ฏ๐ž๐ฅ๐จ๐ฉ๐ž๐:

A diffusion-driven depth completion architecture that:

๐Ÿ”น Integrates sparse depth and RGB inputs to overcome the limitations of pure visual estimation.

๐Ÿ”น Utilizes a Multi-scale Feature Extraction and Depth Gradient Fusion module to capture fine-grained surface orientation and local structure.

๐Ÿ”น Optimizes depth maps iteratively using a conditional diffusion model, refining geometry even in regions with weak textures or reflections.

๐ŸŽฏ ๐Š๐ž๐ฒ ๐‘๐ž๐ฌ๐ฎ๐ฅ๐ญ๐ฌ:

โœ… 25.55% and 9.03% improvement in accuracy on the StereoMIS and C3VD dataset compared to SOTA surgical estimators like EndoDAC.

โœ… 7.35% and 5.28% reduction in RMSE on StereoMIS and C3VD compared to the best depth completion baseline (OGNI-DC).

โœ… Outperformed foundational models (DepthAnything-v2) and standard depth completion (Marigold-DC) methods in both accuracy and robustness.

๐Ÿ’ก ๐–๐ก๐ฒ ๐ข๐ญ ๐ฆ๐š๐ญ๐ญ๐ž๐ซ๐ฌ:

This work demonstrates that diffusion models can effectively solve the “sparse-to-dense” challenge in medical imaging. By providing accurate depth completion despite complex lighting and texture conditions, EndoDDC has the potential to significantly enhance autonomous navigation, procedural safety, and spatial awareness in minimally invasive surgery.

๐Ÿ”– #DepthCompletion #DiffusionModel #EndoscopicSurgery #SurgicalNavigation #ICRA #CUHKEngineering #CUHK

No alternative text description for this image
No alternative text description for this image

๐Ÿš€ ICRA 2026: ๐‘ต๐’†๐’–๐’“๐’๐‘ฝ๐‘ณ๐‘จ: ๐‘บ๐’–๐’“๐’ˆ๐’Š๐’„๐’‚๐’ ๐‘บ๐’„๐’†๐’๐’‚๐’“๐’Š๐’-๐‘จ๐’˜๐’‚๐’“๐’† ๐‘ณ๐’†๐’‚๐’“๐’๐’Š๐’๐’ˆ ๐’๐’‡ ๐‘ซ๐’†๐’ƒ๐’–๐’๐’Œ๐’Š๐’๐’ˆ ๐‘บ๐’Œ๐’Š๐’๐’๐’” ๐’Š๐’ ๐‘ฌ๐’๐’…๐’๐’”๐’„๐’๐’‘๐’Š๐’„ ๐‘น๐’๐’ƒ๐’๐’•๐’Š๐’„ ๐‘ต๐’†๐’–๐’“๐’๐’”๐’–๐’“๐’ˆ๐’†๐’“๐’š ๐’—๐’Š๐’‚ ๐‘ฝ๐’Š๐’”๐’Š๐’๐’-๐‘ณ๐’‚๐’๐’ˆ๐’–๐’‚๐’ˆ๐’†-๐‘จ๐’„๐’•๐’Š๐’๐’ ๐‘ด๐’๐’…๐’†๐’ ๐Ÿค–๐Ÿงฒ

We present ๐๐ž๐ฎ๐ซ๐จ-๐•๐‹๐€, an scenario-aware model designed for the motion control of a parallel continuum neurosurgical robot.

Robotic surgery systems have garnered significant attention for their precision and efficiency, yet achieving autonomous tasks in complex neurosurgical environments remains challenging. Although Vision-Language-Action (VLA) models hold great potential, their development is constrained by the scarcity of data from surgical environments and robotic kinematics. To address this issue, this paper proposes NeuroVLA: a VLA model specifically designed for neurosurgical robotic tumor debulking tasks. Through phantom experiments conducted on a flexible parallel continuum robot, we constructed a dataset and decomposed the debulking task into four skill-based instructions. NeuroVLA utilizes a Vision-Language Model (VLM) as its backbone for scene reasoning, enabling the robot to comprehend the surgical scene and its own state. Experimental results demonstrate that after training on 90 debulking segments, NeuroVLA can infer actions based on images, language instructions, and the robotโ€™s state. It achieved average pixel distance errors of 29.10 pixels and 21.55 pixels for the “alignment” and “transfer” skills, respectively, and success rates of 88.89% and 100% for the “grasping” and “release” skills.

๐Ÿง  Technical Framework:

โ—     End-to-End scenario-aware VLA model

โ—     Skill-based scenario infer mechanism

โ—     Debulking task dataset in neurosurgery

๐ŸŽฏ Experimental Results:

โ—     NeuroVLA demonstrates significantly lower pixel distance (PD) errors in the “alignment” and “transfer” skills (29.10 px / 21.55 px), far surpassing the performance of baseline models (such as Octo’s 79.72 px / 65.46 px).

In the “grasping” and “release” skills, NeuroVLA exhibits greater robustness, achieving a grasping success rate of 88.89% and a release success rate of 100%. In contrast, baseline models often misinterpret incomplete forceps closure as task completion, leading to grasping failures.

#ICRA2026

diagram