๐Ÿš€ From Seeing to Reasoning in Endoscopic Surgery ๐Ÿค–๐Ÿ‘จโ€โš•๏ธ

We are excited to share our latest Comment published in #npj Digital Medicine:
โ€œHow can reasoning capability empower the AI copilot robot in endoscopic surgeryโ€

Current AI copilots in endoscopic surgery are still largely reactive and vision-driven. While they can detect anatomy, instruments, and scene changes, they often still struggle to truly understand surgical intent, infer hidden tissue dynamics, and respond robustly to uncertaintyโ€”all of which are essential for safe and precise intraoperative assistance.

In this article, we highlight how reasoning capability can become a key enabler for the next generation of AI copilot robots in endoscopic surgery:

  1. ๐Ÿง  Reasoning beyond perception: Enabling VLA-based surgical robots to go beyond simple visual recognition and translate high-level surgeon intent into precise, context-aware low-level motion goals.
  2. ๐Ÿ”„ Multimodal and uncertainty-aware intelligence: Fusing endoscopic vision with preoperative imaging, intraoperative sensing, and tracking signals, while dynamically re-weighting information sources under occlusion, bleeding, smoke, and other uncertain conditions.
  3. ๐Ÿค Coordinated multi-instrument collaboration: Supporting synchronized control of multiple tools for subtasks such as traction, dissection, and hemostasis, with reasoning-guided adaptation to tissue deformation and workflow variation.
  4. ๐Ÿ”ฎ Anticipatory and safer decision-making: Using chain-of-thought-style reasoning to forecast tissue response, evaluate possible action outcomes, and generate more conservative and interpretable assistance under risk.
  5. ๐Ÿฅ Surgeon-in-the-loop clinical deployment: Framing the AI copilot robot at LoA 2โ€“3, where the system assists with task generation, monitoring, and bounded low-level execution under explicit safety constraints and continuous surgeon oversight.

We believe the future of endoscopic robotic assistance lies not only in systems that can see, but in systems that can reason, adapt, and collaborate. With reasoning-enabled VLA models, AI copilot robots may evolve from reactive executors into true cognitive partners in the operating room.

๐Ÿ“ƒ Read the full paper here: https://www.nature.com/articles/s41746-026-02827-8

๐Ÿ‘ Kudos to the team: Mr. Guankun Wang, Dr. Long Bai, and Prof. Hongliang Ren.

๐Ÿš€ Advanced Science 2026: Transferable Autonomous Endoscopy Navigation! ๐Ÿค–๐Ÿ’Š

Thrilled to share our latest Advanced Science work on enabling highly transferable, autonomous navigation for wireless capsule endoscopy (WCE)โ€”using a lightweight Edge-Contour-Depth Fusion module and deep reinforcement learning (DRL).

WCE has revolutionized GI diagnostics, but its potential is often restricted by incomplete mucosal coverage and the poor ability of existing AI navigation methods to adapt across different patient anatomies. This motivated us to ditch the heavy, brittle, traditional “end-to-end” visual video streams that cause AI models to overfit to a single patient.

๐Ÿง โœจ What we developed:
A unified, clinically viable framework that features:
๐Ÿ”น Anatomical Landmark Guidance: Operates on stable, low-dimensional coordinates of conserved gastric structures (the fundus and pyloric antrum) rather than high-dimensional raw video.
๐Ÿ”น Lightweight Perception Module: Combines classical Canny edge detection and Hu moments with a compact monocular depth network (DispNet) to run efficiently on low-power clinical hardware.
๐Ÿ”น Robust Sim-to-Real Pipeline: Utilizes a patient-specific digital twin combined with a model-free Adaptive Dynamic Programming (ADP) controller to actively neutralize real-world physical disturbances and actuator latency.

๐ŸŽฏ Key Results:
โœ… >97% mucosal coverage achieved within 50 seconds across 8 diverse, patient-derived stomach models in simulation.
โœ… 87% mean coverage stability and a 53% reduction in procedure time during real-world ex-vivo experiments compared to expert manual control.
โœ… Drastically reduced computational overhead, allowing deployment on low-cost processors (<2 TOPS).

๐Ÿ’ก Why it matters:
This study establishes a scalable paradigm that conquers the “reality gap” and patient anatomical variability in medical robotics. By decoupling perception from control, it removes the need for expensive, massive patient datasets and high-end GPUs, paving the way for operator-independent, intelligent GI diagnostics.

๐ŸŒฑ Whatโ€™s next?
We are expanding our training to encompass extreme pathological distortions (like hiatal hernias) and advancing toward fully wireless clinical deployment with dynamic, target-reaching capabilities for intraoperative pathologies.

๐Ÿ”— Paper Link: https://advanced.onlinelibrary.wiley.com/doi/10.1002/advs.202600008

๐Ÿš€ IEEE TMRB 2026: Soft Pouch Matrix Actuation for Multimodal Intraluminal Locomotion & Liquid Biopsy Robotics ๐Ÿค–๐Ÿ’ง

Thrilled to share our newly accepted paper in IEEE Transactions on Medical Robotics and Bionics, where we introduce a soft, endoscope-deployable microfluidic suction robot that combines multimodal intraluminal locomotion with localized aspiration and sampling for targeted mucus clearance and liquid biopsy.

๐Ÿง โœจ What we developed:

A soft intraluminal robotic platform that:

๐Ÿ”น Integrates Locomotion + Sampling: A pneumatically controlled 2ร—2 pouch matrix for multimodal actuation, paired with an independent microfluidic suction module for active liquid extraction and sample recovery.

๐Ÿ”น Enables Stable Pitch Control: A balloon-based pitch control mechanism improves controllability for intraluminal operation, with the best overall performance observed at an initial pressure range of 2โ€“3 kPa.

๐Ÿ”น Balances Compliance and Safety: Single-pouch characterization guided the selection of a 2 mm pouch radius, achieving up to 246.91% maximum deformation; burst tests show a system safety factor โ‰ˆ 5.27 under the reported operating conditions.

๐Ÿ”น Targets Real Clinical Pain Points: Designed for constrained lumens (e.g., distal airway) where conventional airway clearance approaches struggle with reach and effectiveness.

๐ŸŽฏ Key Results:

โœ… Multimodal mobility: Differential actuation achieves 26.9 mm/min forward speed and 4.86ยฐ yaw per drive cycle.

โœ… Robust suction across viscosities: Efficiently extracts 20โ€“80% glycerol solutions within 10 s (via parameter tuning).

โœ… In vivo feasibility: Endoscope-assisted porcine validation confirmed sequential pouch-driven motion and successful recovery of biological samples containing mucus and tissue fragments after saline irrigation.

๐Ÿ’ก Why it matters:

This work demonstrates a compliant, integrated โ€œmove + anchor + suctionโ€ approach for narrow lumensโ€”supporting safer localized intervention and sampling, with a path toward distal airway translation.

๐ŸŒฑ Whatโ€™s next?

Weโ€™re moving toward more automated closed-loop pneumatic control, improved steerability/navigation, and miniaturization for deeper airway accessโ€”while expanding validation in airway-specific models.

๐Ÿš€ Nature Communication 2026: Single twistable tendon-driven continuum robots

๐Ÿค–๐Ÿชข
Thrilled to share our latest work published on Nature Communication, which redefines actuation for tendon-driven continuum robots โ€” achieving full 3D omnidirectional motion and body twist using only a single tendon.


๐Ÿง โœจ What we developed: A new class of continuum robots that:
๐Ÿ”น Breaks Design Constraints: Eliminates the inherent trade-off between miniaturization and 3D manipulability by replacing multiple tendons with a single eccentric one.
๐Ÿ”น Push-Pull-Twist Actuation: Achieves complex spatial movement through a unique driving mechanism.
๐Ÿ”น High Efficiency: Features an outer diameter of 2.0โ€“3.5 mm with a hollow ratio exceeding 57% โ€” doubling the spatial utilization of traditional designs.
๐Ÿ”น Open-Source Support: Includes a derived kinematics model and an open-source simulator for the robotics community.


๐ŸŽฏ Key Results:

  • โœ… >1,000-fold Improvement: Massive increase in manipulability compared to conventional multi-tendon mechanisms.
  • โœ… High Force Retention: Retains at least 70% of tip force across all directions.
  • โœ… Versatile Demonstration: Proven success in teleoperation, navigation through tortuous environments, and “chopstick-like” continuum grippers.

๐Ÿ’ก Why it matters: This work proves that miniature robots can maintain high dexterity and power without the bulk of traditional hardware, pointing toward the next generation of surgical actuators.


๐ŸŒฑ Whatโ€™s next? We are exploring potential medical applications and the integration of these actuators into complex surgical procedures.

๐Ÿš€ย NVIDIA GTC 2026: Open-H-Embodiment โ€” The World’s First and Largest Open-Source Medical Robotics Dataset

Thrilled to share our latest international collaboration! At NVIDIA GTC 2026 in San Jose, CA, the team led by Professor Hongliang Ren from The Chinese University of Hong Kong (CUHK), in partnership with NVIDIA and 35 leading global institutions, officially released Open-H-Embodiment, the worldโ€™s first and largest open-source dataset for medical robotics, now available on HuggingFace.

During the GTC keynote, Kimberly Powell, NVIDIAโ€™s VP of Healthcare, highlighted this milestone. Our lab is honored to be a primary contributor, filling the critical gap in Embodied AI for medical robotics by providing high-fidelity data for contact dynamics and closed-loop control.

๐Ÿง โœจ What we contributed & developed:

This project breaks the “perception-heavy, execution-light” limitation of traditional medical AI. Key highlights include:

๐Ÿ”น 778 Hours of Massive Multimodal Data: The dataset covers 400 complete clinical surgeries and 9 major robotic platforms (e.g., dVRK, CMR Versius, Kuka). It includes 65% clinical data, 23% bench-top experiments, and 12% simulation data.

๐Ÿ”น Three High-Value Specialized Datasets from Our Lab:

  • Dual-Source Ultrasound Dataset:ย Experts-level trajectories covering in-vivo porcine EUS and human forearm scanning, overcoming complex organ environments and multi-device calibration.
  • Robotic Surgery Skill Dataset:ย Multi-modal data (RGB/RGB-D + Kinematics) for tissue manipulation and suturing, featuring millisecond-level synchronization and dual-mode control (teleoperation & automation).
  • Flexible Endoscope Tracking Baseline:ย A standardized dataset addressing hysteresis and deformation in flexible endoscopy, supporting nanosecond-level time synchronization.

๐Ÿ”น Surgical VLA & World Models:

  • GR00T-H:ย A 3B-parameter Vision-Language-Action model based on NVIDIA Isaac GR00T, capable of long-horizon dexterous tasks like end-to-end suturing.
  • Cosmos-H-Surgical-Simulator:ย An action-conditioned world model that boosts simulation efficiency by over 70x, bridging the sim-to-real gap.

๐ŸŽฏ Key Results: โœ… Global Standardization: First effort to unify medical robotic data across different devices and institutions under CC-BY-4.0. โœ… Efficiency Boost: Accelerated surgical simulation (600 sims in 40 mins) to generate high-fidelity video-action pairs. โœ… Clinical Relevance: Successfully captured nearly 500 hours of real-world clinical data for hernia, gallbladder, and uterine surgeries.

๐Ÿ’ก Why it matters: This initiative provides the foundational “bedrock” for Medical Physical AI. By sharing high-quality, synchronized data for surgery, ultrasound, and endoscopy, we are lowering the barrier for researchers worldwide to develop autonomous surgical agents that are both explainable and adaptive.

๐ŸŒฑ Whatโ€™s next? Our lab is continuing to deepen research in: ๐Ÿ”น Reasoning-based autonomous control for surgical robots. ๐Ÿ”น Cross-platform generalization of Medical VLA models. ๐Ÿ”น Clinical translation of Embodied AI to improve patient outcomes.

Datasets address: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Open-H-Embodiment

Project website: https://github.com/open-h

#NVIDIAGTC2026 #MedicalRobotics #EmbodiedAI #HuggingFace #CUHK #OpenSource #HealthcareInnovation

๐Ÿš€ ๐—ฆ๐—ฐ๐—ถ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—”๐—ฑ๐˜ƒ๐—ฎ๐—ป๐—ฐ๐—ฒ๐˜€ 2026: Inertia-Driven Amphibious โ€œLeglessbotโ€ with Asymmetric Microundulatory Fin Arrays ๐Ÿค–๐ŸŒŠ

Thrilled to share our latest #๐—ฆ๐—ฐ๐—ถ๐—ฒ๐—ป๐—ฐ๐—ฒ๐—”๐—ฑ๐˜ƒ๐—ฎ๐—ป๐—ฐ๐—ฒ work on a ๐—ฐ๐—ฒ๐—ป๐˜๐—ถ๐—บ๐—ฒ๐˜๐—ฒ๐—ฟ-๐˜€๐—ฐ๐—ฎ๐—น๐—ฒ, ๐—ณ๐˜‚๐—น๐—น๐˜† ๐˜€๐—ฒ๐—ฎ๐—น๐—ฒ๐—ฑ, ๐—น๐—ฒ๐—ด๐—น๐—ฒ๐˜€๐˜€ ๐—ฎ๐—บ๐—ฝ๐—ต๐—ถ๐—ฏ๐—ถ๐—ผ๐˜‚๐˜€ ๐—ฟ๐—ผ๐—ฏ๐—ผ๐˜ that can crawl on sand, jump, and steerably swim – powered by a ๐˜€๐—ถ๐—ป๐—ด๐—น๐—ฒ ๐˜ƒ๐—ฎ๐—ฟ๐—ถ๐—ฎ๐—ฏ๐—น๐—ฒ-๐—ผ๐˜‚๐˜๐—ฝ๐˜‚๐˜ ๐˜ƒ๐—ผ๐—ถ๐—ฐ๐—ฒ-๐—ฐ๐—ผ๐—ถ๐—น ๐—บ๐—ผ๐˜๐—ผ๐—ฟ (๐—ฉ๐—–๐— ).

At small scales, reliable ๐˜ธ๐˜ข๐˜ต๐˜ฆ๐˜ณ๐˜ฑ๐˜ณ๐˜ฐ๐˜ฐ๐˜ง ๐˜ด๐˜ฆ๐˜ข๐˜ญ๐˜ช๐˜ฏ๐˜จ is tough: transmissions and active mechanisms mean moving parts and dynamic seals that are fragile and ๐˜ญ๐˜ฆ๐˜ข๐˜ฌ-๐˜ฑ๐˜ณ๐˜ฐ๐˜ฏ๐˜ฆ. We wanted an amphibious robot that stays sealed and robust – yet still supports multiple locomotion modes.

๐Ÿง โœจ ๐—ช๐—ต๐—ฎ๐˜ ๐˜„๐—ฒ ๐—ฑ๐—ฒ๐˜ƒ๐—ฒ๐—น๐—ผ๐—ฝ๐—ฒ๐—ฑ:

An ๐—ถ๐—ป๐—ฒ๐—ฟ๐˜๐—ถ๐—ฎ-๐—ฑ๐—ฟ๐—ถ๐˜ƒ๐—ฒ๐—ป ๐—ฎ๐—ฐ๐˜๐˜‚๐—ฎ๐˜๐—ถ๐—ผ๐—ป + ๐—ฝ๐—ฎ๐˜€๐˜€๐—ถ๐˜ƒ๐—ฒ ๐—ฝ๐—ฟ๐—ผ๐—ฝ๐˜‚๐—น๐˜€๐—ถ๐—ผ๐—ป ๐—ฑ๐—ฒ๐˜€๐—ถ๐—ด๐—ป that:

๐Ÿ”น Uses a variable-output VCM inside a fully sealed rigid shell (no external moving parts)

๐Ÿ”น Switches among three modes: jumping, full-stroke vibration (land), and small-stroke vibration (water)

๐Ÿ”น Uses asymmetric, tilted passive fins for frequency-tuned steering in water (IDMP)

๐Ÿ”น Explains hydrodynamics via aquatic tests, high-speed PIV, and CFD

All of this – ๐—ผ๐—ป๐—ฒ ๐—บ๐—ฎ๐—ถ๐—ป ๐—น๐—ถ๐—ป๐—ฒ๐—ฎ๐—ฟ ๐—ฎ๐—ฐ๐˜๐˜‚๐—ฎ๐˜๐—ผ๐—ฟ + ๐—ฝ๐—ฎ๐˜€๐˜€๐—ถ๐˜ƒ๐—ฒ ๐—ณ๐—ถ๐—ป๐˜€. No exposed legs, gears, or propellers – making sealing and durability much easier at the centimeter scale.

๐ŸŽฏ ๐—ž๐—ฒ๐˜† ๐—ฅ๐—ฒ๐˜€๐˜‚๐—น๐˜๐˜€:

โœ… 24-g prototype (57.5 ร— 36 ร— 36 mm) with a fully enclosed shell

โœ… ~1.4 BL/s (~78 mm/s) on dry sand; 41.6 mm/s on flat ground

โœ… Jump height up to 17.16 mm at 15 V; continuous โ€œtumblerโ€ jumping

โœ… Load carrying: 960 g (~40ร— body weight) and escape under 5-kg loads

โœ… In water: min turning radius 5.6 mm; straight swim at 35/44/60 Hz (~32/35/28 mm/s); peak yaw -22ยฐ/s (30 Hz) or 16ยฐ/s (40 Hz)

๐Ÿ’ก ๐—ช๐—ต๐˜† ๐—ถ๐˜ ๐—บ๐—ฎ๐˜๐˜๐—ฒ๐—ฟ๐˜€:

This work shows how inertia + mode-switchable actuation can bridge the ๐—บ๐—ผ๐—บ๐—ฒ๐—ป๐˜๐˜‚๐—บ-๐—ณ๐—ฟ๐—ฒ๐—พ๐˜‚๐—ฒ๐—ป๐—ฐ๐˜† ๐˜๐—ฟ๐—ฎ๐—ฑ๐—ฒ-๐—ผ๐—ณ๐—ณ, enabling jumping and swimming in the same tiny robot. Passive asymmetric fins turn simple reciprocation into steerable thrust – ๐˜„๐—ถ๐˜๐—ต ๐—ป๐—ผ ๐—ฎ๐—ฑ๐—ฑ๐—ฒ๐—ฑ ๐—ฎ๐—ฐ๐˜๐˜‚๐—ฎ๐˜๐—ผ๐—ฟ๐˜€.

๐ŸŒฑ ๐—ช๐—ต๐—ฎ๐˜โ€™๐˜€ ๐—ป๐—ฒ๐˜…๐˜?

Future work will focus on improving environmental adaptability via ๐˜ง๐˜ฆ๐˜ฆ๐˜ฅ๐˜ฃ๐˜ข๐˜ค๐˜ฌ ๐˜ค๐˜ฐ๐˜ฏ๐˜ต๐˜ณ๐˜ฐ๐˜ญ and adaptive structures, and optimizing energy efficiency + onboard power ๐˜ฎ๐˜ช๐˜ฏ๐˜ช๐˜ข๐˜ต๐˜ถ๐˜ณ๐˜ช๐˜ป๐˜ข๐˜ต๐˜ช๐˜ฐ๐˜ฏ.

Special shoutout to the team –

Lingqi Tang, Yongzun Yang (co-first authors), Bing Li, Bingfu Zhang, Qiguang He, Hongliang Ren, Yao Li – for making this project possible.

๐Ÿ”— Paper link: https://lnkd.in/gFDfZYKq

๐Ÿ”– #ScienceAdvances #Robotics #AmphibiousRobots #Microrobots #InertiaDriven #Locomotion #VCM #PIV #CFD

No alternative text description for this image

๐Ÿ”ฌ We are pleased to announce the publication of our latest research, “๐—˜๐—ป๐—ฑ๐—ผ๐—–๐—ผ๐—ป๐˜๐—ฟ๐—ผ๐—น๐— ๐—ฎ๐—ด: ๐—ฅ๐—ผ๐—ฏ๐˜‚๐˜€๐˜ ๐—ฒ๐—ป๐—ฑ๐—ผ๐˜€๐—ฐ๐—ผ๐—ฝ๐—ถ๐—ฐ ๐˜ƒ๐—ฎ๐˜€๐—ฐ๐˜‚๐—น๐—ฎ๐—ฟ ๐—บ๐—ผ๐˜๐—ถ๐—ผ๐—ป ๐—บ๐—ฎ๐—ด๐—ป๐—ถ๐—ณ๐—ถ๐—ฐ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐˜„๐—ถ๐˜๐—ต ๐—ฝ๐—ฒ๐—ฟ๐—ถ๐—ผ๐—ฑ๐—ถ๐—ฐ ๐—ฟ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—ฟ๐—ฒ๐˜€๐—ฒ๐˜๐˜๐—ถ๐—ป๐—ด ๐—ฎ๐—ป๐—ฑ ๐—ต๐—ถ๐—ฒ๐—ฟ๐—ฎ๐—ฟ๐—ฐ๐—ต๐—ถ๐—ฐ๐—ฎ๐—น ๐˜๐—ถ๐˜€๐˜€๐˜‚๐—ฒ-๐—ฎ๐˜„๐—ฎ๐—ฟ๐—ฒ ๐—ฑ๐˜‚๐—ฎ๐—น-๐—บ๐—ฎ๐˜€๐—ธ ๐—ฐ๐—ผ๐—ป๐˜๐—ฟ๐—ผ๐—น” in ๐—”๐—ฑ๐˜ƒ๐—ฎ๐—ป๐—ฐ๐—ฒ๐—ฑ ๐—˜๐—ป๐—ด๐—ถ๐—ป๐—ฒ๐—ฒ๐—ฟ๐—ถ๐—ป๐—ด ๐—œ๐—ป๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐˜๐—ถ๐—ฐ๐˜€!

๐Ÿฅ Accurate visualization of subtle vascular dynamics remains a significant challenge in minimally invasive surgery, where dynamic complexities often limit decision-making reliability. Our paper introduces ๐—˜๐—ป๐—ฑ๐—ผ๐—–๐—ผ๐—ป๐˜๐—ฟ๐—ผ๐—น๐— ๐—ฎ๐—ด, a framework designed to ๐—ฒ๐—ป๐—ต๐—ฎ๐—ป๐—ฐ๐—ฒ ๐˜ƒ๐—ฎ๐˜€๐—ฐ๐˜‚๐—น๐—ฎ๐—ฟ ๐—บ๐—ผ๐˜๐—ถ๐—ผ๐—ป ๐˜ƒ๐—ถ๐˜€๐—ถ๐—ฏ๐—ถ๐—น๐—ถ๐˜๐˜† in endoscopic videos while preserving surrounding tissue structure. The approach integrates ๐—ฃ๐—ฒ๐—ฟ๐—ถ๐—ผ๐—ฑ๐—ถ๐—ฐ ๐—ฅ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—ฅ๐—ฒ๐˜€๐—ฒ๐˜๐˜๐—ถ๐—ป๐—ด to minimize error accumulation over time and ๐—›๐—ถ๐—ฒ๐—ฟ๐—ฎ๐—ฟ๐—ฐ๐—ต๐—ถ๐—ฐ๐—ฎ๐—น ๐—ง๐—ถ๐˜€๐˜€๐˜‚๐—ฒ-๐—ฎ๐˜„๐—ฎ๐—ฟ๐—ฒ ๐— ๐—ฎ๐—ด๐—ป๐—ถ๐—ณ๐—ถ๐—ฐ๐—ฎ๐˜๐—ถ๐—ผ๐—ป for adaptive vessel tracking.

๐Ÿ“Š To validate robustness, we constructed ๐—˜๐—ป๐—ฑ๐—ผ๐—ฉ๐— ๐— ๐Ÿฎ๐Ÿฐ, a benchmark dataset spanning four surgical specialties and diverse intraoperative scenarios. Quantitative metrics and expert surgeon evaluations indicate improved magnification accuracy and image quality compared to existing methods.

๐Ÿค We extend our sincere gratitude to our collaborators across The Chinese University of Hong Kong (An Wang, Mengya Xu, Yiting Chang, Prof Hongliang Ren), The University of Hong Kong (Rulin Zhou), Southern Medical University (่‹Ÿ้พ™้ฃž, Prof Hao Chen), The First Affiliated Hospital of Wenzhou Medical University (Yiru Ye), Southern University of Science and

Technology (Prof Jiankun Wang), and Singapore General Hospital (Prof Chwee Ming Lim) for their invaluable contributions to this multidisciplinary work.

The paper is available at https://lnkd.in/ggPEFswF

#MotionMagnification #SurgicalAI #Endoscopy

No alternative text description for this image

๐Ÿš€ ICRA 2026: ๐‘ฎ๐’†๐’๐‘ณ๐’‚๐’๐‘ฎ: ๐‘ฎ๐’†๐’๐’Ž๐’†๐’•๐’“๐’š-๐‘จ๐’˜๐’‚๐’“๐’† ๐‘ณ๐’‚๐’๐’ˆ๐’–๐’‚๐’ˆ๐’†-๐‘ฎ๐’–๐’Š๐’…๐’†๐’… ๐‘ฎ๐’“๐’‚๐’”๐’‘๐’Š๐’๐’ˆ ๐’˜๐’Š๐’•๐’‰ ๐‘ผ๐’๐’Š๐’‡๐’Š๐’†๐’… ๐‘น๐‘ฎ๐‘ฉ-๐‘ซ ๐‘ด๐’–๐’๐’•๐’Š๐’Ž๐’๐’…๐’‚๐’ ๐‘ณ๐’†๐’‚๐’“๐’๐’Š๐’๐’ˆ

Thrilled to share our latest work, ๐†๐ž๐จ๐‹๐š๐ง๐†, a unified geometry-aware framework for language-guided robotic grasping.

Language-guided grasping is a key capability for intuitive humanโ€“robot interaction. A robot should not only detect objects but also understand natural instructions such as โ€œpick up the blue cup behind the bowl.โ€ While recent multimodal models have shown promising results, most existing approaches rely on multi-stage pipelines that loosely couple perception and grasp prediction. These methods often overlook the tight integration of geometry, language, and visual reasoning, making them fragile in cluttered, occluded, or low-texture environments. This motivated us to bridge the gap between semantic language understanding and precise geometric grasp execution.

๐Ÿง โœจ ๐–๐ก๐š๐ญ ๐ฐ๐ž ๐๐ž๐ฏ๐ž๐ฅ๐จ๐ฉ๐ž๐:

A novel unified framework for geometry-aware language-guided grasping that includes:

๐Ÿ”น Unified RGB-D Multimodal Representation:

 We embed RGB, depth, and language features into a shared representation space, enabling consistent cross-modal semantic alignment for accurate target reasoning.

๐Ÿ”น Depth-Guided Geometric Module (DGGM):

 Instead of treating depth as auxiliary input, we explicitly inject geometric priors derived from depth into the attention mechanism, strengthening object discrimination under occlusion and ambiguous visual conditions.

๐Ÿ”น Adaptive Dense Channel Integration (ADCI):

 A dynamic multi-layer fusion strategy that balances global semantic cues and fine-grained geometric details for robust grasp prediction.

๐ŸŽฏ  ๐Š๐ž๐ฒ ๐‘๐ž๐ฌ๐ฎ๐ฅ๐ญ๐ฌ:

โœ… GeoLanG significantly outperforms prior multi-stage baselines on OCID-VLG for language-guided grasping.

โœ… Demonstrates strong robustness in cluttered and heavily occluded scenes.

โœ… Successfully validated on real robotic hardware, showing reliable sim-to-real transfer.

๐Ÿ’ก ๐–๐ก๐ฒ ๐ข๐ญ ๐ฆ๐š๐ญ๐ญ๐ž๐ซ๐ฌ:

This work shows that tightly coupling geometric reasoning with multimodal language understanding can significantly enhance robotic grasp reliability. By embedding depth-aware geometric priors directly into attention mechanisms, we reduce ambiguity and improve consistency in grasp decision-making.

GeoLanG provides a pathway toward more intelligent robotic systems that understand not just what object to grasp, but also how to grasp it robustly in complex real-world environments.

๐ŸŒฑ ๐–๐ก๐š๐ญโ€™๐ฌ ๐ง๐ž๐ฑ๐ญ?

We are exploring extending this geometry-aware multimodal reasoning toward:

 ๐Ÿ”น Real-time interactive grasping

 ๐Ÿ”น Multi-step manipulation tasks

 ๐Ÿ”น Integration with motion planning and autonomous robotic control

#ICRA2026 #CUHK

No alternative text description for this image
No alternative text description for this image

๐Ÿš€ ICRA 2026: ๐‘ฌ๐’๐’…๐’๐‘ซ๐‘ซ๐‘ช: ๐‘ณ๐’†๐’‚๐’“๐’๐’Š๐’๐’ˆ ๐‘บ๐’‘๐’‚๐’“๐’”๐’† ๐’•๐’ ๐‘ซ๐’†๐’๐’”๐’† ๐‘น๐’†๐’„๐’๐’๐’”๐’•๐’“๐’–๐’„๐’•๐’Š๐’๐’ ๐’‡๐’๐’“ ๐‘ฌ๐’๐’…๐’๐’”๐’„๐’๐’‘๐’Š๐’„ ๐‘น๐’๐’ƒ๐’๐’•๐’Š๐’„ ๐‘ต๐’‚๐’—๐’Š๐’ˆ๐’‚๐’•๐’Š๐’๐’ ๐’—๐’Š๐’‚ ๐‘ซ๐’Š๐’‡๐’‡๐’–๐’”๐’Š๐’๐’ ๐‘ซ๐’†๐’‘๐’•๐’‰ ๐‘ช๐’๐’Ž๐’‘๐’๐’†๐’•๐’Š๐’๐’ ๐Ÿค–

Thrilled to share our latest work on enabling robust sparse-to-dense reconstruction for endoscopic surgical robots โ€” bridging the gap between ๐ฌ๐ฉ๐š๐ซ๐ฌ๐ž ๐ฌ๐ž๐ง๐ฌ๐จ๐ซ ๐๐š๐ญ๐š ๐š๐ง๐ ๐ก๐ข๐ ๐ก-๐ช๐ฎ๐š๐ฅ๐ข๐ญ๐ฒ ๐Ÿ‘๐ƒ ๐ฆ๐š๐ฉ๐ฉ๐ข๐ง๐  using a novel ๐๐ข๐Ÿ๐Ÿ๐ฎ๐ฌ๐ข๐จ๐ง-๐›๐š๐ฌ๐ž๐ framework.

Fine-tuning foundational models often fails due to a lack of dense ground truth, and self-supervised methods struggle with scale ambiguity, sparse depth sensors offer a reliable geometric prior.

This motivated us to develop EndoDDC, a method that robustly generates dense depth maps by fusing RGB images with sparse depth inputs.

๐Ÿง โœจ ๐–๐ก๐š๐ญ ๐ฐ๐ž ๐๐ž๐ฏ๐ž๐ฅ๐จ๐ฉ๐ž๐:

A diffusion-driven depth completion architecture that:

๐Ÿ”น Integrates sparse depth and RGB inputs to overcome the limitations of pure visual estimation.

๐Ÿ”น Utilizes a Multi-scale Feature Extraction and Depth Gradient Fusion module to capture fine-grained surface orientation and local structure.

๐Ÿ”น Optimizes depth maps iteratively using a conditional diffusion model, refining geometry even in regions with weak textures or reflections.

๐ŸŽฏ ๐Š๐ž๐ฒ ๐‘๐ž๐ฌ๐ฎ๐ฅ๐ญ๐ฌ:

โœ… 25.55% and 9.03% improvement in accuracy on the StereoMIS and C3VD dataset compared to SOTA surgical estimators like EndoDAC.

โœ… 7.35% and 5.28% reduction in RMSE on StereoMIS and C3VD compared to the best depth completion baseline (OGNI-DC).

โœ… Outperformed foundational models (DepthAnything-v2) and standard depth completion (Marigold-DC) methods in both accuracy and robustness.

๐Ÿ’ก ๐–๐ก๐ฒ ๐ข๐ญ ๐ฆ๐š๐ญ๐ญ๐ž๐ซ๐ฌ:

This work demonstrates that diffusion models can effectively solve the “sparse-to-dense” challenge in medical imaging. By providing accurate depth completion despite complex lighting and texture conditions, EndoDDC has the potential to significantly enhance autonomous navigation, procedural safety, and spatial awareness in minimally invasive surgery.

๐Ÿ”– #DepthCompletion #DiffusionModel #EndoscopicSurgery #SurgicalNavigation #ICRA #CUHKEngineering #CUHK

No alternative text description for this image
No alternative text description for this image

๐Ÿš€ ICRA 2026: ๐‘ต๐’†๐’–๐’“๐’๐‘ฝ๐‘ณ๐‘จ: ๐‘บ๐’–๐’“๐’ˆ๐’Š๐’„๐’‚๐’ ๐‘บ๐’„๐’†๐’๐’‚๐’“๐’Š๐’-๐‘จ๐’˜๐’‚๐’“๐’† ๐‘ณ๐’†๐’‚๐’“๐’๐’Š๐’๐’ˆ ๐’๐’‡ ๐‘ซ๐’†๐’ƒ๐’–๐’๐’Œ๐’Š๐’๐’ˆ ๐‘บ๐’Œ๐’Š๐’๐’๐’” ๐’Š๐’ ๐‘ฌ๐’๐’…๐’๐’”๐’„๐’๐’‘๐’Š๐’„ ๐‘น๐’๐’ƒ๐’๐’•๐’Š๐’„ ๐‘ต๐’†๐’–๐’“๐’๐’”๐’–๐’“๐’ˆ๐’†๐’“๐’š ๐’—๐’Š๐’‚ ๐‘ฝ๐’Š๐’”๐’Š๐’๐’-๐‘ณ๐’‚๐’๐’ˆ๐’–๐’‚๐’ˆ๐’†-๐‘จ๐’„๐’•๐’Š๐’๐’ ๐‘ด๐’๐’…๐’†๐’ ๐Ÿค–๐Ÿงฒ

We present ๐๐ž๐ฎ๐ซ๐จ-๐•๐‹๐€, an scenario-aware model designed for the motion control of a parallel continuum neurosurgical robot.

Robotic surgery systems have garnered significant attention for their precision and efficiency, yet achieving autonomous tasks in complex neurosurgical environments remains challenging. Although Vision-Language-Action (VLA) models hold great potential, their development is constrained by the scarcity of data from surgical environments and robotic kinematics. To address this issue, this paper proposes NeuroVLA: a VLA model specifically designed for neurosurgical robotic tumor debulking tasks. Through phantom experiments conducted on a flexible parallel continuum robot, we constructed a dataset and decomposed the debulking task into four skill-based instructions. NeuroVLA utilizes a Vision-Language Model (VLM) as its backbone for scene reasoning, enabling the robot to comprehend the surgical scene and its own state. Experimental results demonstrate that after training on 90 debulking segments, NeuroVLA can infer actions based on images, language instructions, and the robotโ€™s state. It achieved average pixel distance errors of 29.10 pixels and 21.55 pixels for the “alignment” and “transfer” skills, respectively, and success rates of 88.89% and 100% for the “grasping” and “release” skills.

๐Ÿง  Technical Framework:

โ—     End-to-End scenario-aware VLA model

โ—     Skill-based scenario infer mechanism

โ—     Debulking task dataset in neurosurgery

๐ŸŽฏ Experimental Results:

โ—     NeuroVLA demonstrates significantly lower pixel distance (PD) errors in the “alignment” and “transfer” skills (29.10 px / 21.55 px), far surpassing the performance of baseline models (such as Octo’s 79.72 px / 65.46 px).

In the “grasping” and “release” skills, NeuroVLA exhibits greater robustness, achieving a grasping success rate of 88.89% and a release success rate of 100%. In contrast, baseline models often misinterpret incomplete forceps closure as task completion, leading to grasping failures.

#ICRA2026

diagram