We present ๐๐ข๐ซ๐ข-๐๐๐ฉ๐ฌ๐ฎ๐ฅ๐, a swallowable kirigami-inspired capsule robot that enables minimally invasive GI biopsyโpushing capsule endoscopy from imaging to tissue sampling.
Wireless capsule endoscopy is comfortable and accessible, but cannot collect biopsy tissue, while histology is still the gold standard. Our work targets safe, depth-controlled, retrievable sampling in a capsule form factor.
๐ง Technical Framework:
โ Kirigami PI skin: flat during locomotion, deploys sharp protrusions when stretched
Thrilled to share our latest work, ๐๐ฎ๐ซ๐ ๐๐ข๐๐๐, the first video-language model specifically designed to address both full and fine-grained surgical video comprehension.
Surgical scene understanding is critical for training and robotic decision-making. While current Multimodal Large Language Models (MLLMs) excel at image analysis, they often overlook the fine-grained temporal reasoning required to capture detailed task execution and specific procedural processes within a surgery. This motivated us to bridge the gap between global video understanding and micro-action analysis.
๐ง โจ What we developed:
A novel framework and resource for surgical video reasoning that includes:
๐น ๐๐ฐ๐จ-๐ฌ๐ญ๐๐ ๐ ๐๐ญ๐๐ ๐๐ ๐จ๐๐ฎ๐ฌ ๐ฆ๐๐๐ก๐๐ง๐ข๐ฌ๐ฆ: The first stage extracts global procedural context, while the second stage performs high-frequency local analysis for fine-grained task execution.
๐น ๐๐ฎ๐ฅ๐ญ๐ข-๐๐ซ๐๐ช๐ฎ๐๐ง๐๐ฒ ๐ ๐ฎ๐ฌ๐ข๐จ๐ง ๐๐ญ๐ญ๐๐ง๐ญ๐ข๐จ๐ง (๐๐ ๐): Effectively integrates low-frequency global features with high-frequency local details to ensure comprehensive scene perception.
๐น ๐๐๐-๐๐๐ ๐๐๐ญ๐๐ฌ๐๐ญ: We constructed a large-scale dataset with over 31,000 video-instruction pairs, featuring hierarchical knowledge representation for enhanced visual reasoning.
๐ฏ Key Results:
โ SurgVidLM significantly outperforms existing models (like Qwen2-VL) in multi-grained surgical video understanding tasks.
โ Capable of inferring anatomical landmarks (e.g., Denonvilliers’ fascia) and providing clinical motivation, moving beyond simple visual description.
โ Demonstrated strong performance on unseen surgical tasks, proving the robustness of our hierarchical training approach.
๐ก Why it matters:
This work shows that by combining global context with localized high-frequency focus, we can significantly reduce “hallucinations” in surgical AI. It provides a pathway toward more intelligent, context-aware surgical assistants that can understand not just what is happening, but how and why specific steps are performed.
๐ฑ Whatโs next?
We are exploring how to extend this multi-grained understanding to real-time intraoperative guidance and integrating it with physical robotic control for autonomous sub-tasks.
We present TMR-VLA, an end-to-end framework designed for the motion control of tri-leg silicone-based soft robots.
Miniature magnetic robots face a hardware bottleneck where the robot body is too small to integrate onboard sensors or power. This creates a gap between actuation and perception, often requiring human experts to manually adjust magnetic fields based on visual feedback. Our work aims to bridge this gap by enabling autonomous control through a multi-modal system.
๐ง Technical Framework:
โ End-to-End Mapping: The policy translates sequential endoscope images and natural language instructions directly into low-level coil voltage commands.
โ Action Adaptor: We utilized an EndoVLA-initialized backbone with an Action LoRA Adaptor that allows the model to autoregressively emit voltage increments.
โ TrilegMR-Motion Dataset: The model was trained on a new dataset containing 15,793 image-action pairs across 60 episodes.
โ Diverse Locomotion: The system controls five motion primitives: squatting, leg-lifting, rotation, forward movement, and recovery.
๐ฏ Experimental Results:
โ Success Rate: TMR-VLA achieved an average success rate of 74% across tested motion types.
โ Performance: The model outperformed general-purpose multimodal models (such as Qwen2.5-VL and LLaVA-1.6) in both instruction interpretation and action execution.
โ Inference Speed: Real-time control was demonstrated at approximately 2 Hz using an NVIDIA RTX 5090 GPU.
๐ก Significance: This study addresses the challenge of autonomous control in untethered soft robots without increasing their structural complexity. It provides a foundational baseline for intelligent navigation in complex in-vivo environments.
Big news! ๐ Our lab is proud to announce that 6 of our latest papers have been accepted! ๐
We are incredibly proud of the teamโs hard work and innovation. To give each project the spotlight it deserves, we will be sharing details about each breakthrough one by one over the coming days.
We are thrilled to share our latest Comment published in #NatureReviewsBioengineering: “๐๐ฟ๐๐ถ๐ณ๐ถ๐ฐ๐ถ๐ฎ๐น ๐ธ๐ถ๐ป๐ฎ๐ฒ๐๐๐ต๐ฒ๐๐ถ๐ฎ ๐ถ๐ป ๐ฎ๐๐๐ผ๐ป๐ผ๐บ๐ผ๐๐ ๐ฟ๐ผ๐ฏ๐ผ๐๐ถ๐ฐ ๐๐๐ฟ๐ด๐ฒ๐ฟ๐”.
Current autonomous surgical robots are ๐ต๐ฒ๐ฎ๐๐ถ๐น๐ ๐๐ถ๐๐ถ๐ผ๐ป-๐ฐ๐ฒ๐ป๐๐ฟ๐ถ๐ฐ. While they ๐ฐ๐ฎ๐ป “๐๐ฒ๐ฒ” anatomy, they ๐น๐ฎ๐ฐ๐ธ ๐๐ต๐ฒ ๐ถ๐ป๐๐ฟ๐ถ๐ป๐๐ถ๐ฐ ๐ฎ๐ฏ๐ถ๐น๐ถ๐๐ ๐๐ผ “๐ณ๐ฒ๐ฒ๐น” tissue interactionsโ๐ฎ ๐ฐ๐ฟ๐๐ฐ๐ถ๐ฎ๐น ๐๐ธ๐ถ๐น๐น ๐๐ต๐ฎ๐ ๐ต๐๐บ๐ฎ๐ป ๐๐๐ฟ๐ด๐ฒ๐ผ๐ป๐ ๐ฟ๐ฒ๐น๐ ๐ผ๐ป ๐ณ๐ผ๐ฟ ๐๐ฎ๐ณ๐ฒ๐๐ ๐ฎ๐ป๐ฑ ๐ฑ๐ฒ๐ ๐๐ฒ๐ฟ๐ถ๐๐.
In this article, we propose a hierarchical framework for ๐๐ฟ๐๐ถ๐ณ๐ถ๐ฐ๐ถ๐ฎ๐น ๐๐ถ๐ป๐ฎ๐ฒ๐๐๐ต๐ฒ๐๐ถ๐ฎ to bridge this gap:
1. ๐๐ง๐ต๐ฒ ๐ฃ๐ต๐๐๐ถ๐ฐ๐ฎ๐น ๐๐ฒ๐๐ฒ๐น: Integrating proprioception and exteroception for high-res physical sensing.
2. ๐ฌ๐ง๐ต๐ฒ ๐๐น๐ด๐ผ๐ฟ๐ถ๐๐ต๐บ๐ถ๐ฐ ๐๐ฒ๐๐ฒ๐น: Moving from raw signal processing to semantic understanding of contact.
We believe the future of autonomous surgery lies in systems that can synergistically fuse vision and kinaesthesia to not just see, but truly feel, think, and act.
๐ Read the ๐ณ๐๐น๐น ๐ฝ๐ฎ๐ฝ๐ฒ๐ฟ here: [https://lnkd.in/gqTEpYjs]
Thrilled to share our latest The International Journal of Robotics Research (IJRR) work on enabling ๐ด๐ฟ๐ฎ๐๐ถ๐๐โ๐ฎ๐๐ฎ๐ฟ๐ฒ ๐ฐ๐ผ๐ป๐๐ฟ๐ผ๐น for ๐ฝ๐ผ๐ฟ๐๐ฎ๐ฏ๐น๐ฒ ๐ฐ๐ฎ๐ฏ๐น๐ฒโ๐ฑ๐ฟ๐ถ๐๐ฒ๐ป ๐๐ผ๐ณ๐ ๐๐น๐ฒ๐ป๐ฑ๐ฒ๐ฟ ๐ฟ๐ผ๐ฏ๐ผ๐๐ โ using ๐ผ๐ป๐น๐ ๐ฎ ๐๐ถ๐ป๐ด๐น๐ฒ ๐๐ ๐จ and a powerful ๐ฟ๐ผ๐ฏ๐ผ๐ฝ๐ต๐๐๐ถ๐ฐ๐ฎ๐น simulationโdriven framework.
Soft robots are lightweight and flexible, but their high aspect ratios make them ๐ฆ๐น๐ต๐ณ๐ฆ๐ฎ๐ฆ๐ญ๐บ sensitive to gravity, causing passive deformation that traditional kinematics just canโt handle. This motivated us to rethink how soft robots can ๐ด๐ฆ๐ฏ๐ด๐ฆ and ๐ค๐ฐ๐ฎ๐ฑ๐ฆ๐ฏ๐ด๐ข๐ต๐ฆ for gravity โ without bulky sensors or complex hardware.
This work shows that soft robots can maintain ๐๐๐ฎ๐ฏ๐น๐ฒ, ๐ฐ๐ผ๐ป๐๐ถ๐๐๐ฒ๐ป๐ ๐ฐ๐ผ๐ป๐ณ๐ถ๐ด๐๐ฟ๐ฎ๐๐ถ๐ผ๐ป๐ ๐๐ป๐ฑ๐ฒ๐ฟ ๐ฐ๐ต๐ฎ๐ป๐ด๐ถ๐ป๐ด ๐ด๐ฟ๐ฎ๐๐ถ๐๐ by integrating ๐๐ถ๐ฟ๐๐๐ฎ๐น ๐๐ฒ๐ป๐๐ถ๐ป๐ด + ๐๐ถ๐บ๐๐น๐ฎ๐๐ถ๐ผ๐ปโ๐ฑ๐ฟ๐ถ๐๐ฒ๐ป ๐ถ๐ป๐๐ฒ๐ฟ๐๐ฒ ๐ฐ๐ผ๐บ๐ฝ๐๐๐ฎ๐๐ถ๐ผ๐ป. It reduces reliance on physical sensors and opens a pathway toward ๐๐ฐ๐ฎ๐น๐ฎ๐ฏ๐น๐ฒ, ๐ด๐ฒ๐ป๐ฒ๐ฟ๐ฎ๐น๐ถ๐๐ฎ๐ฏ๐น๐ฒ ๐ด๐ฟ๐ฎ๐๐ถ๐๐โ๐ฎ๐๐ฎ๐ฟ๐ฒ ๐๐ผ๐ณ๐ ๐ฟ๐ผ๐ฏ๐ผ๐๐.
๐ฑ ๐ช๐ต๐ฎ๐โ๐ ๐ป๐ฒ๐ ๐?
Weโre exploring how to extend this architecture to virtualizable external force sensing and richer environmental interactions.
Endoscopic submucosal dissection (ESD) is a key technique for early GI cancer treatment, requiring high dexterity and precision.
๐ ๐ช๐ต๐ฎ๐ ๐๐ฒ ๐ฑ๐ถ๐ฑ:
We developed the first ๐ต๐ฒ๐๐ฒ๐ฟ๐ผ๐ด๐ฒ๐ป๐ฒ๐ผ๐๐ ๐ณ๐น๐ฒ๐ ๐ถ๐ฏ๐น๐ฒ ๐บ๐ฎ๐ป๐ถ๐ฝ๐๐น๐ฎ๐๐ผ๐ฟ๐ (๐๐๐ ๐) for bimanual ESD, integrating:
โ Kinematic modeling using DenavitโHartenberg & Cosserat rod methods
โ Workspace & dexterity analysis via simulation
โ Validation through 16 ex vivo ESD tests
๐ก This work demonstrates a novel strategy for surgical roboticsโleveraging heterogeneous structures to enhance flexibility, stiffness, and accuracy in minimally invasive procedures.
๐ Kudos to our amazing team and collaborators from CUHK (Prof. Huxin Gao, Tao Zhang, Prof. Hongliang Ren), Qilu Hospital (Xiaoxiao Yang, Prof. ๅทฆ็งไธฝ, Prof. Yanqing Li), Southern University of Science and Technology (Xiao Xiao, Prof. Qinghu Meng), and Beijing Institute of Technology (Prof. Changsheng Li)!
Navigating flexible robotic endoscopes in the dynamic, deformable stomach environment is a grand challenge. Our proposed Contact-Aided Navigation (CAN) strategy, powered by deep reinforcement learning and force-feedback, achieved:
โข 100% success rate in both static and dynamic simulated stomach environments
โข Average navigation error of just ๐ญ.๐ฒ ๐บ๐บ
โข Robust generalization even under strong external disturbances
This work highlights how ๐ฒ๐บ๐ฏ๐ผ๐ฑ๐ถ๐ฒ๐ฑ ๐๐ ๐ฎ๐ป๐ฑ ๐ฏ๐ถ๐ผ๐บ๐ฒ๐ฐ๐ต๐ฎ๐ป๐ถ๐ฐ๐-๐ถ๐ป๐๐ฝ๐ถ๐ฟ๐ฒ๐ฑ ๐๐๐ฟ๐ฎ๐๐ฒ๐ด๐ถ๐ฒ๐ can transform surgical robotics, enabling safer and more precise navigation in complex clinical environments.
Check the paper at https://lnkd.in/g6KgZTdD
๐ Huge thanks to the team, collaborators, and the broader robotics community for the support and inspiration.
This work introduces a comprehensive dataset designed to advance AI-driven surgical robotics and medical imaging. By capturing detailed ๐ฎ๐ป๐ฎ๐๐ผ๐บ๐ถ๐ฐ๐ฎ๐น ๐น๐ฎ๐ป๐ฑ๐บ๐ฎ๐ฟ๐ธ๐ ๐ผ๐ณ ๐๐ต๐ฒ ๐๐ฝ๐ฝ๐ฒ๐ฟ ๐ฎ๐ถ๐ฟ๐๐ฎ๐, we aim to support safer, more accurate ๐ฏ๐ฟ๐ผ๐ป๐ฐ๐ต๐ผ๐๐ฐ๐ผ๐ฝ๐ ๐ฎ๐ป๐ฑ ๐ถ๐ป๐๐๐ฏ๐ฎ๐๐ถ๐ผ๐ป procedures โ paving the way for improved patient outcomes and robust benchmarking in clinical AI.
๐ ๐๐ถ๐ด๐ต๐น๐ถ๐ด๐ต๐๐:
– First-of-its-kind dataset focused on airway anatomical landmarks
– Enables benchmarking for automated navigation and intubation tasks
– Openly available to foster collaboration across robotics, AI, and healthcare communities
We hope this resource will accelerate innovation in ๐ฒ๐บ๐ฏ๐ผ๐ฑ๐ถ๐ฒ๐ฑ ๐ถ๐ป๐๐ฒ๐น๐น๐ถ๐ด๐ฒ๐ป๐ฐ๐ฒ ๐ณ๐ผ๐ฟ ๐ต๐ฒ๐ฎ๐น๐๐ต๐ฐ๐ฎ๐ฟ๐ฒ and inspire new interdisciplinary collaborations.
๐ Read the full paper: https://rdcu.be/eS0d5
Grateful to all co-authors and collaborators from The Chinese University of Hong Kong (Ruoyi Hao, Zhiqing Tang, Catherine Po Ling Chan, Jason Ying Kuen Chan, Prof. Hongliang Ren), Hubei University of Technology (Zhang Yang), Huazhong University of Science and Technology (Yang Zhou), National University of Singapore (Lalithkumar Seenivasan), and Singapore General Hospital (Shuhui
Xu, Neville Wei Yang Teo, Kaijun Tay, Vanessa Yee Jueen Tan, Jiun Fong Thong, Kimberley Liqin Kiong, Shaun Loh, Song Tar Toh, and Prof. Chwee Ming Lim), for making this possible. Excited to see how others build upon this foundation!
๐กThis work introduces ๐ฉ๐-๐ฆ๐๐ฟ๐ด๐ฃ๐ง, the first large-scale multimodal dataset that ๐ช๐ฏ๐ต๐ฆ๐จ๐ณ๐ข๐ต๐ฆ๐ด ๐ท๐ช๐ด๐ถ๐ข๐ญ ๐ต๐ณ๐ข๐ซ๐ฆ๐ค๐ต๐ฐ๐ณ๐ช๐ฆ๐ด ๐ธ๐ช๐ต๐ฉ ๐ด๐ฆ๐ฎ๐ข๐ฏ๐ต๐ช๐ค ๐ฑ๐ฐ๐ช๐ฏ๐ต ๐ด๐ต๐ข๐ต๐ถ๐ด ๐ฅ๐ฆ๐ด๐ค๐ณ๐ช๐ฑ๐ต๐ช๐ฐ๐ฏ๐ด in surgical environments.
๐Alongside the dataset, we propose ๐ง๐-๐ฆ๐๐ฟ๐ด๐ฃ๐ง, a text-guided point tracking method that consistently outperforms vision-only approaches, especially under challenging intraoperative conditions such as smoke, occlusion, and tissue deformation.
๐ We are deeply grateful to all coauthors and especially our clinical collaborators at Shenzhen Peopleโs Hospital for their invaluable contributions. Looking forward to engaging with the community at AAAI in Singapore and advancing the conversation on multimodal surgical AI!
๐ Excited to share that the 3rd C4SR+ Workshop: ๐พ๐ค๐ฃ๐ฉ๐๐ฃ๐ช๐ช๐ข, ๐พ๐ค๐ข๐ฅ๐ก๐๐๐ฃ๐ฉ, ๐พ๐ค๐ค๐ฅ๐๐ง๐๐ฉ๐๐ซ๐, ๐พ๐ค๐๐ฃ๐๐ฉ๐๐ซ๐ ๐๐ช๐ง๐๐๐๐๐ก ๐๐ค๐๐ค๐ฉ๐๐ ๐๐ฎ๐จ๐ฉ๐๐ข๐จ ๐๐ฃ ๐ฉ๐๐ ๐๐ข๐๐ค๐๐๐๐ ๐ผ๐ ๐๐ง๐ took place during #IROS2025 in Hangzhou, China.
This yearโs workshop attracted ๐ผ๐๐ฒ๐ฟ ๐ผ๐ป๐ฒ ๐ต๐๐ป๐ฑ๐ฟ๐ฒ๐ฑ ๐ฝ๐ฎ๐ฟ๐๐ถ๐ฐ๐ถ๐ฝ๐ฎ๐ป๐๐ ๐ณ๐ฟ๐ผ๐บ ๐ฎ๐ฐ๐ฟ๐ผ๐๐ ๐๐ต๐ฒ ๐ด๐น๐ผ๐ฏ๐ฒ โ a fantastic turnout that reflects the growing momentum in ๐๐๐ฟ๐ด๐ถ๐ฐ๐ฎ๐น ๐ฟ๐ผ๐ฏ๐ผ๐๐ถ๐ฐ๐ ๐ฎ๐ป๐ฑ ๐ฒ๐บ๐ฏ๐ผ๐ฑ๐ถ๐ฒ๐ฑ ๐๐.
๐ค Distinguished Speakers
We were honored to host leading experts who shared their groundbreaking research and perspectives, including Prof. Nassir Navab from Technical University of Munich, Prof. Leonardo Mattos from Italian Institute of Technology, Prof. Mingchuan Zhou from Zhejiang University, Prof. Dandan Zhang from Imperial College London, Prof. Yunjie Yang from University of Edinburgh, and Prof. Guoying Gu from Shanghai Jiao Tong University.
๐ Key Themes Discussed
โข Embodied AI in surgery and intelligent operating rooms
โข Soft & continuum robotics for minimally invasive procedures
โข Humanโrobot collaboration in clinical practice
โข Cognitive surgical systems and decision-making
๐ Workshop Contributions
– 9 oral and 9 poster paper presentations from emerging researchers
– Best Paper & Best Presentation Awards recognizing outstanding contributions
๐ A heartfelt thank you to all speakers, participants, and organizers who made this workshop such a success. The discussions and collaborations will continue to shape the future of surgical robotics.
๐ Learn more about the workshop and its highlights on the official page: https://lnkd.in/gswzMFAy