Imitation learning has seen great success, but IL policies still struggle with OOD observations such as changes in camera pose. We designed a 3D backbone, Adapt3R, that can combine with your favorite Imitation Learning algorithm to enable zero-shot generalization to unseen embodiments and camera viewpoints! Learning 3D representations is hard without 3D data 💡 The key idea is to use a 2D foundation model to extract semantic features, and use 3D information to localize those features in a canonical 3D space without extracting any semantic information from the 3D data. Adapt3r unprojects 2D features into a point cloud, transforms them into the end effector’s coordinate frame, and uses attention pooling to condense them into a single conditioning vector for IL. Notice that Adapt3R attends to the same points before and after the camera change! 2D features lifted into 3D are an effective representation for this scenario and Adapt3R makes good use of them! So, what did we observe empirically? - Adapt3R is just as proficient as RGB based baselines in case of in distribution evaluations. - But, Adapt3R is very good at embodiment transfers - Most importantly, Adapt3R handles viewpoint changes at test time! No More fixing the camera to match training distribution! Overall this means Adapt3R provides 3D representations as a drop in replacement for 2D RGB baselines. this was led by Albert Wilcox with help from Mohamed Ghanem, Masoud Moghani, Pierre Barroso, Benjamin Joffe and Animesh Garg Check out more at and play with the code in your next robot learning project. 🌐 Website: https://lnkd.in/dWneBJ5d 📄 Paper: https://lnkd.in/dvcbA_22 🖥️ Code: https://lnkd.in/dFKYjym7
Building Dynamic Camera Perspectives in Robotics
Explore top LinkedIn content from expert professionals.
Summary
Building dynamic camera perspectives in robotics means enabling robots to adapt their visual viewpoint in real time, so they can see and understand their environment from multiple angles, rather than relying on a fixed camera view. This flexibility helps robots handle complex tasks, avoid blind spots, and better respond to changing situations.
- Integrate multiple viewpoints: Use several cameras or moveable camera mounts to let robots gather detailed visual data from different positions, improving accuracy and coverage.
- Encourage active observation: Build robot systems that can move their cameras or heads to search for objects and navigate obstacles, rather than sticking to a static setup.
- Adapt to new scenarios: Train robots to adjust their camera perspectives and interpret new viewpoints without starting the learning process over, so they stay reliable even when conditions change.
-
-
In the world of #binpicking and Vision-Guided Robotics (VGR), a single perspective isn't always enough. Traditional single-scan setups often struggle with "blind spots," but the Photoneo, now part of Zebra Technologies' MultiView technology, is changing the game, engineered to bypass the physical limitations of a single-scan perspective. 🔴 Why MultiView? 🔵 By combining 3D data from multiple viewpoints into a single, high-resolution 3D point cloud, we eliminate the hurdles that stop most automation lines: ✅ Occlusions: See what others miss by looking around obstacles. ✅ Complex Surfaces: Perfect for thin, reflective, or irregularly shaped parts. ✅ Large Objects: Achieve full coverage of oversized components without sacrificing detail. 💡 Two Ways to Deploy: 👉 Static: Utilizing multiple PhoXi 3D Scanners for rock-solid reliability in fixed environments: https://lnkd.in/eaNPAdr6 👉In Motion: Using the MotionCam-3D in a hand-eye configuration to capture data dynamically while the robot moves: https://lnkd.in/eT5B7Bbz The result? An unmatched localization success rate, even in the most complex industrial scenarios. If you want to increase your throughput and decrease error rates, it's time to look at the bigger picture. Stop settling for one point of view. See everything. Universal Robots #digitaltwin #physicalAI #3DVision #Robotics #Automation #BinPicking #Photoneo #ZebraTechnologies #SmartManufacturing #VGR #Innovation #automotive #Tier1 #3Dscanning #3Dcamera #machinetending
-
Vision language action models finally get to look around on their own. Through the paper SaPaVe, researchers allowed robots to actively move their head-mounted camera to find objects completely out of view. Instead of the usual fixed near-optimal camera setup, SaPaVe adds a 2-DoF active head (pitch + yaw) that the model learns to control semantically from language instructions alone. They accomplished this using: - A two separate action spaces, splitting the final model output into two MLP heads, one for camera movement (2-DoF) and one for manipulation (26-DoF), so the two skills don't interfere during training - ActiveViewPose-200K, a synthetically generated dataset of 200k image-instruction-camera movement triplets built cheaply in simulation A two-stage training strategy where the first train only the camera head on ActiveViewPos and the second trains the manipulation head on real robot data while mixing in camera examples to prevent forgetting - Spatial Knowledge Injection via MapAnything, feeding 3D geometry (depth, camera poses) into the action decoder so the policy stays stable even with viewpoint changes (Day 2 of highlighting interesting CVPR 2026 papers about VLAs)
-
A robot that follows "camera + instruction" falls apart the moment you move the camera. This paper fixes that — with no retraining. The problem: a vision-language-action policy bakes one fixed setup — camera angle, robot shape — into its weights. Move the camera to an angle it never trained on and real-robot success collapses from 68% to 17%. From a single frame it has no way to even tell which setup it's in. In-Context World Modeling (ICWM) gives the robot a habit people already have — wiggle an unfamiliar joystick a few times before you trust it. 3 ideas: 1. Probe first. Before the task, the robot makes a few random, goal-free test moves and records how the camera view responds. 2. Infer, don't retrain. The same network — no new weights — reads those clips to work out the current camera + robot setup, then acts on it. 3. It generalizes. The same trick handles new viewpoints, clutter, even a physically lengthened arm — +71% on average on a real UR5e arm over a multi-view-trained baseline. I made a short visual breakdown of the whole thing. Swipe through. 👇 --- paper - https://lnkd.in/ggXM_3Ks full writeup - https://lnkd.in/gfPUfvQH Write AI research papers 👉 https://lnkd.in/dQ2qXqaw
Explore categories
- Hospitality & Tourism
- Productivity
- Finance
- Soft Skills & Emotional Intelligence
- Project Management
- Education
- Technology
- Leadership
- Ecommerce
- User Experience
- Recruitment & HR
- Customer Experience
- Real Estate
- Marketing
- Sales
- Retail & Merchandising
- Science
- Supply Chain Management
- Future Of Work
- Consulting
- Writing
- Economics
- Employee Experience
- Healthcare
- Workplace Trends
- Fundraising
- Networking
- Corporate Social Responsibility
- Negotiation
- Communication
- Engineering
- Career
- Business Strategy
- Change Management
- Organizational Culture
- Design
- Innovation
- Event Planning
- Training & Development