π€ Robotics - Robotics and Code-Based World Models
Postdoctoral researcher position available at Princeton University, working at the intersection of robotics and code-based world models. The position is part of a research group led by @basisorg and involves collaboration with a team of researchers and engineers. The ideal candidate should have expertise in robotics, planning, and abstractions, as well as experience working with real robots.
Key Points:
Robotics and Code-Based World Models: The position involves working on the intersection of robotics and code-based world models, which is a rapidly growing area of research. The goal is to develop more efficient and effective ways of modeling and interacting with the physical world.
Abstractions and Planning: The ideal candidate should have expertise in abstractions and planning, as well as experience working with real robots. This includes understanding how to design and implement algorithms that can efficiently and effectively interact with the physical world.
Collaboration and Teamwork: The position involves collaboration with a team of researchers and engineers, so the ideal candidate should have strong communication and teamwork skills.
π Resources:
- Original post β
- Original source
- BasisOrg (@basisorg)
- Princeton University
π Robotics - ROSCON and OpenRoboticsOrg
If you see someone with a https:// prefix.dev badge at ROSCON, they are likely from the OpenRoboticsOrg community. Come say hi, we're friendly!
Key Points:
ROSCON and OpenRoboticsOrg: ROSCON is a conference for robotics researchers and engineers, and OpenRoboticsOrg is a community of robotics researchers and engineers. The badge is a way of identifying members of the community.
Community and Collaboration: The OpenRoboticsOrg community is a group of researchers and engineers who are passionate about robotics and are working together to advance the field.
Friendliness: The community is friendly and welcoming, so if you see someone with the badge, don't be afraid to come say hi.
π Resources:
- Original post β
- Original source
- TobiasRobotics (@TobiasRobotics)
- OpenRoboticsOrg (@OpenRoboticsOrg)
πΉ Video Editing - Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing
Our paper "Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing" was accepted to the AACL-IJCNLP 2026 Demo Track! Crayotter is an open-source multimodal multi-agent system for prompt-driven long-form video editing, where every step β retrieval, analysis, and generation β is performed by a separate agent.
Key Points:
Crayotter: Crayotter is an open-source multimodal multi-agent system for prompt-driven long-form video editing. It is designed to be highly flexible and customizable, allowing users to easily create and modify workflows.
Multi-Agent Workflows: The system uses a multi-agent approach, where each step of the video editing process is performed by a separate agent. This allows for highly efficient and scalable workflows.
Long-Form Video Editing: Crayotter is designed specifically for long-form video editing, where the video is typically several hours long. It is capable of handling large amounts of data and can perform complex editing tasks.
π Resources:
- Original post β
- Original source
- Liruizhe94 (@liruizhe94)
- Chenyang_Lyu (@Chenyang_Lyu)
π€ World Models - Jev: A Visual Explanation
For anyone curious how Jev works, I made a visual explanation using @claudeai. This is based on the Qwen2.5-RLCD model which @harshagundal released on @huggingface. The idea is to replace autoregressive LLM generation by a single Transformer decoder (of a pre-trained LLM), which is more efficient and scalable.
Key Points:
Jev: Jev is a visual explanation of how the Jev model works. It is designed to be highly intuitive and easy to understand, even for those without a strong background in machine learning.
Qwen2.5-RLCD Model: The Jev model is based on the Qwen2.5-RLCD model, which is a pre-trained language model that has been released on the Hugging Face model hub.
Transformer Decoder: The Jev model uses a single Transformer decoder, which is more efficient and scalable than traditional autoregressive LLM generation.
π Resources:
- Original post β
- Original source
- Raitweet10 (@Raitweet10)
- NielsRogge (@NielsRogge)
- ClaudeAI (@claudeai)
- Harshagundal (@harshagundal)
- Hugging Face (@huggingface)
πΈ World Models - RGB vs DINO Features
World-action models typically imagine the future in RGB β but are pixels really the right representation for robotics? Our bet is no: RGB spends capacity on fine-grained details and variation that are often irrelevant for robot policies. DINO features, point tracks, and depth maps are more efficient and effective.
Key Points:
RGB vs DINO Features: RGB is a common representation used in world-action models, but it may not be the most effective choice for robotics. DINO features, point tracks, and depth maps are more efficient and effective.
Fine-Grained Details: RGB spends capacity on fine-grained details and variation that are often irrelevant for robot policies. This can lead to inefficiencies and decreased performance.
DINO Features: DINO features, point tracks, and depth maps are more efficient and effective because they capture the essential information needed for robot policies.
π Resources:
- Original post β
- Original source
- GeYan_21 (@GeYan_21)
- Adamjhung (@Adamjhung)
πΊ AI - Siri and Apple
New John Ternus interview as Apple CEO! Ternus says Siri βis not designed to be your friendβ and that Apple wants to βhelp people be a better version of themselves.β Very interesting and candid conversation.
Key Points:
Siri and Apple: Siri is a virtual assistant developed by Apple, and it is not designed to be a friend. Its purpose is to provide helpful information and assistance to users.
Apple's Goals: Apple's goal is to help people be a better version of themselves, and Siri is a tool that can help achieve this goal.
Candid Conversation: The interview with John Ternus is a candid and interesting conversation that provides insight into Apple's goals and values.
π Resources:
- Original post β
- Original source
- JackyWangAI (@JackyWangAI)
- Samifathi (@samifathi)
π» Code - Disconnect Between Dev and Code
It creates a disconnect between the dev and the code. This is fine for small projects that never see production, and even on larger projects at first, but over time (a couple months) performance degrades and the dev cannot fix the issue, as they do not understand the code.
Key Points:
Disconnect Between Dev and Code: The disconnect between the dev and the code can lead to performance degradation and decreased ability to fix issues.
Small Projects: For small projects, the disconnect may not be a significant issue, but for larger projects, it can become a major problem.
Code Understanding: The dev may not understand the code, which can make it difficult to fix issues and maintain the project.
π Resources:
- Original post β
- Original source
- The Root User (@The_Root_User)
πΈ Computer Vision - RIGOR: Rig-Informed Geometry for Omnidirectional Reconstruction
Huang et al., "RIGOR: Rig-Informed Geometry for Omnidirectional Reconstruction" Here's how you can use existing feed-forward models with 360 cameras. A tightly knit system that treats 360 images as four virtual "rigs" that must agree.
Key Points:
RIGOR: RIGOR is a system for omnidirectional reconstruction that uses existing feed-forward models with 360 cameras.
360 Cameras: 360 cameras are used to capture images from all directions, and RIGOR treats these images as four virtual "rigs" that must agree.
Tightly Knit System: The system is tightly knit, meaning that it is designed to work together seamlessly to produce high-quality results.
π Resources:
- Original post β
- Original source
- Kwangmoo_yi (@kwangmoo_yi)
π€ Robotics - MARA Project
Very excited to ramp up our work working with Tom! We're looking for an exceptional postdoctoral researcher (details below) to work jointly with Tom and Basis on MARA project--our effort to build robotic agents that actively learn models of the physical world.
Key Points:
MARA Project: The MARA project is an effort to build robotic agents that actively learn models of the physical world.
Postdoctoral Researcher: The project is looking for an exceptional postdoctoral researcher to work jointly with Tom and Basis.
Robotic Agents: The project involves building robotic agents that can learn and adapt to their environment.
π Resources:
- Original post β
- Original source
- LintonVision (@LintonVision)
- ZennaTavares (@ZennaTavares)
π AI - Scaling Compute
So they say "three things we scaled". Let me translate: 1. "compute" - yep, that's compute. 2. "environments and harnesses" - actually, also compute. 3. "and grader compute" - you guessed it, that's also compute. joke aside, pretty cool to see their public live dashboard.
Key Points:
Scaling Compute: The post discusses scaling compute, which is a critical aspect of AI development.
Compute: Compute is a key component of AI development, and scaling it is essential for achieving high performance.
Public Live Dashboard: The post mentions a public live dashboard, which provides insight into the scaling process.
π Resources:
- Original post β
- Original source
- Giffmana (@giffmana)