💡 Academia - Funding Challenges
This article discusses the challenges faced by academics in securing research funding, focusing on the limitations imposed by the preference for "safe" research proposals.
Key Points:
• Difficulty securing research grants due to limited funding.
• Predominance of "safe" proposals that continue existing research.
• Limited opportunities for pursuing novel research topics.
🔗 Resources:
• Anshul Heaven ↗ - Academic perspective
• Anon Rand ↗ - Academic perspective
• SteveStuWill ↗ - Related Image
Image
🤖 AI Models - World Models and Perfect Prediction
This article explores the concept of an AI model that can make perfect predictions while possessing a flawed understanding of the underlying world. The discussion centers on a research paper presented at ICML which demonstrates this phenomenon using a solar system model.
Key Points:
• AI models can achieve perfect prediction without accurate world models.
• A transformer model accurately predicted planetary orbits despite lacking a complete understanding of gravitational laws.
• This highlights the potential disconnect between predictive accuracy and true comprehension.
Image
🔗 Resources:
• Kyle Cranmer ↗ - Research contribution
• Keyon V ↗ - ICML paper details
🚀 AI Acquisitions - Meta Acquires PlayAI
This article reports on Meta's acquisition of the voice technology startup, PlayAI. The acquisition brings the entire PlayAI team to Meta.
Key Points:
• Meta expands its AI capabilities.
• Acquisition of PlayAI enhances voice interaction technologies.
• PlayAI team joins Meta next week.
🔗 Resources:
• Munshi PremChnd ↗ - Acquisition announcement
• Meta Acquisition News ↗ - More details
🤖 Data Filtering - Low-Motion Data
This article examines the counterintuitive effects of low-motion data filtering on model performance, specifically in the context of fine-tuning and language adherence.
Key Points:
• Low-motion data filtering is beneficial for task-specific fine-tuning.
• Counterintuitively, it negatively impacts model generalization and language adherence.
• Understanding the impact requires a case by case examination.
Image
🔗 Resources:
• Shreyas Gite ↗ - Discussion on data filtering
• M Zubair Irshad ↗ - Related insights
🤖 Large Behavior Models (LBMs) - Toyota Research Institute (TRI)
This article announces the release of the Toyota Research Institute's first Large Behavior Models (LBMs), highlighting the contribution of multi-task policy learning and post-training efforts.
Key Points:
• TRI releases its first Large Behavior Models (LBMs).
• LBMs leverage multi-task policy learning and post-training techniques.
• Focus on researching LBM applications.
🔗 Resources:
• M Zubair Irshad ↗ - Development details
💡 Reinforcement Learning (RL) - Overview and Resources
This article provides a pointer to resources for understanding the current state of reinforcement learning (RL), including GRPO, agents, VERL, and asynchronous GRPO/DAPO.
Key Points:
• A small talk provides an overview of RL.
• Covers GRPO, agents, VERL, asynchronous GRPO, and DAPO.
• Suitable for individuals new to RL.
🔗 Resources:
• Rohan Awhad ↗ - Small talk video
🤖 Imitation Learning - Action Chunking in RL
This article describes how action chunking, successful in imitation learning, can be extended to reinforcement learning to improve exploration and sample efficiency.
Key Points:
• Action chunking enhances imitation learning.
• Extending its benefits to RL improves exploration and sample efficiency.
• Simple implementation.
Image
🔗 Resources:
• Colin Qiyang Li ↗ - Research paper
• Qiyang Li ↗ - Discussion on action chunking
💡 Legal - Australian Protests and Legal Consequences
This article discusses the author's experience with Australian authorities, focusing on a $24,000 fine for holding a blank sign and contrasting it with the lack of police action following a car break-in.
Key Points:
• $24,000 fine for holding a blank sign outside the Brisbane Chinese Consulate.
• Lack of police action regarding a car break-in.
• Disparity in legal response.
Image
🔗 Resources:
• Drew Pavlou ↗ - Account of events
🤖 AI Evaluation - METR Evals Study
This article shares reflections on the author's participation in the METR Evals study, focusing on how AI speedup gains can be offset by factors such as prompt monitoring and contextual interruptions.
Key Points:
• AI speedup gains can be reduced by distractions during prompt execution.
• Author's experience in the METR Evals study highlighted these tradeoffs.
• Careful consideration of contextual factors is crucial for accurate speedup assessment.
🔗 Resources:
• Ruben Bloom ↗ - METR Evals Study reflections
• METR Evals ↗ - Study information
🤖 AI Next-Move Prediction - Othello Experiment
This article describes an experiment where an Othello next-move predictor was fine-tuned to reconstruct the game board from its internal state. The results show that while the reconstructed boards were often inaccurate, the predicted next moves were correct.
Key Points:
• Finetuning an Othello next-move predictor to reconstruct boards.
• Reconstructed boards were often incorrect but next moves were accurate.
• Suggests that next-token prediction may be a relatively easy task.
Image
🔗 Resources:
• Rajiv Movva ↗ - Othello experiment details
• Keyon V ↗ - Related image
⭐️ Support
If you liked reading this report, please star ⭐️ this repository and follow me on Github ↗, 𝕏 (previously known as Twitter) ↗ to help others discover these resources and regular updates.