🤖 GPU Performance - RDNA4 vs. RDNA3 and Blackwell vs. Ada Lovelace
This article compares the Instruction Per Clock (IPC) performance of AMD's RDNA 4 and RDNA 3 architectures, and Nvidia's Blackwell and Ada Lovelace architectures, based on a ComputerBase analysis. The focus is on rasterizer and raytracing performance differences.
Key Points:
• RDNA 4 shows noticeable improvements in rasterization and, especially, raytracing performance.
• Blackwell architecture shows limited performance gains compared to Ada Lovelace.
• The ComputerBase analysis provides a detailed comparison of these GPU architectures.
🔗 Resources:
• ComputerBase Article ↗ - Detailed GPU performance comparison
Image
Image
Image
Image
🤖 Robotics - RSS Conference and Recruitment
This article announces the author's attendance at the RSS conference and details their recruitment needs for a founding scientist and new models/tools for their robot.
Key Points:
• Seeking a founding scientist to join their core team.
• Looking for new models and tools to enhance their robot for customers.
• Open to discussing robotic insights at the RSS conference.
🤖 Reward Function Design - Avoiding Reward Hacking in Chemistry
This article discusses a 3.5k-word essay on designing reward functions for chemistry applications while mitigating reward hacking. The essay details methods used to prevent reward hacking during the training of a scientific reasoning model called ether0.
Key Points:
• Provides comprehensive strategies for designing reward functions in chemistry.
• Explores various reward hacking avoidance techniques.
• Details experiences in training the ether0 scientific reasoning model.
🔗 Resources:
Image
💡 Personal Reflection - Mountain Trip and Mental Well-being
This article describes a personal experience of a weekend trip to the mountains, emphasizing the restorative effect of nature on mental well-being.
Key Points:
• Highlights the positive impact of nature on mental well-being.
• Shares a personal experience of a mountain trip.
🔗 Resources:
Image
🤖 Reinforcement Learning - Asynchronous RL for Improved GPU Utilization
This article contrasts synchronous and asynchronous reinforcement learning (RL), highlighting the advantages of asynchronous RL for maximizing GPU utilization. It focuses on AReaL-boba²'s asynchronous approach.
Key Points:
• Synchronous RL wastes GPU resources due to batch waiting times.
• Asynchronous RL in AReaL-boba² improves GPU utilization by decoupling generation and training.
• Asynchronous RL represents a significant advancement in reinforcement learning.
🔗 Resources:
Image
🤖 NLP - Personalizing AI Assistants using User Feedback (SynthesizeMe)
This article introduces SynthesizeMe, a new approach for personalizing AI assistants by leveraging user feedback. The method uses both implicit and explicit feedback to tailor the AI assistant to individual users.
Key Points:
• SynthesizeMe personalizes AI assistants based on user feedback.
• It utilizes both implicit and explicit feedback mechanisms.
• The approach aims to create more personalized and effective AI interactions.
🔗 Resources:
Image
🤖 Large Language Model Fine-tuning - Arithmo-Mistral-7B Results
This article suggests adding the results of Arithmo-Mistral-7B (4-bit QLoRA on 1x4090) to a research paper. It highlights the model's achievement as the first to demonstrate improved results from fine-tuning on the Mistral-7B model.
Key Points:
• Arithmo-Mistral-7B shows improved results from fine-tuning on Mistral-7B.
• The model uses 4-bit QLoRA on a single 4090 GPU.
• The results are available in the linked Github repository.
🔗 Resources:
• Arithmo-Mistral-7B Details ↗ - More details on the model
• Arithmo-Mistral-7B Github ↗ - Github repository
Image
Image
🤖 Reward Model Training - Comparison with Prior Work
This article discusses the conceptual differences between a specific reward model training approach and the work presented in https://arxiv.org/abs/2110.14168 ↗. The key difference lies in training a per-"reasoning step" model versus a per-token reward model.
Key Points:
• Compares a new reward model training approach with prior work.
• Highlights the key difference in training a per-reasoning step model.
• The discussion focuses on the conceptual distinctions between the two methods.
🔗 Resources:
• Related Work ↗ - Prior work on reward model training
🤖 Robotics - RSS2025 Conference Attendance and Talk Announcement
This article announces the author's attendance at the RSS2025 conference in Los Angeles and their planned presentation on Gemini Robotics. The author also mentions participation in the SemRob and RoboEval workshops.
Key Points:
• Announcement of attendance at RSS2025 in Los Angeles.
• Presentation on Gemini Robotics planned.
• Participation in SemRob and RoboEval workshops.
🤖 Geometric Deep Learning - Book Chapter Release
This article announces the release of the first draft of chapter 'g' of a book on geometric deep learning. The chapter covers graphs, graph neural networks (GNNs), and large language models (LLMs).
Key Points:
• First draft of chapter 'g' on geometric deep learning released.
• Covers graphs, GNNs, and LLMs.
🔗 Resources:
Image
⭐️ Support
If you liked reading this report, please star ⭐️ this repository and follow me on Github ↗, 𝕏 (previously known as Twitter) ↗ to help others discover these resources and regular updates.