🤖 Fine-tuning LLMs - GRPO with SmolLM
This article details a notebook demonstrating the fine-tuning of a small language model (SmolLM-135M) using Gradient-based Reward Optimization (GRPO) and a filtered smoltldr dataset. The goal is to generate concise summaries ("TL;DR").
Key Points:
• Fine-tuning a small LLM model is efficient.
• GRPO enhances model output quality.
• The filtered smoltldr dataset is used for training.
🔗 Resources:
• Pranab Kumar's Twitter ↗ - Lead researcher
• Maxime Labonne's Twitter ↗ - Contributor
• Hugging Face ↗ - Model hosting platform
• Ben Burtenshaw's Twitter ↗ - Contributor
Image
🚀 Kitware Hiring - Computer Vision and AI Roles
Kitware is recruiting for roles in AI, machine learning, computer vision, and NLP. They are attending WACV 2025 and inviting attendees to visit their booth.
Key Points:
• Opportunities in AI, machine learning, computer vision, and NLP.
• Booth 202 at WACV 2025.
• Application form available online.
🔗 Resources:
• WACV 2025 ↗ - Conference
• Kitware ↗ - Company website
• Application Form ↗ - Job application
Image
🚀 AI Agent Evaluation - Agent Leaderboard
This article introduces the Agent Leaderboard, a tool designed to evaluate Large Language Models (LLMs) based on their performance in real-world tool-calling tasks crucial for AI agent development.
Key Points:
• Provides insights into LLM performance for AI agents.
• Evaluates LLMs on real-world tool-calling tasks.
• Offers comparative results for various LLMs.
🔗 Resources:
• Clement Delangue's Twitter ↗ - Relevant context
• Omar Sar0's Twitter ↗ - Relevant context
• RunGalileo ↗ - Agent Leaderboard creator
Image
✨ Axelera AI at MWC 2025 - Edge AI
Axelera AI is showcasing their Edge AI solutions at Mobile World Congress (MWC) 2025. Jean Vieville will participate in an expert panel.
Key Points:
• Presence at MWC 2025, Hall 8.1 – 4YFN – Stand 8.1B52
• Jean Vieville participating in expert panel.
• Focus on Edge AI technologies.
🔗 Resources:
• Axelera AI ↗ - Company website
Image
💡 LLM Evaluation Metrics - Current Challenges
This article discusses the challenges in evaluating Large Language Models (LLMs), highlighting the limitations of existing metrics like MMLU and SWE-Bench.
Key Points:
• Existing metrics like MMLU are outdated.
• SWE-Bench is too narrow in scope.
• A broader, more comprehensive evaluation framework is needed.
🚀 Prompt Engineering Book - EPub Release
This article announces the release of a Prompt Engineering book in EPUB format, compatible with various ebook readers.
Key Points:
• Available in EPUB format.
• Compatible with Kindle, iPad, smartphones, and other e-readers.
• 204 purchases in 3 days.
🤖 Muon Optimizer - Theory and Implementation
This article points to resources explaining the theory and implementation of the Muon optimizer, a new optimization technique.
Key Points:
• Detailed explanation of Muon optimizer theory.
• Links to related work on Modula and modular duality.
• Demonstrated effectiveness in practice.
🔗 Resources:
• Phillip Isola's Twitter ↗ - Relevant context
• Modula Documentation ↗ - Background information
• Jeremy Howard's Twitter ↗ - Relevant context
Image
🤖 Magma Codebase - Transformers Library Integration
This article discusses the Magma codebase, built upon the transformers library, focusing on its features for efficient LLM usage.
Key Points:
• Uses the transformers library extensively.
• Supports elastic image token insertion.
• Allows easy switching between different transformers LLMs.
🔗 Resources:
• Mu Cai's Twitter ↗ - Relevant context
• JW2Yang4AI's Twitter ↗ - Relevant context
• Merve Noyann's Twitter ↗ - Relevant context
Image
💡 Critical Thinking - Beyond Team Loyalty
This article encourages critical thinking and independent judgment, even within one's own group or organization.
Key Points:
• Blind agreement is counterproductive.
• Critical evaluation of one's own perspectives is important.
• Independent thought is essential for problem-solving.
⭐️ Support
If you liked reading this report, please star ⭐️ this repository and follow me on Github ↗, 𝕏 (previously known as Twitter) ↗ to help others discover these resources and regular updates.