Agent Plasticity is introduced as a metric to measure how efficiently agents improve on held-out environments. The study focuses on agents that amortize experience into reusable artifacts. It notes that the top performer ≠ most efficient learner.
Key points
Amortization: agents convert experience into reusable artifacts
Metric: Agent Plasticity measures efficiency on held-out environments
Comparison: top performer ≠ most efficient learner
Sources
Benchmark Designers Should "Train on the Test Set” Poster at COLM2026
The poster titled “Benchmark Designers Should ‘Train on the Test Set’ to Expose Exploitable Non-Visual Shortcuts” will be presented at COLM2026. It warns that benchmarks can be gamed and suggests gaming them first. The presentation is scheduled for Poster #124 in the Grand Ballroom at 4:40 pm.
Key points
Poster #124 will be shown in the Grand Ballroom at 4:40 pm.
The poster’s claim: “if your benchmark can be gamed, it will be… so you’d better game it first!”
Sources
Updated paper and code released for NYU vision collaboration
The collaboration with @jihanyang13, @shushengyang, @rob_fergus, and @sainingxie at NYU released an updated paper. Full code for the work is available at the provided website.
Key points
Paper: Updated version released.
Code: Full implementation accessible at https://vision-x-nyu.github.io/test-set-training.
Sources
- Original post
- Linked resource - Linked in the post
Odyssey-3 foundation world model launched
Odyssey-3 is a foundation world model that generates embodied environments in real time for policy training. It predicts physics and action effects instantly. The team reports a new state‑of‑the‑art result on a physics benchmark and first place in three WorldMark categories.
Key points
Launch: Odyssey-3 foundation world model released.
Benchmark: New SOTA on Physics benchmark.
Ranking: Ranked 1st in three WorldMark categories.
Sources
Level-of-Token Diffusion announced for adaptive image/video generation
Level-of-Token (LoT) Diffusion is a framework that allocates finer tokens to detailed regions and coarser tokens elsewhere for image and video generation. It adapts pretrained diffusion transformers to multiresolution token layouts, reducing token sequence length while preserving full‑resolution flow prediction. The announcement includes a link to a demo site with an arcade‑style interface.
Key points
LoT Diffusion uses rectangular patches of varying sizes as tokens, assigning finer tokens where detail is needed.
The method fine‑tunes pretrained diffusion transformers with patch‑wise asymmetric flow matching to recover dense velocity fields from the reduced token sequence.
Sources
- Original post
- Level-of-Token Diffusion - Linked page

Google EmbeddingGemma 2 finds video from single sentence
Google’s EmbeddingGemma 2 can locate a video using one sentence. In a test, “A cat in a hat” returned the top result out of 40. A reverse‑order test on four clips showed the model’s forward and backward scores were 99.2–99.96% identical, but only one of four guesses was correct.
Key points
Test result: “A cat in a hat” ranked #1 of 40 videos.
Reverse test: forward vs backward similarity 99.2–99.96%, correct guesses 1/4.

