Humanity’s Sixth Sense (HSS) is a new benchmark for intuitive visual reasoning. It includes 522 open‑ended tasks covering images and video. The tasks test spatial, causal, and social understanding.
Key points
Benchmark name: Humanity’s Sixth Sense (HSS)
Number of tasks: 522 open‑ended tasks spanning images and video
Sources
Mistral Large 4 model announced with 1T parameters and upcoming open weights
Mistral Large 4 is a 1‑trillion‑parameter model with 49 billion active parameters. Open weights are scheduled to be released at the end of the month. The model was run through real‑world tests and compared with leading proprietary models.
Key points
Model size: 1 trillion parameters (49 billion active)
Open weights release: end of the month
Sources
MiniMax AI M3 Prompt Caching Live on SambaCloud
MiniMax AI M3 now supports prompt caching on SambaCloud. Cached prefixes reduce time‑to‑first‑token by 35–88% and increase throughput up to 4.7×. Input cost drops about 90% and cached tokens cost $0.06 per million. No code changes are needed.
Key points
Performance: TTFT improves 35–88% and speed up to 4.7×.
Cost: Input cost lowers 90% and cached tokens cost $0.06 per million.
Sources
- Original post
- Linked resource - Linked in the post
CoreWeave launches Agent Lens for full‑corpus trace analysis
CoreWeave introduced Agent Lens, a tool that logs an agent’s traces and reads the entire production conversation corpus. It groups conversations that fail for the same reason. The post cites 40,000 overnight conversations with 230 failures that went unnoticed.
Key points
Agent Lens logs traces and processes the full production corpus, not a sample.
It groups conversations that fail for the same reason.
Sources
Replit builds on Windows using Microsoft Execution Containers and Nvidia OpenShell
Replit now builds and runs apps locally on Windows. Each build runs in its own sandbox. The sandbox is powered by Microsoft Execution Containers and Nvidia OpenShell.
Key points
Local Windows builds: Replit builds and runs apps on Windows.
Sandbox technology: Builds run in sandboxes using Microsoft Execution Containers and Nvidia OpenShell.
Sources
Google indexing timelines – discovery, indexing, core update recovery
Google published timing data for SEO processes. New URL discovery takes about 20 hours or may never happen. Full indexing ranges from roughly 1.5 hours up to several months, and recovering from a core update takes 3 to 6 months.
Key points
New URL discovered: ~20 hours, or never
Indexing end‑to‑end: ~1.5 hours to months
Core update recovery: 3 to 6 months
Sources
Liquid AI launches Open d1 decision models
Liquid AI released two open-weight decision models called Open d1. d1-3B supports text and vision, and d1-omni-600M supports text plus image or text plus audio. The models read a state plus typed questions and return calibrated probabilities in one forward pass.
Key points
Model: d1-3B (text + vision)
Model: d1-omni-600M (text + image or text + audio)
