Skip to content
Drix10 Blog

Humanity’s Sixth Sense benchmark announced

, 7 items in AI Companies and Ventures, 3 min read

In this digest (7 items)

Humanity’s Sixth Sense (HSS) is a new benchmark for intuitive visual reasoning. It includes 522 open‑ended tasks covering images and video. The tasks test spatial, causal, and social understanding.

Key points

  • Benchmark name: Humanity’s Sixth Sense (HSS)

  • Number of tasks: 522 open‑ended tasks spanning images and video

Sources

Mistral Large 4 model announced with 1T parameters and upcoming open weights

Mistral Large 4 is a 1‑trillion‑parameter model with 49 billion active parameters. Open weights are scheduled to be released at the end of the month. The model was run through real‑world tests and compared with leading proprietary models.

Key points

  • Model size: 1 trillion parameters (49 billion active)

  • Open weights release: end of the month

Sources

MiniMax AI M3 Prompt Caching Live on SambaCloud

MiniMax AI M3 now supports prompt caching on SambaCloud. Cached prefixes reduce time‑to‑first‑token by 35–88% and increase throughput up to 4.7×. Input cost drops about 90% and cached tokens cost $0.06 per million. No code changes are needed.

Key points

  • Performance: TTFT improves 35–88% and speed up to 4.7×.

  • Cost: Input cost lowers 90% and cached tokens cost $0.06 per million.

Sources

CoreWeave launches Agent Lens for full‑corpus trace analysis

CoreWeave introduced Agent Lens, a tool that logs an agent’s traces and reads the entire production conversation corpus. It groups conversations that fail for the same reason. The post cites 40,000 overnight conversations with 230 failures that went unnoticed.

Key points

  • Agent Lens logs traces and processes the full production corpus, not a sample.

  • It groups conversations that fail for the same reason.

Sources

Replit builds on Windows using Microsoft Execution Containers and Nvidia OpenShell

Replit now builds and runs apps locally on Windows. Each build runs in its own sandbox. The sandbox is powered by Microsoft Execution Containers and Nvidia OpenShell.

Key points

  • Local Windows builds: Replit builds and runs apps on Windows.

  • Sandbox technology: Builds run in sandboxes using Microsoft Execution Containers and Nvidia OpenShell.

Sources

Google indexing timelines – discovery, indexing, core update recovery

Google published timing data for SEO processes. New URL discovery takes about 20 hours or may never happen. Full indexing ranges from roughly 1.5 hours up to several months, and recovering from a core update takes 3 to 6 months.

Key points

  • New URL discovered: ~20 hours, or never

  • Indexing end‑to‑end: ~1.5 hours to months

  • Core update recovery: 3 to 6 months

Sources

Liquid AI launches Open d1 decision models

Liquid AI released two open-weight decision models called Open d1. d1-3B supports text and vision, and d1-omni-600M supports text plus image or text plus audio. The models read a state plus typed questions and return calibrated probabilities in one forward pass.

Key points

  • Model: d1-3B (text + vision)

  • Model: d1-omni-600M (text + image or text + audio)

Sources