Skip to content
Drix10 Blog

Google Flow Music (formerly Producer AI) launch

, 6 items in AI Generated Music and Audio, 4 min read

In this digest (6 items)

Google Flow Music is an AI music platform that combines generation, conversational editing, and instrument/effect building in a single workspace. The service was previously called Producer AI. It is available via free and premium plans and includes watermarking and modular tool features.

Key points

  • Rename: Google Flow Music is the new name for Producer AI.

  • Features: The platform offers music generation, conversational editing, and tools for building instruments and effects in one workspace.

Sources

SteerablePlex full-duplex model steering announced

SteerablePlex is a full-duplex speech model that can be steered with textual instructions. The authors introduced SimIF-Bench to evaluate scenario adherence and a GDPO-based training recipe to improve control. They connected the model to an asynchronous backend language model for more reliable multi‑stage constraint following.

Key points

  • SimIF-Bench evaluates whether a conversational model stays within a prescribed scenario and completes multiple goals in order.

  • GDPO-based training enables a full-duplex model to follow textual instructions while maintaining turn‑taking ability.

Sources

Speech-Rewarded Style Planning (SRSP) for Conversational TTS

SRSP is a text‑based style planner trained through a frozen downstream TTS model. It generates style instructions from dialogue history and response text and is optimized with group‑relative policy optimization using teacher‑forced likelihood of target speech tokens as reward. Experiments on an English subset of the ISCSLP 2026 TTS corpus show SRSP improves speech‑style and emotion similarity and reduces mel‑cepstral distortion compared to baseline methods.

Key points

  • Method: SRSP uses group‑relative policy optimization (GRPO) with teacher‑forced likelihood as reward.

  • Results: SRSP achieves higher speech‑style and emotion similarity and lower mel‑cepstral distortion than the Base LLM and target‑audio‑informed captioning baselines.

Sources

SepGen model announced for audio‑video stem generation and separation

SepGen extends a pretrained audio‑video generator to output video, mixed soundtrack, and individual waveforms per captioned source in one sampling run. It operates in generation mode, creating stems from source captions, and separation mode, extracting stems from a clean mix using free‑text captions. The paper reports that SepGen outperforms language‑conditioned separators on speech tasks and provides code, checkpoints, and datasets.

Key points

  • SepGen supports two modes: generation (caption‑driven stem creation) and separation (caption‑driven extraction).

  • SepGen outperforms language‑conditioned separators, especially on speech, according to the evaluation.

Sources

DuplexAgent-RSI announced for full‑duplex voice agent collaboration

DuplexAgent-RSI is a system that coordinates full‑duplex voice agents with asynchronous search, reasoning, and coding agents. It expresses the workflow as six editable modules and uses a closed‑loop improvement process that revises them from interaction traces. Experiments on several benchmarks report stronger spoken‑knowledge and executable‑tool scores than compared delegated systems while keeping interruption response strong.

Key points

  • Six editable modules form the harness workflow.

  • A simulator automatically generates timed test conversations and produces failure traces.

  • Experiments on intelligence, agentic, and duplex benchmarks show stronger spoken‑knowledge and executable‑tool scores than compared delegated systems.

Sources

Accurate speaker inventory method for online diarization

A new method improves speaker count accuracy in online diarization. It changes registration triggers and uses a candidate pool with commit‑gated label assignment. The approach reduces macro speaker‑count error from 225 % to 21 % and macro DER from 17.59 % to 15.95 % with 0.109 s added latency.

Key points

  • Macro speaker‑count error reduced from 225 % to 21 % of reference.

  • Macro DER reduced from 17.59 % to 15.95 % across nine datasets.

Sources