Google Flow Music is an AI music platform that combines generation, conversational editing, and instrument/effect building in a single workspace. The service was previously called Producer AI. It is available via free and premium plans and includes watermarking and modular tool features.
Key points
Rename: Google Flow Music is the new name for Producer AI.
Features: The platform offers music generation, conversational editing, and tools for building instruments and effects in one workspace.
Sources
SteerablePlex full-duplex model steering announced
SteerablePlex is a full-duplex speech model that can be steered with textual instructions. The authors introduced SimIF-Bench to evaluate scenario adherence and a GDPO-based training recipe to improve control. They connected the model to an asynchronous backend language model for more reliable multi‑stage constraint following.
Key points
SimIF-Bench evaluates whether a conversational model stays within a prescribed scenario and completes multiple goals in order.
GDPO-based training enables a full-duplex model to follow textual instructions while maintaining turn‑taking ability.
Sources
- Original post
- SteerablePlex: Can We Steer Full-Duplex Models? - Linked page
Speech-Rewarded Style Planning (SRSP) for Conversational TTS
SRSP is a text‑based style planner trained through a frozen downstream TTS model. It generates style instructions from dialogue history and response text and is optimized with group‑relative policy optimization using teacher‑forced likelihood of target speech tokens as reward. Experiments on an English subset of the ISCSLP 2026 TTS corpus show SRSP improves speech‑style and emotion similarity and reduces mel‑cepstral distortion compared to baseline methods.
Key points
Method: SRSP uses group‑relative policy optimization (GRPO) with teacher‑forced likelihood as reward.
Results: SRSP achieves higher speech‑style and emotion similarity and lower mel‑cepstral distortion than the Base LLM and target‑audio‑informed captioning baselines.
Sources
SepGen model announced for audio‑video stem generation and separation
SepGen extends a pretrained audio‑video generator to output video, mixed soundtrack, and individual waveforms per captioned source in one sampling run. It operates in generation mode, creating stems from source captions, and separation mode, extracting stems from a clean mix using free‑text captions. The paper reports that SepGen outperforms language‑conditioned separators on speech tasks and provides code, checkpoints, and datasets.
Key points
SepGen supports two modes: generation (caption‑driven stem creation) and separation (caption‑driven extraction).
SepGen outperforms language‑conditioned separators, especially on speech, according to the evaluation.
Sources
DuplexAgent-RSI announced for full‑duplex voice agent collaboration
DuplexAgent-RSI is a system that coordinates full‑duplex voice agents with asynchronous search, reasoning, and coding agents. It expresses the workflow as six editable modules and uses a closed‑loop improvement process that revises them from interaction traces. Experiments on several benchmarks report stronger spoken‑knowledge and executable‑tool scores than compared delegated systems while keeping interruption response strong.
Key points
Six editable modules form the harness workflow.
A simulator automatically generates timed test conversations and produces failure traces.
Experiments on intelligence, agentic, and duplex benchmarks show stronger spoken‑knowledge and executable‑tool scores than compared delegated systems.
Sources
Accurate speaker inventory method for online diarization
A new method improves speaker count accuracy in online diarization. It changes registration triggers and uses a candidate pool with commit‑gated label assignment. The approach reduces macro speaker‑count error from 225 % to 21 % and macro DER from 17.59 % to 15.95 % with 0.109 s added latency.
Key points
Macro speaker‑count error reduced from 225 % to 21 % of reference.
Macro DER reduced from 17.59 % to 15.95 % across nine datasets.